LLM Routing: How It Works and What the Benchmarks Show
LLM routing sends each request to the model best suited to it. The technique is real and the savings can be real, but the newest benchmark finds several commercial routers fail to beat a simple baseline. Here is how routing works and where it stops helping.

Key takeaways
- LLM routing is a prediction made before inference about which model in a pool can answer a given request acceptably at the lowest cost.
- Routers decide four ways: rules on request metadata, a trained classifier, a cascade that escalates after a cheap attempt, and semantic similarity to previously seen queries.
- The peer-reviewed RouteLLM abstract claims cost reduction of "over 2 times in certain cases." The widely quoted 85 percent figure lives in the project's GitHub README, not the paper.
- LLMRouterBench, the largest public routing benchmark, found that several recent approaches including commercial routers fail to reliably outperform a simple baseline.
- Routing lowers the price of the inference you run. It does not lower how much inference you run.
Last updated: 16 September 2026.
What LLM routing actually is
LLM routing is a decision made before inference. Given an incoming request and a pool of two or more models at different price and capability tiers, a routing policy predicts which model will produce an acceptable answer for that specific request at the lowest cost, then sends the request only there.
Three properties of that definition do real work. The decision is per request, so two users asking different things in the same session can land on different models. The decision is a prediction, which means it can be wrong, and the cost of being wrong is asymmetric: routing a hard query to a cheap model produces a bad answer, while routing an easy query to an expensive model only produces a bad invoice. And the decision is made from a pool, which means routing only exists when you have already accepted the operational cost of depending on more than one model.
IBM Research's 2024 write-up on routers draws the cleanest available line through the design space, splitting routers into non-predictive ones that run inference on several models and pick the best result, and predictive ones that decide from information gathered before any inference happens (research.ibm.com). Almost everything sold as a router today is predictive. The framing is two years old now and predates every current model tier, but the distinction has held up better than the pricing in it.
Why teams reach for routing now
The spread between tiers inside a single vendor's lineup is wide enough to be worth engineering against. Anthropic's published rate card puts Haiku 4.5 at $1 per million input tokens and $5 per million output, Sonnet 5 at $2 and $10, and Opus 5 at $5 and $25 (claude.com/pricing). Google's spread is wider still: Gemini 2.5 Flash-Lite lists at $0.10 input and $0.40 output, against $2.00 and $12.00 for Gemini 3.1 Pro under 200k context (ai.google.dev). Twenty times, inside one API.
The second condition is that production traffic is rarely uniformly hard. A support deployment answers a long tail of genuinely difficult questions and a fat head of "where is my invoice." A coding assistant handles architecture questions and import fixes through the same endpoint.
Here is one decision worked all the way through, with real rate cards. Say you have a ticket-classification step: 800 input tokens of ticket text and instructions, 20 output tokens of label, two million calls a month. On Opus 5, that is 1,600 million input tokens at $5 and 40 million output tokens at $25, so $8,000 plus $1,000, or $9,000 a month. On Haiku 4.5 the same volume costs $1,600 plus $200, or $1,800. A delta of $7,200 a month on one step.
That number is arithmetic on two published price lists, not a benchmark result, and it holds only on one condition: that Haiku actually labels those tickets at parity with Opus on your taxonomy. Nobody can tell you that from outside your data. It is a two-day eval, and it is the eval that determines whether the rest of this article matters to you. Classification with a fixed label set is the friendliest possible case for routing, because correctness is checkable against ground truth you already have in your ticket history. Most routing decisions are not that clean.
The four ways routers decide
| Approach | How it decides | What it costs you | When it fits | |---|---|---|---| | Rule and metadata | Hand-written conditions on properties you already know: endpoint, customer tier, prompt length, requested output format | Engineering time to maintain rules, and drift as traffic changes | Traffic segments are already known and structurally distinct | | Classifier | A small trained model scores predicted difficulty or predicted win rate, then thresholds it | Training data, retraining when models change, plus its own inference hop | High volume, heterogeneous traffic, and labelled history to train on | | Cascade | Tries the cheap model first, evaluates the answer, escalates on failure | Double inference on every escalated query, and a hard dependency on the quality check | Verifiable outputs, where a cheap self-check or validator is trustworthy | | Semantic | Embeds the request and routes by similarity to a library of past queries with known good handlers | Embedding cost per request, plus maintenance of the reference library | Repetitive traffic clustered around a stable set of intents |
Each of these is a strategy with its own tuning surface, failure behaviour, and evaluation method. We have written up the routing strategies in detail separately, so the paragraphs below are deliberately brief.
Rule and metadata routing
Conditions written by hand against properties available before the model sees anything. This is the approach teams underrate because it looks unsophisticated, and it is the one the robustness literature likes most, since a router with no learned parameters has nothing for an attacker to manipulate.
Classifier routing
A small model, often a fine-tuned encoder, predicts whether the cheap model will succeed. This is where the published research concentrates, and it is the approach that most needs your own traffic to train on.
Cascade routing
Answer cheaply, check the answer, escalate if the check fails. The economics are entirely determined by escalation rate and by whether your check is better than a coin flip. A cascade with a weak verifier pays twice for the same wrong answer.
Semantic routing
Embed the query and route by nearest neighbour in a curated library. Fast and interpretable when traffic clusters. It degrades quietly when a new intent appears that resembles nothing in the library.
What the research actually found
What RouteLLM showed
RouteLLM, from Ong and colleagues at Berkeley's Sky Computing Lab, trained routers on human preference data from Chatbot Arena and added data augmentation to improve them. The arXiv abstract claims cost reduction "over 2 times in certain cases" without compromising response quality, and reports that the routers transfer, holding performance when the strong and weak models are swapped at test time (arXiv:2406.18665). The abstract names no specific benchmark for the cost figure.
The number you have probably seen quoted is different. "Reduce costs by up to 85%" while maintaining "95% GPT-4 performance," attributed to benchmarks like MT Bench, is a bullet point in the project's README (github.com/lm-sys/RouteLLM). It is a legitimate reported result from the project. It is also not what the paper claims, and vendor pages routinely cite it as though it were. Anyone building a business case on 85 percent should know which artifact it came from.
What LLMRouterBench found
The most rigorous public evaluation of routing arrived in July 2026. Li and colleagues built LLMRouterBench across more than 400,000 instances from 21 datasets and 33 models, and ran 10 representative routing baselines through a single unified evaluation (Findings of ACL 2026).
Four of their findings should change how you plan a routing project.
The premise of routing holds. They confirm strong model complementarity, which is to say different models really do win on different queries, and there is genuine headroom for a router to capture.
The methods are less differentiated than the literature implies. Under unified evaluation, many routing methods perform similarly. The specific technique you pick matters less than the fact that you routed at all.
Several recent approaches, commercial routers included, fail to reliably outperform a simple baseline. That is the sentence to read twice. It does not say routing fails. It says that buying a sophisticated router is not a reliable improvement over routing crudely, and that the burden of proof sits with the vendor.
And a substantial gap remains to the Oracle, the hypothetical perfect router, driven mainly by persistent model-recall failures. The available upside is real and largely uncaptured. They also report that the choice of backbone embedding model has limited effect, and that enlarging the model pool shows diminishing returns against careful curation of which models are in it.
Read together, the two papers give an honest picture that no page currently ranking for this term presents. Routing earns money in measured cases. Which router you use is weakly determined by the research. Treat it as an experiment on your own traffic with a control arm, not as a discount you procure.
Routing versus an LLM gateway
These three get conflated in procurement conversations, and they solve unrelated problems.
| Mechanism | What it decides | Primary purpose | |---|---|---| | Routing | Which model answers this request | Cost and quality per request | | Load balancing | Which instance, region, or key serves this request | Throughput and rate-limit headroom | | Gateway | Nothing about model choice by itself | One API surface with key management, logging, retries, and fallbacks |
The practical consequence: a gateway is where a router gets deployed, and most gateway products ship some routing, but adopting a gateway does not mean you are routing. If every request still goes to the same model through a unified API, you bought observability and key hygiene. Both are worth having. Neither moves your bill. For readers comparing products rather than techniques, we maintain a survey of gateways and routers on the market.
Where routing stops helping
A predictive router adds a hop before every request. A classifier or embedding step is cheap in absolute terms and still lands on the critical path of every user-facing call. A cascade is worse, because escalated queries pay full latency twice.
The router is also a new component that can fail or misclassify, and its failures are quieter than an outage. Quality drift shows up as a slow increase in the share of hard queries answered by the cheap model, which looks like nothing on a cost dashboard and looks like everything to the affected users. Any routing deployment needs a quality metric watched as closely as the spend metric.
Then there is the attack surface, which is specific to learned routers. Researchers have shown that query-independent token sequences, appended to any input, can force a router to select the expensive model, working in both white-box and black-box settings against open-source and commercial routers (arXiv:2501.01818). A 2026 follow-up optimises adversarial suffixes against black-box routers through a surrogate ensemble (arXiv:2604.15022). A systematisation of the threat categories reports that real routing systems are most vulnerable to cost escalation specifically (arXiv:2601.21380). A broader robustness study found DNN-based routers weakest and training-free routers most robust, since they have no learned parameters to manipulate (arXiv:2503.08704). If your router faces untrusted input and your savings thesis depends on most traffic staying cheap, an adversary can delete the thesis.
The ceiling is the part worth sitting with. Routing changes which model answers a request. It has no opinion about how many requests exist. A workflow that makes forty model calls per run makes forty routed calls per run. Route perfectly and you have bought a discount on a volume you never questioned. Routing sits below the orchestration layer above routing, and it is one lever among the wider set of cost levers rather than the whole answer.
This article does not cover router evaluation methodology, which is the harder problem and deserves its own treatment. It also does not address multi-objective routing where latency budgets and compliance constraints compete with cost.
What we think, and what we build
Our position: routing is a real optimisation and a bounded one. It cannot stop a decision your system made last week from being made again from scratch today. Teams reach for a router because the bill is growing, and the bill is usually growing because the same reasoning is being repeated across thousands of runs. No routing policy addresses repetition. It prices repetition more cheaply.
So we built for the other half. Major is the enterprise platform where agents build the software they run on. When one of our agents works out how to handle a repeatable step, it builds an app that performs that step in deterministic code, with its own managed database, storage, and logs, and from then on it runs the app instead of reasoning through the step again. The model stays in the loop for the judgment calls. It is doing less of the work, and the work it stopped doing generates no tokens to route. That step also becomes inspectable, because it now lives in code with permissions and an audit trail rather than inside a prompt. Reason once, run forever.
The two levers are complements, and we would rather say that plainly than pretend otherwise. Routing lowers the price of the inference you do. Moving repeatable work into apps lowers the amount of inference you need. Use both. The honest scoping: this only pays off where the work genuinely repeats. A workflow that is different every run has nothing to push down into code, and for that workflow a good router is the better investment. Most enterprise workflows are not that workflow, which is why we think the app layer is where the larger saving lives and why coordinating models, agents and the app layer is the design problem worth your attention.
If your routing project is really a cost project, it is worth checking which of your model calls are decisions and which are repetitions you are paying to re-derive. You can see how Major turns the repeatable ones into deterministic apps.
Related articles
Frequently asked questions
- What is LLM routing?
- LLM routing is a decision made before inference. Given an incoming request and a pool of models at different price and capability tiers, a routing policy predicts which model can answer that specific request acceptably at the lowest cost, then sends the request only to that model. The decision is per request and can be wrong.
- What is the difference between LLM routing and an LLM gateway?
- Routing decides which model answers a given request. A gateway provides one API surface across providers, handling key management, logging, retries, and fallbacks. The practical consequence: a gateway is where a router gets deployed, but adopting a gateway does not mean you are routing. If every request still reaches the same model, you gained observability without changing spend.
- Does LLM routing actually reduce costs?
- In measured cases, yes. The RouteLLM paper reports cost reduction of over two times in certain cases. But LLMRouterBench, which evaluated 10 routing baselines across 400,000-plus instances, found several recent approaches including commercial routers fail to reliably outperform a simple baseline. Treat routing as an experiment on your own traffic with a control arm, not a procured discount.
- When should I use per-query routing instead of fixed rules?
- Per-query routing earns its complexity when traffic is heterogeneous in ways your metadata cannot see, volume is high enough that a few percentage points of cost matter, and you have labelled history to train on. When your traffic already splits along known structural lines, such as endpoint or customer tier, rules are cheaper, faster, and harder to attack.
- What are the risks of LLM routing?
- The router adds a latency hop to every request and becomes a component that can fail or misclassify. Quality drift is the quiet risk: more hard queries answered by the cheap model, invisible on a cost dashboard. Learned routers are also an attack surface, with published work showing adversarial inputs that force selection of the expensive model.