Route by whether the query needs multi-step reasoning, not by how important it feels. Reasoning models spend inference budget generating intermediate steps before answering; fast models respond immediately. Sending everything to the reasoning model wastes money on lookups, and sending everything to the fast one fails on the queries that matter most.
Most production traffic is easy. A minority is genuinely hard. Serving both from one model means paying reasoning prices for trivial questions or accepting reasoning failures on hard ones.
Routing solves this, and it is one of the few optimisations that improves cost and quality simultaneously rather than trading them.
- 2 axesLatency and cost both scale with reasoning depth
- 1 questionDoes answering require intermediate steps?
- 3 tiersFast, standard and reasoning covers most real traffic
- Router costMust stay well below the saving it produces
- EscalationReacting beats predicting on ambiguous queries
What actually distinguishes the two?
Reasoning models are trained to spend additional inference compute producing intermediate reasoning before committing to an answer. That extra computation is where their advantage on multi-step problems comes from.
The trade is direct: more tokens generated, higher cost, longer latency. On a question with no intermediate steps to take, that budget produces nothing except a slower, more expensive version of the same answer.
| Query type | Route to | Why |
|---|---|---|
| Factual lookup from retrieved context | Fast | Extraction, not inference |
| Format conversion or rewriting | Fast | Mechanical transformation |
| Classification into known labels | Fast | Bounded decision space |
| Multi-hop questions across sources | Reasoning | Requires chaining intermediate conclusions |
| Arithmetic and quantitative comparison | Reasoning | Errors compound without explicit steps |
| Code debugging and generation | Reasoning | Benefits from working through state |
| Planning and constraint satisfaction | Reasoning | Multiple interacting requirements |
Read the pattern rather than memorising the rows: the split is between retrieving an answer and deriving one. Anything already present in the context is a fast-model job regardless of how complex the subject matter sounds.
How do you build the router?
Figure 1 — Predictive routing versus reactive escalation
Predict up front
A classifier inspects the query and picks a model before generation begins.
Cheapest when correct, and it must be. A misroute to the fast model produces a wrong answer nobody catches; a misroute upward wastes the saving entirely.
Cheap, brittle
Escalate on failure
Always try the fast model, then detect low confidence or a failed validation and retry with the reasoning model.
Costs two calls on hard queries, but the decision is made with the answer in hand rather than guessed from the question.
Dearer, robust
Use both, in that order. Predictive routing handles the clear cases cheaply; escalation catches what the classifier got wrong. Systems relying on prediction alone fail silently on exactly the queries the classifier found ambiguous.
Signals worth using in the classifier, roughly in order of usefulness: presence of numbers or comparisons, question length and clause count, whether retrieved context already contains a direct answer, task type where the interface knows it, and explicit user selection where the product exposes it.
What makes escalation work?
A detectable failure signal. Without one, escalation cannot trigger.
- Schema violation where structured output was required
- Failed deterministic validation, such as arithmetic that does not check out
- Explicit uncertainty where the model was instructed to signal when unsure
- Self-consistency divergence across a small number of samples
- Guardrail rejection from the output validation layer
The first two are the most reliable because they are deterministic. Build escalation on those before reaching for confidence estimates, which are noisier than they appear.
Watch for this
The router must cost far less than the saving it generates. A classifier that is itself a model call can consume most of the benefit, particularly when routing between two models whose price gap is modest. Measure end-to-end cost including routing overhead, not just the per-query price of the chosen model.
How should you measure whether routing works?
- Route accuracy on a labelled sample. Have humans label 100 queries by whether they needed reasoning, then check what the router chose.
- Cost per query end to end, including router overhead and escalation retries.
- Quality on the hard slice specifically. Aggregate quality hides degradation on the minority of queries routing was meant to protect.
- Escalation rate over time. A rising rate means the classifier is drifting relative to your traffic.
Point 3 is the one that catches bad routing. A router sending everything to the fast model looks excellent on cost and average quality while quietly failing every genuinely hard query, and only a hard-slice metric reveals it.
Where do teams go wrong?
Routing on importance instead of difficulty
A question from an enterprise customer is not thereby a reasoning task. Routing on business importance sends easy questions to expensive models and produces no quality benefit at all.
Building a router before establishing the need
If 95% of traffic is easy and the fast model handles it well, routing adds engineering surface for a modest gain. Measure the distribution of query difficulty first. Sometimes the answer is to use one model and spend the effort elsewhere.
Letting the classifier ossify
Query distributions drift as products change. A router trained on last quarter’s traffic degrades gradually, and because misroutes downward fail silently, nothing announces it. Re-evaluate on a schedule.
What do experienced teams do differently?
They expose the choice to users where the product allows it.
A visible control letting someone request more careful analysis is often more accurate than any classifier, because users know their own intent. It also converts a hidden engineering decision into a transparent product feature, which handles the ambiguous middle far better than guessing.
They also log the routing decision alongside every response. When output quality is questioned later, knowing which model answered turns an unanswerable question into a two-second lookup, and it is the data any future router improvement will need.
A short glossary
- Reasoning model
- A model trained to spend additional inference compute generating intermediate steps before producing a final answer.
- Predictive routing
- Choosing a model from the query alone, before any generation happens.
- Reactive escalation
- Trying a cheaper model first and retrying with a stronger one when a failure signal appears.
- Escalation rate
- The proportion of queries retried at a higher tier, useful as a drift indicator.
- Hard slice
- The subset of evaluation queries genuinely requiring multi-step reasoning, reported separately from aggregates.
The decision summary
| If your situation is | Do this |
|---|---|
| Traffic is overwhelmingly simple lookups | Use one fast model. Skip routing entirely. |
| Clear split between easy and hard, easily classified | Predictive routing on cheap deterministic signals |
| Difficulty is hard to judge from the question | Escalation on validation failure, no classifier |
| Mixed, high volume, cost matters | Predictive routing plus escalation fallback |
| Users know their own intent | Expose the choice in the interface |
| Output is structured and machine-checkable | Escalate on schema or validation failure, which is highly reliable |
Key takeaways
- Route by whether a query needs derivation or merely retrieval, not by how important it seems.
- Reasoning models spend inference budget on intermediate steps, which is wasted on lookups.
- Predictive routing is cheap and brittle; reactive escalation is dearer and robust. Use both in that order.
- Escalation needs a detectable failure signal, and deterministic ones are far more reliable than confidence estimates.
- The router must cost substantially less than the saving it produces, measured end to end.
- Track quality on the hard slice separately, or a lazy router will look excellent while failing the queries that matter.
- Where the product allows it, letting users request deeper analysis beats any classifier.
Frequently asked questions
When should I use a reasoning model?
When answering requires deriving something through intermediate steps: multi-hop questions across sources, arithmetic and quantitative comparison, code debugging, and planning under multiple constraints. If the answer is already present in the supplied context, extraction is the task and a fast model handles it better.
Is routing worth the added complexity?
Only when your traffic genuinely splits. Measure the difficulty distribution first. If the overwhelming majority of queries are simple lookups a fast model handles well, routing adds engineering surface for a modest gain, and the effort is better spent elsewhere.
Should I predict the right model or escalate on failure?
Both, in that order. Predictive routing handles clearly easy and clearly hard queries cheaply. Escalation catches what the classifier got wrong, which matters because misroutes downward produce wrong answers that nothing else in the system will detect.
What signals should trigger escalation?
Deterministic ones first: schema violations, failed arithmetic checks, and guardrail rejections. These are reliable and free. Confidence-based signals such as self-consistency divergence are noisier and should supplement deterministic checks rather than replace them.
How much does the router itself cost?
It must cost far less than it saves, and this is easy to get wrong. A classifier implemented as its own model call can consume most of the benefit, especially when the price gap between tiers is modest. Measure end-to-end cost including routing and retries.
How do I know if my router is misrouting?
Label a sample of 100 queries by hand for whether they needed reasoning, then compare against what the router chose. Also track quality on the hard slice separately, since a router that sends everything to the fast model looks excellent on every aggregate metric.
Should users be able to choose the model?
Where the product allows it, yes. Users know their own intent better than a classifier can infer it, and a visible control for requesting more careful analysis handles the ambiguous middle well. It also turns a hidden decision into a transparent feature.
Does routing need to be re-tuned over time?
Yes. Query distributions drift as products change, and a classifier trained on older traffic degrades gradually. Because misroutes to the cheaper model fail silently, nothing announces the decline, so schedule periodic re-evaluation rather than waiting for a complaint.
References
- Wei, J., et al. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. arXiv:2201.11903.
- Wang, X., et al. Self-Consistency Improves Chain of Thought Reasoning in Language Models. arXiv:2203.11171. Basis for the divergence signal used in escalation.
- Dettmers, T., et al. QLoRA: Efficient Finetuning of Quantized LLMs. arXiv:2305.14314. Source for caution on trusting single benchmark numbers when comparing model tiers.
