Test-time compute scaling buys accuracy by spending more inference on a hard problem instead of training a larger model. Sampling several answers, verifying them, and searching over reasoning paths all trade tokens for correctness. Every method is capped by whether the base model can solve the problem at all.
The most surprising result in this area is how far a modest model gets when allowed to try repeatedly and check its work. On problems with verifiable answers, repeated sampling with selection closes a meaningful part of the gap to a much larger model.
The equally important result is where that stops working, which is the part most coverage omits.
| Method | Mechanism | Needs | Cost shape |
|---|---|---|---|
| Chain of thought | Generate intermediate steps before answering | Nothing extra | Longer single generation |
| Self-consistency | Sample n answers, take the majority | A comparable answer format | n generations |
| Best-of-n with a verifier | Sample n, score each, pick the best | A verifier or reward model | n generations plus n scorings |
| Deterministic checking | Sample until output passes a hard check | A machine-checkable answer | Variable, often cheapest |
| Tree search over steps | Explore and prune partial reasoning paths | Step-level scoring | Highest, and most complex |
The row that pays best in practice is deterministic checking, because the verifier is free and perfectly reliable. If your task has a machine-checkable answer, this dominates every other method on cost per unit of accuracy gained.
Why does sampling work at all?
Because generation is stochastic, so a model that fails a problem 60% of the time still solves it sometimes.
Wang et al. showed that sampling several chain-of-thought paths and taking the majority answer improves reasoning accuracy substantially over a single greedy generation (arXiv:2203.11171). The mechanism is straightforward: errors are varied and distributed, while correct answers agree with each other.
That mechanism also explains the limit. Majority voting only works when correct answers converge and wrong answers scatter. On problems where the model is systematically wrong in a consistent way, every sample agrees and the majority is confidently incorrect.
Figure 1 — When sampling helps and when it cannot
Model is capable but unreliable
It reaches the right answer some fraction of the time and errs differently on each attempt.
Correct answers cluster, wrong ones scatter, so majority voting or verification recovers the right one.
Compute buys accuracy
Model is systematically wrong
It misunderstands the problem the same way on every attempt, or lacks the knowledge entirely.
All samples agree on the wrong answer. Sampling a hundred times produces a hundred consistent errors and high apparent confidence.
Compute buys nothing
This is the ceiling that defines the technique. Test-time compute amplifies capability the model already has; it does not create capability it lacks. No inference budget substitutes for a base model that cannot reach the answer.
Where do the returns run out?
Gains are steep at first and flatten quickly. Moving from one sample to a handful typically delivers most of the available improvement, and each subsequent doubling delivers progressively less.
Two things determine where your curve flattens. The first is the base model’s per-attempt success rate on the task: if that is near zero, no amount of sampling helps. The second is verifier quality: with a perfect verifier, more samples keep helping for longer, while with a noisy one additional samples eventually introduce more selection error than they remove.
Watch for this
A weak verifier can make more sampling actively worse. If the scorer systematically prefers a certain style of wrong answer, generating more candidates gives it more opportunities to select one. Measure verifier accuracy independently before scaling n, and treat a noisy verifier as a reason to keep n small.
How does this compare with training a bigger model?
They trade along different axes, and the right choice depends on volume and task shape.
- Test-time compute is per-query. Costs scale with usage, forever, and can be applied selectively to hard queries only.
- A bigger model is a fixed choice. It costs more on every query including the easy ones, but needs no orchestration.
- Test-time compute is targetable. Combined with routing, you spend the budget only where difficulty warrants it, which is where the economics get genuinely favourable.
That third point is the practical synthesis. Test-time scaling and query routing are complementary: route easy queries to a single fast generation, and spend sampling budget on the minority that need it.
What does the latency cost look like?
Worse than the token cost suggests, and this constrains where the technique is usable.
Sequential methods such as tree search multiply latency directly. Parallel sampling can hide much of it, since n generations can run concurrently, but that requires capacity to burst and adds the verification pass afterwards.
For interactive products the practical envelope is small: a handful of parallel samples with a fast verifier. Deep search belongs in asynchronous and offline work where nobody is watching a cursor.
Where do teams go wrong?
Scaling n before checking the base rate
If the model solves a problem type essentially never, sampling is pure expenditure. Measure single-sample accuracy on the task first. A near-zero base rate means the answer is a better model or better context, not more attempts.
Using the generator as its own verifier
A model asked to judge its own output tends to approve it, and its errors correlate with the errors it just made. Where a deterministic check exists, use it. Where one does not, prefer a different model family for scoring.
Applying it uniformly
Sampling every query five times because it helps on hard ones multiplies cost across traffic that never needed it. This technique belongs behind a difficulty gate, not in the default path.
What do experienced teams do differently?
They spend the budget adaptively rather than at a fixed n.
Generate once, check, and stop if the answer passes. Generate more only while the check keeps failing, up to a cap. Because most queries pass on the first attempt, average cost stays close to single-generation cost while hard queries receive the budget they need.
They also log how many attempts each query consumed. That distribution is one of the most useful signals available: a rising average means either the traffic is getting harder or the model has regressed, and it surfaces both far earlier than accuracy metrics do.
A short glossary
- Test-time compute
- Computation spent during inference rather than training, traded for improved accuracy on hard problems.
- Self-consistency
- Sampling several reasoning paths and selecting the majority answer among them.
- Best-of-n
- Generating n candidates and selecting the highest-scoring one using a verifier or reward model.
- Verifier
- A model or deterministic check that scores candidate answers, whose quality caps how far sampling can help.
- Adaptive budget
- Spending additional attempts only while a check keeps failing, rather than at a fixed sample count.
Key takeaways
- Test-time compute buys accuracy by spending inference on hard problems instead of training a larger model.
- Sampling works because errors scatter while correct answers converge, which is also exactly why it fails on systematic errors.
- Every method is capped by the base model: compute amplifies existing capability rather than creating new capability.
- Deterministic checking dominates on cost per unit of accuracy wherever the answer is machine-checkable.
- A weak verifier can make additional sampling actively worse by giving it more chances to select a bad candidate.
- Returns flatten quickly, so measure single-sample accuracy before scaling the sample count.
- Spend adaptively behind a difficulty gate rather than sampling uniformly across all traffic.
Frequently asked questions
What is test-time compute scaling?
Spending additional computation during inference, rather than during training, to improve accuracy on difficult problems. Methods include generating intermediate reasoning steps, sampling several candidate answers, verifying and selecting among them, and searching over partial reasoning paths.
Does sampling more answers always improve accuracy?
No. It helps when the model is capable but unreliable, because correct answers converge while errors scatter. When the model is systematically wrong in a consistent way, every sample agrees on the same wrong answer, and additional sampling produces confident errors rather than corrections.
How many samples should I generate?
Fewer than instinct suggests, and measured rather than guessed. Returns flatten quickly, with most of the gain arriving in the first few samples. Check single-sample accuracy first: a near-zero base rate means sampling cannot help regardless of how many attempts you allow.
What makes a good verifier?
Determinism, wherever possible. A machine-checkable answer such as arithmetic, code that compiles, or output matching a schema gives you a free and perfectly reliable verifier. Model-based verifiers are noisier and cap how far additional sampling can usefully go.
Can test-time compute replace using a larger model?
Partly, and the economics depend on volume. Test-time compute costs per query and can be applied selectively to hard queries only, while a larger model charges more on every query including easy ones. Combining sampling with query routing is usually better than either alone.
Does this work for open-ended generation?
Much less well. The techniques rely on being able to compare candidate answers, which is straightforward when there is a correct answer and difficult when quality is subjective. Verifiable tasks such as mathematics, code and structured extraction benefit most.
What is the latency impact?
Significant, and it constrains where the technique fits. Sequential methods multiply latency directly. Parallel sampling hides some of it but needs burst capacity plus a verification pass. Interactive products can usually afford a handful of parallel samples; deep search belongs in offline work.
Should I use the same model to generate and verify?
Only if no alternative exists. A model judging its own output tends to approve it, and its scoring errors correlate with the generation errors it just made. Prefer a deterministic check, or failing that a different model family for verification.
References
- Wang, X., et al. Self-Consistency Improves Chain of Thought Reasoning in Language Models. arXiv:2203.11171.
- Wei, J., et al. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. arXiv:2201.11903.
- Cobbe, K., et al. Training Verifiers to Solve Math Word Problems. arXiv:2110.14168. Foundational work on verifier-based selection.
