A vertical model does not beat a generalist overall; it beats one on specific axes. Specialists win on domain vocabulary, format conformance and cost per call. They lose on anything outside the domain, and they lose badly. The useful question is which axis your product actually depends on.
The question teams usually ask is “should we build a domain model.” The sharper version is: which of our failures come from the model not knowing our vocabulary, and which come from it not reasoning well? Those need different answers.
General models have closed much of the gap on domain knowledge. What has not closed is the gap on cost, latency and rigid format conformance at high volume.
At a glance
- Domain vocabulary is where specialists still reliably win
- Continued pretraining differs sharply from fine-tuning in cost and effect
- Retrieval closes knowledge gaps far more cheaply than training does
- Specialists degrade sharply just outside their domain boundary
- Evaluation must be built by domain experts, not borrowed from benchmarks
- Maintenance is the cost nobody budgets for
What does a specialist actually do better?
| Axis | Specialist advantage | Why |
|---|---|---|
| Domain vocabulary | Strong | Tokenisation and representation of specialist terms improve with exposure |
| Format conformance | Strong | Rigid output conventions are learned behaviour, which trains well |
| Cost per call at volume | Strong | A smaller model serving one narrow task is far cheaper |
| Latency | Strong | Smaller models generate faster |
| Data residency and control | Strong | Self-hosting removes third-party processing entirely |
| Factual domain knowledge | Weak to moderate | Retrieval supplies this more cheaply and stays current |
| Reasoning within the domain | Weak | Reasoning is general capability; domain training does not add it |
| Anything adjacent to the domain | Very weak | Sharp degradation at the boundary, often silent |
Read the last two rows together, because they are the trap. Teams justify a vertical model on the belief it will reason better about their domain. It will not. Domain training teaches vocabulary and form, not inference.
Which approach, and at what cost?
Figure 1 — Three routes to a domain model
- CHEAPESTRetrieval + promptingDomain knowledge supplied at query time. Days of work. Always current.
- MODERATEAdapter fine-tuningTeaches format and tone on a frozen base. Hundreds of examples.
- EXPENSIVEContinued pretrainingBillions of domain tokens. Shifts vocabulary representation itself.
- RARETrain from scratchAlmost never justified outside frontier labs
Most teams that say “we need a domain model” need the first option. The step from retrieval to continued pretraining is several orders of magnitude in cost, and it buys vocabulary representation rather than knowledge.
Continued pretraining is the only route that genuinely changes how a model represents specialist terms, because it updates the base weights against domain text at scale. That is a real capability, and it requires a corpus most organisations do not have.
Watch for this
Continued pretraining on a narrow corpus causes measurable loss of general capability. The model gets better at your domain and worse at instruction-following, reasoning and everything adjacent. Budget for a general-capability regression suite alongside your domain evaluation, and expect to mix general data back into the training set to limit the damage.
Where does the boundary problem bite?
At the edges of the domain, where nobody tested.
A clinical model handles clinical text well and encounters a billing question, a scheduling request, or a patient asking something conversational. General models degrade gracefully on unfamiliar input. Specialists produce confident domain-shaped answers to questions outside their domain.
This matters because real traffic is never as clean as the training distribution. The practical mitigation is a router: classify whether the request is in-domain, serve the specialist when it is, and fall back to a general model when it is not.
How should you evaluate one?
Not on public benchmarks, which measure general capability on general text.
- Have domain experts write the evaluation. The distinctions that matter in a specialist field are invisible to anyone outside it.
- Include boundary cases deliberately. Requests just outside the domain, where the specialist is most dangerous.
- Run a general-capability regression suite before and after training, and treat degradation as a cost rather than a footnote.
- Compare against a general model with retrieval, not against a general model alone. That is the real alternative and the honest baseline.
Point 4 is the comparison most projects avoid, because a strong general model with good retrieval frequently matches a specialist on quality while costing far less to build and nothing to maintain.
What is the maintenance cost?
Higher than the build, and almost never budgeted.
- Base model drift. Your specialist is pinned to a base version. When better bases ship, you either stay behind or repeat the work.
- Domain drift. Terminology, regulation and practice change, so a model trained on last year’s corpus quietly ages.
- Evaluation upkeep. The expert-authored eval needs expert time to extend, which competes with those experts’ actual jobs.
- Serving infrastructure. Self-hosting is the point for data control, and it is also an ongoing operational commitment.
A general model plus retrieval has none of these. That asymmetry should feature in the decision far more than it usually does.
Where do teams go wrong?
Building for knowledge rather than form
The most common error, and the most expensive. Facts belong in a retrieval index where they can be updated, cited and access-controlled. Training them into weights makes them stale, unattributable and unmodifiable.
Skipping the regression suite
Domain metrics improve, everyone celebrates, and nobody notices that instruction-following degraded until production surfaces it. Measure general capability before and after, every time.
Underestimating corpus requirements
Continued pretraining needs domain text at a scale most organisations discover they lack only after committing. Audit the corpus honestly before choosing the approach, not after.
What do experienced teams do differently?
They stage the decision rather than making it once.
Start with retrieval and prompting, measure, then add an adapter for format and tone if the gap is behavioural, and only consider continued pretraining if vocabulary representation is demonstrably the bottleneck. Each stage is cheap enough to abandon, which keeps the expensive option genuinely optional.
They also keep the general model in production alongside the specialist, routed by an in-domain classifier. That covers the boundary problem structurally rather than hoping traffic stays clean.
A short glossary
- Continued pretraining
- Further training a base model on large volumes of domain text, updating its weights and vocabulary representation.
- Vertical model
- A model specialised to one industry or field rather than intended for general use.
- Boundary degradation
- Sharp quality loss on inputs just outside a specialist’s training domain, often without any signal.
- General-capability regression
- Loss of broad ability such as instruction-following caused by narrow specialisation.
- In-domain classifier
- A router deciding whether a request falls inside the specialist’s competence before dispatching it.
Where to go deeper
Three directions, in order of practical payoff.
Start with the parameter-efficient fine-tuning literature, since adapters are where most domain adaptation should begin and the cost profile is completely different from full training. Then look at retrieval evaluation, because the honest baseline for any vertical model is a general model with good retrieval, and knowing how to measure that fairly settles most of these debates. Finally, read the catastrophic forgetting literature, which explains why specialisation costs general capability and what mixing strategies limit the damage.
The through-line: domain adaptation is a spectrum of increasingly expensive interventions, and most teams should stop earlier along it than they initially plan to.
Key takeaways
- Specialists win on vocabulary, format conformance, cost and latency, not on reasoning.
- Domain training teaches form and terminology; it does not add inference ability.
- Facts belong in retrieval, where they stay current, citable and access-controlled.
- Continued pretraining measurably degrades general capability and needs a regression suite alongside it.
- Specialists fail sharply and silently just outside their domain, so route with an in-domain classifier.
- The honest baseline is a general model with retrieval, not a general model alone.
- Maintenance exceeds build cost, and a general model with retrieval carries almost none of it.
Frequently asked questions
Do domain-specific models beat general models?
On specific axes, yes: domain vocabulary, rigid format conformance, cost per call at volume, latency, and data control. On reasoning and on anything outside the domain, no. The correct question is which axis your product actually depends on, since the answer differs per application.
Should I fine-tune or use retrieval for domain knowledge?
Retrieval, almost always. Facts stored in weights cannot be updated without retraining, cannot be cited, and cannot respect per-user access control. Fine-tuning earns its cost for format, tone and task behaviour, which is a different problem from knowing things.
What is continued pretraining?
Further training a base model on large volumes of domain text, which updates its weights and improves how it represents specialist vocabulary. It is the only route that genuinely changes vocabulary representation, and it requires a corpus at a scale most organisations do not have.
Does specialising a model make it worse at other things?
Yes, measurably. Narrow training degrades instruction-following, general reasoning and adjacent-domain performance. Run a general-capability regression suite before and after, and consider mixing general data back into the training set to limit the loss.
How do I stop a specialist failing outside its domain?
Route around it. Classify whether a request is in-domain and serve the specialist only when it is, falling back to a general model otherwise. Specialists produce confident domain-shaped answers to out-of-domain questions, and that failure is usually silent.
What should I compare a vertical model against?
A strong general model with good retrieval, which is the real alternative. Comparing against a bare general model flatters the specialist and skips the option most likely to match it at a fraction of the cost and with no maintenance burden.
Who should write the evaluation set?
Domain experts. The distinctions that matter in a specialist field are invisible to anyone outside it, and a benchmark borrowed from general evaluation will not capture them. Include boundary cases deliberately, since those are where specialists are most dangerous.
What is the hidden cost of a vertical model?
Maintenance. The model is pinned to a base version that will be superseded, the domain terminology drifts, the expert-authored evaluation needs expert time to extend, and self-hosting is an ongoing operational commitment. A general model with retrieval carries almost none of this.
References
- Hu, E. J., et al. LoRA: Low-Rank Adaptation of Large Language Models. arXiv:2106.09685.
- Gururangan, S., et al. Don’t Stop Pretraining: Adapt Language Models to Domains and Tasks. arXiv:2004.10964.
- Lewis, P., et al. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. arXiv:2005.11401.
- Muennighoff, N., et al. MTEB: Massive Text Embedding Benchmark. arXiv:2210.07316. Evidence that no single model dominates across task types.
