October 4, 2026
Scaling Inference

Continual Learning for LLMs: Updating Models Without Catastrophic Forgetting

Continual Learning for LLMs Updating Models Without Catastrophic Forgetting

Training a model on new data degrades what it previously learned, because the same weights encode both. There is no update that touches only the new knowledge. Mitigations reduce the damage without eliminating it, which is why most production systems keep knowledge outside the weights entirely.

The intuition that misleads people is that a model stores facts in separate places, so adding one should not disturb another. Weights are shared and distributed, so every gradient step that improves the new task moves parameters that the old task also depended on.

That is the mechanism. Everything else follows from it.

Common beliefWhat is actually true
New training only adds knowledgeIt redistributes shared parameters, so old capability shifts too
Forgetting shows up in the lossTraining loss falls while unmeasured capabilities degrade invisibly
Small updates are safeSmall narrow updates can cause disproportionate general degradation
Adapters prevent forgettingThey limit it by freezing the base, but the adapter still overrides behaviour
Rehearsal fully solves itIt substantially reduces it and requires retaining the original data
Frequent small updates are betterRepeated cycles compound drift; batched deliberate updates are safer

Why is the damage invisible?

Because you measure the thing you are training and rarely the things you are not.

Figure 1 — What your metrics see during an update

What you measure

Loss on the new task, accuracy on the new evaluation set, and the specific behaviour you set out to change.

All of it improves. The update looks like an unambiguous success and ships.

Improving

What you do not measure

Instruction-following on unrelated tasks, reasoning on long chains, rare vocabulary, formats used by other product surfaces.

Some of it degrades. Nothing errors, no dashboard moves, and the regression surfaces weeks later as a user complaint.

Eroding

The asymmetry is structural, not accidental. Evaluation sets are built around intended changes, so they systematically cannot detect unintended ones. A permanent broad regression suite is the only instrument that sees the right column.

What are the mitigation families?

  1. Rehearsal. Mix examples from previous tasks into the new training data. Simple, effective, and requires you to have kept the original data.
  2. Parameter isolation. Freeze the base and train separate adapters per task, so no update overwrites another. This is why adapter-based approaches dominate production.
  3. Regularisation. Penalise movement in parameters identified as important to earlier tasks, constraining where the update can go.
  4. Architectural separation. Keep the changing knowledge outside the weights entirely, in a retrieval index.

The fourth is not really a continual learning technique, and it is what most production teams should use. If knowledge lives in an index, updating it is a write operation with no gradient descent, no forgetting and no revalidation cycle.

Watch for this

Rehearsal requires retaining your original training data indefinitely, which is a data governance commitment rather than a technical footnote. Teams that delete training sets after a run, for storage or privacy reasons, lose the ability to rehearse later. Decide the retention policy before the first training run, not when you need it.

How do adapters change the picture?

They confine the damage rather than removing it.

Because a low-rank adapter leaves base weights frozen, the original model is always recoverable by detaching it. That is a genuine and underrated safety property: your rollback is exact rather than approximate.

What adapters do not prevent is behavioural override. An active adapter still changes how the model responds across inputs, including inputs it was never trained on. The base is preserved on disk while the served behaviour still shifts.

The practical consequence is that one adapter per task, swapped at request time, is safer than one adapter trained to handle several tasks. Narrow adapters are easier to evaluate and cheaper to retrain when a requirement changes.

How should you sequence updates?

Deliberately and infrequently, treating each as a release rather than a continuous drip.

  1. Freeze a regression suite covering general capability, unrelated product surfaces, and previously fixed bugs. This is the instrument, and it must predate the update.
  2. Batch changes. One update containing five improvements is safer than five sequential updates, because each cycle compounds drift.
  3. Measure before and after on both the target and the regression suite, and treat regression as a cost to be weighed rather than a footnote.
  4. Keep the previous version servable. Rollback must be a routing change, not a retraining project.
  5. Record what each version changed, so a behaviour question months later has an answer.

What about keeping a model current on facts?

Do not use training for this. It is the single clearest case in the whole topic.

Facts change on their own schedule, need to be citable, often need per-user access control, and occasionally need deleting on request. Weights satisfy none of those requirements: an update is expensive, unattributable, uncontrollable and effectively irreversible at the fact level.

A retrieval index satisfies all four trivially. The only knowledge that belongs in weights is knowledge that is stable, general and not subject to correction.

Where do teams go wrong?

Updating continuously

Frequent small updates feel responsive and compound drift faster than batched ones. Each cycle carries a little unmeasured degradation, and twelve monthly updates accumulate more than one annual update containing the same changes.

Building the regression suite after the first regression

By then you have no baseline for the capabilities that already degraded. The suite has to exist before the first update to be useful, which means building it while nothing appears to be wrong.

Training facts that should have been retrieved

The most common and most expensive error in this area. It creates a permanent maintenance obligation to solve a problem an index solves with a write.

What do experienced teams do differently?

They keep the base model frozen for as long as possible and push all volatility outward.

Prompts, retrieval indexes and adapters all change at their own cadence without touching the base. The base becomes a stable dependency updated on a deliberate schedule, which turns a continuous risk into a periodic one that can be planned and tested.

They also version adapters against specific base versions and refuse to mix them. An adapter trained against one base produces unpredictable behaviour on another, and that failure is subtle enough to survive casual testing.

A short glossary

Catastrophic forgetting
Loss of previously learned capability when a model is trained on new data, caused by shared parameters being updated.
Rehearsal
Mixing examples from earlier tasks into new training data so old capability is reinforced during the update.
Parameter isolation
Keeping task-specific parameters separate, typically via adapters, so updates cannot overwrite each other.
Regression suite
A frozen evaluation set covering capabilities you are not trying to change, used to detect unintended degradation.
Base pinning
Recording and enforcing which base model version an adapter was trained against.

Key takeaways

  • Forgetting happens because weights are shared, so no update touches only the new knowledge.
  • Damage is invisible by construction, since evaluation sets are built around intended changes.
  • Four mitigation families exist: rehearsal, parameter isolation, regularisation and architectural separation.
  • Adapters confine damage and give exact rollback, but still override behaviour on untrained inputs.
  • Rehearsal requires retaining original training data, which is a governance decision made early or not at all.
  • Batched deliberate updates drift less than frequent small ones.
  • Volatile facts belong in a retrieval index, where updating is a write rather than a training run.

Frequently asked questions

What causes catastrophic forgetting?

Shared parameters. The same weights encode many capabilities, so gradient steps that improve a new task necessarily move parameters the old task relied on. There is no update that isolates new knowledge, which is why the effect is structural rather than a tuning problem.

Do adapters prevent forgetting?

They limit it substantially and do not eliminate it. Because base weights stay frozen, the original model is exactly recoverable by detaching the adapter. But an active adapter still changes behaviour on inputs it was never trained on, so served behaviour shifts even though the base is preserved.

How do I detect forgetting?

With a frozen regression suite covering capabilities you are not trying to change: general instruction-following, unrelated product surfaces, and previously fixed bugs. Build it before the first update, since afterwards you have no clean baseline for whatever already degraded.

Is rehearsal worth the storage cost?

Usually, but it is a governance decision rather than a technical one. Rehearsal requires keeping original training data indefinitely, which conflicts with deletion policies some teams operate under. Settle the retention question before the first training run rather than when you need to rehearse.

Should I update my model frequently or in batches?

Batches. Each training cycle carries some unmeasured degradation, so twelve monthly updates accumulate more drift than one update containing the same twelve changes. Treat updates as releases with a full regression check rather than as a continuous process.

How do I keep a model current on changing facts?

Do not train them in. Facts need to be updatable, citable, access-controlled and sometimes deletable, and weights satisfy none of those. Put volatile knowledge in a retrieval index where updating is a write operation with no forgetting and no revalidation cycle.

Can I use an adapter with a different base model version?

No. Adapters are tied to the exact weights and layer shapes they were trained against, so pairing one with a different base produces unpredictable output. Pin and record the base version alongside every adapter, and refuse mixed combinations in deployment.

What should I do before any model update?

Freeze a regression suite, batch your changes into one deliberate update, measure both target improvement and regression, keep the previous version servable so rollback is a routing change, and record what the version altered so future behaviour questions have an answer.

References

    Rafael Ortega
    Rafael holds a B.Eng. in Mechatronics from Tecnológico de Monterrey and an M.S. in Robotics from Carnegie Mellon. He cut his teeth building perception pipelines for mobile robots in cluttered warehouses, tuning sensor fusion and debugging time-sync issues the hard way. Later, as an edge-AI consultant, he helped factories deploy real-time models on modest hardware, balancing accuracy with latency and power budgets. His writing brings that shop-floor pragmatism to topics like robotics safety, MLOps for embedded devices, and responsible automation. Expect diagrams, honest trade-offs, and “we tried this and it failed—here’s why” energy. Rafael mentors robotics clubs, contributes to open-source tooling for dataset versioning, and speaks about the human implications of automation for line operators. When he’s offline, he roasts coffee, calibrates a temperamental 3D printer, and logs trail-running miles with friends who tolerate his sensor jokes.

      Leave a Reply

      Your email address will not be published. Required fields are marked *