AI Post Transformers — Episode Companion

Continual Learning in LLMs: Beyond Catastrophic Forgetting

A survey maps continual learning across three LLM training stages — pretraining, fine-tuning, alignment — onto three classical mitigation families: rehearsal, regularization, and architecture-based methods. This page visualizes the taxonomy, the method mechanics, illustrative benchmark behavior, and where the survey's own rigor claims don't hold up.

arXiv:2603.12658 Posted March 13, 2026 Chen, Sun, Ye, Li, Lin — Zhejiang Lab

The Three-Stage Pipeline

Hover a stage for what "continual" means at that point in the LLM lifecycle.

Stage × Mechanism Coverage Matrix

The survey's core organizing move: subdivide rehearsal / regularization / architecture by which training stage they're applied at. Cell shade = relative density of cited methods (mock illustrative weighting). Hover a cell for named methods.

Note: the alignment row is split by the survey into RL-free vs. RL-based rather than this exact axis — mapped here to the closest mechanism analogue.

Three Mitigation Families

Same skeleton, reused at every stage: replay old data, penalize drift on important weights, or bolt on new capacity and freeze the rest.

Isolation vs. Transfer: O-LoRA vs. SAPT

Same architecture-based family, opposite design bet. Toggle to compare.

Forgetting Rate vs. Backward Transfer

Illustrative benchmark comparison across method families. Toggle metric.

The Stability Gap During Continual Pretraining

Accuracy relative to the pre-CPT baseline (100%) over training steps. Performance dips before it recovers — sometimes below baseline, sometimes past it, depending on mitigation.

Rigor Scorecard

What the survey's abstract and structure claim, versus what the text actually delivers.

Cited But Not Discussed, or Missing Entirely

Papers directly relevant to the taxonomy that are absent from the bibliography or never engaged with in-text.

References