AI Post Transformers · Episode Companion

Language Model Continual Learning: Do Written Facts Survive Repeated Weight Updates?

arXiv:2607.11020 Qwen3-4B · LoRA adapters Charles O'Neill (Baseten) · posted 2026-07-14 View paper on arXiv →

The question in one line

A fact can reach a model two ways: placed in the prompt (instant, but gone when the conversation ends) or written into the weights (permanent, in principle). This paper asks whether the written version stays usable — not just present — after dozens, even a hundred, later training writes.

Theoretical lineage

Forgetting in neural nets predates transformers by decades. The line below traces the thread this paper pulls together.

Experimental pipeline

To isolate what a written fact retains — independent of anything the model already knew — O'Neill invents fictional entities and facts, writes them in via per-fact LoRA adapters, then grades recall against two fixed reference points.

Lenient vs strict grading → the entailment gap

Lenient grading credits anything that logically entails the answer; strict credits only the answer itself. Subtract strict from lenient and you get the entailment gap — how often a method just restates the trained premise instead of drawing the conclusion.

Hover a bar for the exact score.
26 pt
Bare-statement gap on use questions (single write)
1–5 pt
Study-data gap — a 22.8 point contrast
22.9 pt
8B replication of the same contrast (single-write only)

Write 20 facts, one at a time. Re-test the first one.

Each new fact is merged in via LoRA before the next arrives. Bare-statement training collapses almost immediately; study data survives longer, then plateaus — it never reaches zero.

● Bare-Statement ● Study Data
Click a legend chip to isolate a line. Hover a point for the exact value.
1%
Bare-statement retention after 20 sequential writes
46%
Study-data retention after 20 sequential writes
25–28%
Study-data plateau at 100 sequential writes

Access, not erasure — the sharpest result in the paper

Facts that fail every question at write 20 still hold most of their original log-probability lift. The knowledge never left; the route to it did.

Retention depends on the incoming method, not the stored one

Rows = how the original fact was written. Columns = how later facts were written. Hover a cell.

What a "forgotten" wrong answer actually contains

Bare-statement training, asked about fact #1 after 20 writes.
57–67%
Log-prob lift retained on facts that fail every test (drift-corrected)
70%
Wrong answers that actually contain the most recently written fact
77–80%
Accuracy recovery when the forgotten statement is re-shown in-prompt

This echoes a much older distinction: Tulving & Pearlstone's 1966 split between availability (is it in memory at all) and accessibility (can it be retrieved right now) — sixty years before it shows up again in LoRA adapters.

Capability loss tracks drift from the original model

Correlation between capability loss and KL divergence from the original model runs ρ≈0.83 across twelve conditions, up to 0.95 in the larger factorial. Distilling against a frozen copy of the original model keeps drift from compounding.

Hover a point. Green = frozen-teacher distillation. Purple = accumulating own-merges.

What the causal tests actually isolated

Prompt breadth itself — diverse recitation, no worked implications required — is the causal variable. Three targeted rescues aimed at the interference mechanism directly all failed to move retention.

0.795
Linearized Adam-update correlation with the next update's effect
−0.258
Same metric's correlation with eventual forgetting — predicts the step, not the trajectory

References

1Can a Language Model Learn Facts Continually in Its Weights? — O'Neill, 2026
arxiv.org/abs/2607.11020
2Catastrophic Interference in Connectionist Networks — McCloskey, Cohen, 1989
Google Scholar
3Overcoming Catastrophic Forgetting in Neural Networks — Kirkpatrick et al. (DeepMind), 2017
Google Scholar
4An Empirical Investigation of Catastrophic Forgetting in Gradient-Based Neural Networks — Goodfellow et al., 2013
Google Scholar
5An Empirical Study of Catastrophic Forgetting in LLMs During Continual Fine-tuning — Luo et al., 2023
Google Scholar
6Locating and Editing Factual Associations in GPT — Meng, Bau, Andonian, Belinkov, 2022
Google Scholar
7Mass-Editing Memory in a Transformer — Meng, Sharma, Andonian, Belinkov, Bau, 2023
Google Scholar
8Fast Model Editing at Scale — Mitchell, Lin, Bosselut, Finn, Manning, 2022
Google Scholar
9MQuAKE: Assessing Knowledge Editing via Multi-Hop Questions — Zhong, Wu, Manning, Potts, Chen, 2023
Google Scholar
10RL's Razor: Why Online RL Forgets Less — Shenfeld, Pari, Agrawal (MIT), 2025
Google Scholar
11Does Localization Inform Editing? — Hase, Bansal, Kim, Ghandeharioun (UNC/Google), NeurIPS 2023
Google Scholar
12Towards Mechanistically Understanding Why Memorized Knowledge Fails to Generalize — Dai et al., 2026
Google Scholar
13Model Editing at Scale Leads to Gradual and Catastrophic Forgetting — Gupta, Rao, Anumanchipalli, 2024
Google Scholar
14AI Models Collapse When Trained on Recursively Generated Data — Shumailov et al., 2024
Google Scholar
15LoRA vs Full Fine-Tuning: An Illusion of Equivalence — Shuttleworth, Andreas, Torralba, Sharma (MIT), 2025
Google Scholar
16Availability Versus Accessibility of Information in Memory for Words — Tulving, Pearlstone, 1966
Google Scholar