A fact can reach a model two ways: placed in the prompt (instant, but gone when the conversation ends) or written into the weights (permanent, in principle). This paper asks whether the written version stays usable — not just present — after dozens, even a hundred, later training writes.
Forgetting in neural nets predates transformers by decades. The line below traces the thread this paper pulls together.
To isolate what a written fact retains — independent of anything the model already knew — O'Neill invents fictional entities and facts, writes them in via per-fact LoRA adapters, then grades recall against two fixed reference points.
Lenient grading credits anything that logically entails the answer; strict credits only the answer itself. Subtract strict from lenient and you get the entailment gap — how often a method just restates the trained premise instead of drawing the conclusion.
Each new fact is merged in via LoRA before the next arrives. Bare-statement training collapses almost immediately; study data survives longer, then plateaus — it never reaches zero.
Facts that fail every question at write 20 still hold most of their original log-probability lift. The knowledge never left; the route to it did.
This echoes a much older distinction: Tulving & Pearlstone's 1966 split between availability (is it in memory at all) and accessibility (can it be retrieved right now) — sixty years before it shows up again in LoRA adapters.
Correlation between capability loss and KL divergence from the original model runs ρ≈0.83 across twelve conditions, up to 0.95 in the larger factorial. Distilling against a frozen copy of the original model keeps drift from compounding.
Prompt breadth itself — diverse recitation, no worked implications required — is the causal variable. Three targeted rescues aimed at the interference mechanism directly all failed to move retention.
| 1 | Can a Language Model Learn Facts Continually in Its Weights? — O'Neill, 2026 arxiv.org/abs/2607.11020 |
| 2 | Catastrophic Interference in Connectionist Networks — McCloskey, Cohen, 1989 Google Scholar |
| 3 | Overcoming Catastrophic Forgetting in Neural Networks — Kirkpatrick et al. (DeepMind), 2017 Google Scholar |
| 4 | An Empirical Investigation of Catastrophic Forgetting in Gradient-Based Neural Networks — Goodfellow et al., 2013 Google Scholar |
| 5 | An Empirical Study of Catastrophic Forgetting in LLMs During Continual Fine-tuning — Luo et al., 2023 Google Scholar |
| 6 | Locating and Editing Factual Associations in GPT — Meng, Bau, Andonian, Belinkov, 2022 Google Scholar |
| 7 | Mass-Editing Memory in a Transformer — Meng, Sharma, Andonian, Belinkov, Bau, 2023 Google Scholar |
| 8 | Fast Model Editing at Scale — Mitchell, Lin, Bosselut, Finn, Manning, 2022 Google Scholar |
| 9 | MQuAKE: Assessing Knowledge Editing via Multi-Hop Questions — Zhong, Wu, Manning, Potts, Chen, 2023 Google Scholar |
| 10 | RL's Razor: Why Online RL Forgets Less — Shenfeld, Pari, Agrawal (MIT), 2025 Google Scholar |
| 11 | Does Localization Inform Editing? — Hase, Bansal, Kim, Ghandeharioun (UNC/Google), NeurIPS 2023 Google Scholar |
| 12 | Towards Mechanistically Understanding Why Memorized Knowledge Fails to Generalize — Dai et al., 2026 Google Scholar |
| 13 | Model Editing at Scale Leads to Gradual and Catastrophic Forgetting — Gupta, Rao, Anumanchipalli, 2024 Google Scholar |
| 14 | AI Models Collapse When Trained on Recursively Generated Data — Shumailov et al., 2024 Google Scholar |
| 15 | LoRA vs Full Fine-Tuning: An Illusion of Equivalence — Shuttleworth, Andreas, Torralba, Sharma (MIT), 2025 Google Scholar |
| 16 | Availability Versus Accessibility of Information in Memory for Words — Tulving, Pearlstone, 1966 Google Scholar |