1 00:00:01,000 --> 00:00:42,825 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today's paper is "Can a Language Model Learn Facts Continually in Its Weights?" — a solo paper, one author, Charles O'Neill, out of Baseten, posted to arXiv on July 14th, 2026. The title's asking something specific: a new fact can reach a model two ways — placed in the prompt, usable instantly but gone once the conversation ends, or written into the weights, permanently. Does that written version actually stay usable once you keep training the model afterward, write after write? Or does it just get buried? 2 00:00:42,825 --> 00:01:11,349 [Dr. Ada Shannon] What pulled me into this one is that it's not another 'we ran fine-tuning, here's an accuracy number' paper. O'Neill builds an actual theory of what a written fact even is inside the weights — a fully usable piece of knowledge, or something narrower that only answers the exact question it was trained on — and follows that object through twenty, even a hundred, later writes. That's a more careful question than most of the knowledge-editing literature bothers asking, and it's what hooked me before I even got to the results. 3 00:01:11,349 --> 00:01:34,399 [Hal Turing] Okay, let's ground this for anyone who hasn't followed the knowledge-editing world closely. We talk about training models all the time on this show, but 'continual learning' specifically means something harder than just training once and calling it done, right? It's the model picking up brand-new facts after it's already fully trained, and holding onto them through round after round of further updates. 4 00:01:34,399 --> 00:02:13,849 [Dr. Ada Shannon] Right, and that's hard because of catastrophic forgetting, and it predates transformers entirely. Michael McCloskey and Neal Cohen, out of Johns Hopkins, showed it back in 1989 with plain backprop networks: train one on task A, then task B, and instead of gently getting rusty on A, it collapses on A almost completely. That's structural — everything a network knows lives in one shared set of weights, and every gradient step for something new nudges those same weights with zero visibility into what it does to something learned ten updates ago. This paper reframes forgetting specifically: not 'did accuracy drop,' but 'is the fact still reachable at all' — a very different question. 5 00:02:13,849 --> 00:02:29,875 [Hal Turing] So if fine-tuning the whole model on a new fact is that blunt an instrument, is that where 'knowledge editing' comes in — some more surgical way to change one fact without repeating that Johns Hopkins collapse? And where does LoRA fit into that? 6 00:02:29,875 --> 00:03:23,199 [Dr. Ada Shannon] Exactly — knowledge editing means directly changing a specific fact in the weights, through fine-tuning, distillation, or locate-and-edit methods that find where a fact lives and patch just that spot. LoRA is what this paper actually uses: freeze the base model, train a small low-rank adapter instead of touching every parameter — the approach Hu and colleagues published out of Microsoft in 2021. Each new fact gets its own adapter, merged in before the next arrives. Underneath that sits a distinction that matters more than people assume: factual recall, stating a trained fact, versus using it — paraphrase, application, combining it with something else. Fine-tuning is notoriously bad at the second one. Gekhman and colleagues showed in 2024 that fine-tuning learns new facts slowly and fails reversals and multi-hop use, even when the same fact works fine in a prompt. Berglund and colleagues found the reversal curse — a model trained on 'A is B' often can't answer 'what is B,' even though it's— 7 00:03:23,199 --> 00:03:40,149 [Hal Turing] Oh wait, wait, hold on — that's the one where they trained on stuff like 'Tom Cruise's mother is Mary Lee Pfeiffer,' and the model could say who Tom Cruise's mother is, but flip it to 'who is Mary Lee Pfeiffer's son' and it just can't get there — same fact, wrong direction? 8 00:03:40,149 --> 00:04:16,199 [Dr. Ada Shannon] That's the one. And here's why it matters for today's paper: put that same fact in the prompt instead of the weights, and the direction problem disappears — the model handles it fine either way. So whatever breaks is specific to writing into weights, not some general reasoning limit. Allen-Zhu and Li added the other half of this picture in 2024: facts trained on lots of varied restatements are far easier to actually extract and use later than a fact trained on one bare sentence repeated over and over. That's exactly the thread O'Neill pulls on — not just whether a fact gets written, but what kind of data writes it. 9 00:04:16,199 --> 00:04:29,250 [Hal Turing] So how do you actually test this without the model's own existing knowledge muddying the results? You can't just ask trivia — you wouldn't know if a right answer came from the write or from something it already knew. 10 00:04:29,250 --> 00:05:12,100 [Dr. Ada Shannon] Right, so they invent facts wholesale — 'Zorvathine is a metal that melts when cooled below minus ten degrees Celsius,' that kind of thing — entities that don't exist and usually contradict the model's default assumption. Those get written into Qwen3-4B with the LoRA adapters we just covered, then tested with five question types: recall, paraphrase, application, composition, and counterfactual — state it, restate it, use it, combine it with outside knowledge, or choose it over the model's default. Two fixed reference points anchor everything: the untouched original model as the floor, since it's never seen the fact, and that same model with the fact placed in its prompt as the ceiling. Every result in this paper measures one thing — how far a written fact falls short of that ceiling. 11 00:05:12,100 --> 00:05:37,650 [Dr. Ada Shannon] —then tested against the five question types, graded two ways: lenient credits anything that logically entails the right conclusion, strict credits only the conclusion itself. Subtract strict from lenient and you get what O'Neill calls the entailment gap — how often a method stops at reciting the premise instead of stating the answer. Say 'its melting point is minus ten' instead of 'it melts' — lenient passes that, strict doesn't. That gap is the diagnostic for everything I'm about to walk you through. 12 00:05:37,650 --> 00:05:52,375 [Hal Turing] So it's less 'does it know the fact' and more 'does it draw the conclusion or just point at the evidence.' Given that, what actually separates a method that builds usable knowledge from one that just builds a very confident parrot? 13 00:05:52,375 --> 00:06:34,350 [Dr. Ada Shannon] Almost entirely the data's breadth. Bare-statement training — one sentence, drilled — leaves a 22-to-42-point gap on the use questions. Swap in what they call 'study' data, two dozen paraphrases and worked implications and contrasts with whatever default the fact violates, same step budget, and the gap collapses to 1 to 5 points. Same holds for context distillation — matching the student to a teacher with the fact in its prompt, offline on fixed samples or online on the student's own generations, either KL direction. Bare-statement's the outlier at 26 points; every richer method sits at 2 to 4, a 22.8-point contrast. Flipping forward to reverse KL, betting it'd force more committed knowledge — no effect once the data's already broad. 14 00:06:34,350 --> 00:06:51,775 [Hal Turing] Oh wait, hold on — you mentioned a whole factorial crossing objective, data, and update method, plus an 8B replication. Does that 8B run actually confirm what's coming — the retention collapse — or is that still riding entirely on the 4B numbers? 15 00:06:51,775 --> 00:07:27,725 [Dr. Ada Shannon] Just the single-write gap — 26 points for bare-statement again, a 22.9-point contrast, matching 22.8 at 4B. Nothing past that got rerun — no sequential retention, no storage probes, no causal tests. Everything from here is 4B-and-LoRA territory. Which matters, because: write twenty facts one at a time, merging each into the model before the next arrives, then re-test the earlier ones. Bare-statement retains 1 percent. Study retains 46. Push it to a hundred sequential writes and study plateaus around 25 to 28 percent instead of hitting zero — some facts just survive, not because they were written any better. 16 00:07:27,725 --> 00:07:39,275 [Hal Turing] 1 percent — so it's wiped, not gradually eroded, in twenty rounds. When a fact 'fails' like that, is the information actually gone, or is something else happening underneath? 17 00:07:39,275 --> 00:08:18,450 [Dr. Ada Shannon] Access, not erasure — the sharpest result in the whole paper. They save every adapter and track the probability assigned to the exact written statement across later writes. Facts that fail every question at write twenty still hold 57 to 67 percent of their original log-probability lift, drift-corrected. Under bare-statement training, 70 percent of the wrong answers about a forgotten fact actually contain whatever was written most recently — ask about fact one after twenty writes, get fact twenty back instead. And hand the model its own forgotten statement in the prompt, and accuracy recovers to 77 to 80 percent. The knowledge never left. The route to it did. 18 00:08:18,450 --> 00:08:33,550 [Hal Turing] A filing cabinet where the label keeps getting swapped out, but the folder's still sitting in the drawer. Separate question then — what does all this repeated writing cost the model's general abilities? You can't bolt facts on forever for free. 19 00:08:33,550 --> 00:09:16,275 [Dr. Ada Shannon] Capability loss tracks KL divergence from the original model almost directly — rho around 0.83 across their twelve conditions, up to 0.95 in the bigger factorial. Which is why distilling against a frozen copy of the original, instead of the model's own accumulated merges, writes facts at near-zero capability cost — the drift never gets a chance to compound. An explicit penalty pulling every write back toward the current base rescues capability almost completely, without the measured KL actually dropping — the correlation holds, but it's not a lever you can just pull directly. It's basically extending Shenfeld, Pari, and Agrawal's 'RL's Razor,' out of MIT, 2025 — KL from the base policy predicting forgetting under reinforcement learning — over to plain knowledge writing. 20 00:09:16,275 --> 00:09:30,050 [Hal Turing] Drift predicts the damage, but the fix doesn't actually have to reduce the drift to work — strange result on its own. Last piece then: what did the causal tests find was actually causing the interference? 21 00:09:30,050 --> 00:10:37,650 [Dr. Ada Shannon] Prompt breadth itself is the causal variable, not some derived reasoning step — diverse recitation alone, no worked implications at all, drops the gap from 27.4 to 5.4 points. And interference comes from the incoming write, not from how the earlier fact was stored — cross bare-statement and study as both the stored and the incoming method, and only the incoming method moves retention. A linearized version of the Adam update correlates 0.795 with the next update's immediate effect, but negative 0.258 with eventual forgetting — it predicts the next step, not the trajectory. Three separate rescues — bridging data, activation-guided patching, gradient projection away from the conflict — all failed to move retention, echoing Hase, Bansal, Kim, and Ghandeharioun out of UNC Chapel Hill, NeurIPS 2023: localizing where a fact lives doesn't tell you where to intervene. And worth flagging — Kirkpatrick's elastic weight consolidation, DeepMind, 2017, the textbook fix for exactly this kind of forgetting, gets cited here but never actually run against that KL penalty. 22 00:10:37,650 --> 00:11:15,800 [Hal Turing] So here's what nagged at me through the back half of this paper. Every piece of the pipeline — the invented facts, the training data, the held-out questions, the grading — is generated and judged by LLMs from basically the same model family the paper is testing. O'Neill's own Table 2 lists six ways that setup went wrong before anyone caught it: shared-generator leakage, a length cap punishing one answer style, a checker blind to negation, a judge that silently declined on hard items. If it took six manual audits to catch the instrument lying to itself, how much do we trust the entailment gap and retention numbers? 23 00:11:15,800 --> 00:11:56,250 [Dr. Ada Shannon] Less than the headline suggests, but more than a naive self-referential setup usually earns. Every one of those six defect classes got caught by a certification gate before touching a real result — floor under 20%, ceiling above 80%, truncation under 5% per condition, an adversarial leak audit, and a manual pass-audit, not just a failure audit. They even re-certified under the exact dual-grading policy that produces the entailment gap. That's real discipline, not a disclaimer buried in limitations. But certification only catches what someone thought to test for — a shared-generator artifact nobody designed an audit for just quietly inflates a number that looks clean. Residual risk: real, bounded, not zero. 24 00:11:56,250 --> 00:12:23,975 [Hal Turing] Oh wait, hold on — bounded by what, though? Because what actually worries me more is that nearly every headline claim — the KL-drift ordering, the frozen-teacher rescue, the whole retention collapse — runs on Qwen3-4B with LoRA. The 8B model shows up exactly once, and it only replicates the single-write entailment gap. Not the sequential retention, not storage-versus-access, none of the causal interventions. That's a lot resting on one small model at one scale. 25 00:12:23,975 --> 00:13:04,750 [Dr. Ada Shannon] Right, and it's worse than 'small model, one scale' — the rank results contradict prior literature. Higher LoRA rank helps retention here but hurts capability, the reverse of Biderman et al.'s 2024 Databricks finding that low-rank training forgets less. O'Neill flags it as inconsistent rather than burying it. One candidate explanation: Shuttleworth, Andreas, Torralba, and Sharma's 2025 MIT paper, 'LoRA vs Full Fine-Tuning: An Illusion of Equivalence,' pins LoRA's distinct forgetting on intruder dimensions, not rank. If that's driving the interference story here, the headline advice — keep facts in context, not weights — might describe adapter-merging behavior specifically, not weight-writing in general. 26 00:13:04,750 --> 00:13:35,000 [Hal Turing] That connects to something else that bugged me — though credit where it's due, most papers would just bury this. The KL-to-current-base penalty rescues capability without even lowering the measured drift — strange on its own, and they say so plainly. But the textbook forgetting fix is regularizing toward the original parameters, Kirkpatrick, Pascanu, Rabinowitz and colleagues out of DeepMind, 2017, PNAS — elastic weight consolidation. It's cited in their own related work. Why not just run it? 27 00:13:35,000 --> 00:14:21,175 [Dr. Ada Shannon] No idea, honestly, and it's a strange gap since EWC is cheap and the obvious adjacent baseline. My guess is they're committed to KL-divergence space instead, which traces to Shenfeld, Pari, and Agrawal's 2025 MIT paper, 'RL's Razor: Why Online Reinforcement Learning Forgets Less' — KL from the base policy predicts forgetting under RL. O'Neill extends that from RL to knowledge writing, correlation 0.83 to 0.95 across the factorial. Nice lineage-tracing. But every intervention they test lives in that one family — penalize or freeze against a KL reference — the classical parameter-space alternative never gets a fair fight. Same story with the activation-patching failures: those accord with Hase, Bansal, Kim, and Ghandeharioun's 2023 NeurIPS paper, out of UNC and Google, warning that localizing a fact doesn't tell you where to fix it. 28 00:14:21,175 --> 00:14:41,250 [Hal Turing] Given all that — unresolved reachability, one model and scale carrying the weight, a fix that doesn't touch what it's supposedly fixing — what should a practitioner actually change today? And where does O'Neill point for what's next, because 'reachability remains unsolved' isn't exactly a satisfying place to stop. 29 00:14:41,250 --> 00:15:21,525 [Dr. Ada Shannon] If a fact has to be retrieved on demand, composed, or survive further training — put it in context, not weights. That's the one recommendation the evidence supports end to end. If you do write to weights, use broad data every time, never bare statements, and distill against a frozen original model, never the model's own accumulated merges — that's the gap between near-zero capability cost and the worst condition in the paper. As for what's next, O'Neill wants a mechanistic account that predicts final behavior past the next update and actually gives a fact an address instead of storing it somewhere reachable by luck. And the reinforcement-learning condition barely produced a signal, so whether RL can write knowledge into weights at all is still wide open. 30 00:15:21,525 --> 00:15:46,925 [Hal Turing] So the takeaway isn't 'never write to weights' — it's that weights are a content store without an address book, and nobody's built one yet. Broad data creates something usable, a frozen teacher keeps the rest of the model intact, but getting an old fact back once new ones have piled on top — that stayed unsolved through every intervention they tried, and that's the piece worth sitting with. Thanks for listening, everyone — we'll see you next time.