1 00:00:01,000 --> 00:00:39,325 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today's paper is "Towards Mechanistically Understanding Why Memorized Knowledge Fails to Generalize in Large Language Model Finetuning" — first author Lu Dai, et al., six co-authors total, out of HKUST's Guangzhou campus and HKUST proper, posted to arXiv on July 9th this year. And Ada, the hook for me was simple: fine-tune a model on one new fact, it nails direct recall instantly, then completely fails the moment you ask it to actually reason with that fact. 2 00:00:39,325 --> 00:00:59,725 [Dr. Ada Shannon] What sold me wasn't the failure itself — we've seen 'memorizes but can't use it' before. It's that these authors don't stop at documenting the gap. They bring in a causal intervention tool and trace, layer by layer, checkpoint by checkpoint, exactly where the fact sits inside the network when it fails versus when it finally works. That's a mechanistic account, not another benchmark just restating the symptom. 3 00:00:59,725 --> 00:01:22,025 [Hal Turing] So the real question driving this paper: once fine-tuning has a fact memorized cold, when — and why — does it become something the model's own reasoning circuits can actually pick up? Before we get into their method, though, it's worth stepping back, because this sits inside a bigger question of how you even get new knowledge into a deployed LLM in the first place. 4 00:01:22,025 --> 00:01:59,875 [Dr. Ada Shannon] There are basically three routes. Retrieval-augmented generation just hands the model the fact fresh in the prompt — nothing touches the weights, attention does the binding at inference time. Model editing, like ROME and MEMIT, out of Meng, Bau and colleagues' work at MIT and Northeastern, 2022 and 2023, goes surgical: locate the MLP weights encoding a fact and overwrite them directly, no gradient descent involved. Fine-tuning, what this paper studies, just keeps training the model on the fact with ordinary gradient updates, and it's the one method that's supposed to integrate new knowledge into the model's existing reasoning — not bolt it on like RAG, or edit it in isolation like ROME. 5 00:01:59,875 --> 00:02:34,650 [Hal Turing] And that integration is exactly what falls apart. This 'remembers it but can't use it' pattern isn't new — MQuAKE, from Zhong and colleagues out of Princeton, 2023, built a benchmark testing whether an edited fact survives being chained into a multi-hop question, and found editors like ROME ace single-hop recall but collapse on the second hop. RippleEdits, from Cohen and colleagues out of Tel Aviv University, 2024, found the same kind of failure to propagate. So the symptom's well known. What's actually new here? 6 00:02:34,650 --> 00:03:11,475 [Dr. Ada Shannon] What's new is they open the hood instead of just measuring the outside. That's mechanistic interpretability — judging a model by reading its internals rather than just its output. Two ideas anchor it: the linear representation hypothesis, from Park, Choe and Veitch out of the University of Chicago, ICML 2024, which says LLMs encode facts as literal directions in activation space. And the ROME paper itself, Meng, Bau and colleagues, MIT and Northeastern, NeurIPS 2022, which showed facts often live specifically in MLP weights, acting like a key-value lookup — subject goes in, attribute comes out. 7 00:03:11,475 --> 00:03:20,850 [Hal Turing] Oh wait, wait — hold on, that's the same ROME you just mentioned as a model-editing method a second ago? Same paper, doing two jobs? 8 00:03:20,850 --> 00:03:43,100 [Dr. Ada Shannon] Same paper, exactly. ROME's 'causal tracing' is what let them locate facts in the first place — corrupt a prompt, restore hidden states one layer at a time, see which restoration brings back the right answer. That causal-tracing lineage, plus later activation-patching work, is the direct ancestor of this paper's own method, self-patching. Same intervention flavor, different question: not just where does a fact live, but can that stored fact be made usable if you move it somewhere else. 9 00:03:43,100 --> 00:04:23,025 [Hal Turing] Okay, let's pin down vocabulary, because it carries the rest of the episode. The core phenomenon is the Knowing-Using Gap, and it has two faces. An accuracy gap — nails direct recall, fails at downstream use. And a temporal lag — even when use eventually emerges, it shows up many epochs after memorization already flattened out. They test it with two task types: chaining, where the model must resolve a bridge entity from one memorized fact before answering a second, linked question; and intersection, where it has to spot the one entity shared between two memorized facts while filtering out noise. 10 00:04:23,025 --> 00:04:58,525 [Dr. Ada Shannon] The tool for looking inside is self-patching — conceptually, at this stage, just take the internal representation from one point in the network where a fact might be sitting, splice it into another point, and check whether the answer suddenly improves. If it does, the fact was there, just not being read from where it needed to be. That's what leads to their central claim, the knowledge-circuit misalignment hypothesis: the fact does get written into the weights, but it lands in layers built for storage and recall, not the mid-layer circuits that actually do multi-hop reasoning. It's not missing. It's misfiled. 11 00:04:58,525 --> 00:05:20,700 [Hal Turing] Which is a genuinely different framing than I expected going in — not 'the model doesn't know it,' but 'it knows it, filed in the wrong place.' So can you actually watch that misfiling happen over the course of training, and can you fix it? That's exactly where the paper goes next, with some pretty striking numbers on how much of that lost reasoning turns out to be recoverable. 12 00:05:20,700 --> 00:06:01,025 [Dr. Ada Shannon] This is where self-patching earns its keep. Causal tracing needs a clean, correct run to corrupt and restore — no good here, since chaining fails outright. PatchScope, from Ghandeharioun and colleagues at Google DeepMind, ICML 2024, decodes a hidden state into language via an auxiliary prompt — it tells you what a representation means, not whether it's usable. Self-patching skips interpretation: grab the entity representation from one prompt and layer, drop it into a different layer mid-forward-pass, check if the answer flips from wrong to right. No decoding, no clean baseline needed — just a before-and-after accuracy delta, which lets you study failures directly instead of only successes. 13 00:06:01,025 --> 00:06:11,225 [Hal Turing] So it's testing usability, not legibility. Fine — what did that scan actually find once they ran the numbers, LoRA versus full fine-tuning? 14 00:06:11,225 --> 00:06:44,175 [Dr. Ada Shannon] Under LoRA, chaining shows a real gap — about four-and-a-half epochs of lag, landing at just thirty percent final accuracy. Intersection's the easy case: half an epoch of lag, over ninety percent — no bridge entity required. Full fine-tuning memorizes far faster, single-digit epochs instead of double, but that speed doesn't transfer: chaining under FFT lands at nearly identical accuracy to LoRA, and intersection actually gets worse — bigger lag, lower ceiling. Scale doesn't rescue it either: bigger models never eliminate the lag, and more injected facts widen the final gap even while raw recall stays strong. 15 00:06:44,175 --> 00:06:54,650 [Hal Turing] Okay, structural, not undertrained. So walk me through it — you said you can watch this misfiling happen live, epoch by epoch. What does that actually look like? 16 00:06:54,650 --> 00:07:17,750 [Dr. Ada Shannon] They checkpoint every epoch, self-patch at the head-entity position, and build a heatmap of source layer against target layer. Before memorization, it's solid blue — nothing to relocate. Right as memorization saturates, an off-diagonal red patch appears: the fact's stored somewhere, but not routed to where it needs to be. From there it splits — successful cases widen that red region until it reaches the diagonal; failures keep expanding for a while, then just stall, short of the diagonal, permanently. 17 00:07:17,750 --> 00:07:24,425 [Hal Turing] Oh wait, wait — stalls *why*, though? What actually stops the crawl if training's still running? 18 00:07:24,425 --> 00:08:00,600 [Dr. Ada Shannon] Because the loss has nothing left to push against — once the model matches the answer, loss collapses toward zero and the gradient with it. No error signal left to nudge the representation forward, so it's stranded where the pressure stopped, not where reasoning needs it. Which is the case for forcing it: patch at the single best layer pair per instance — the oracle — and chaining jumps one-point-five to six-x across every model and both domains tested. Intersection's smaller gap nearly closes, mid-to-high nineties. Holds from a one-point-five billion Qwen up to an eight billion LLaMA — uniform lift across six models isn't what a random artifact looks like. 19 00:08:00,600 --> 00:08:10,225 [Hal Turing] So geographically, where do the winning patches actually come from? Is there a real pattern to which layer pairs work, or is it scattered per instance? 20 00:08:10,225 --> 00:08:49,625 [Dr. Ada Shannon] Clean pattern, half intuitive. Effective sources cluster into two groups — early layers and late layers — both targeting the same middle band. Late-to-middle makes sense: you're moving a representation already enriched by everything computed so far, which lines up with the 'Hopping too late' result from Biran and colleagues at Tel Aviv University and Google, EMNLP 2024 — more on that later. Early-to-middle is the surprise: the model can skip several layers entirely and still succeed, so the fact just never gets picked up. Late-to-late does nothing — dead zone — tracking with Geva and colleagues' three-step account of factual recall, Google Research and Tel Aviv University, EMNLP 2023. 21 00:08:49,625 --> 00:08:59,851 [Hal Turing] Before I'm fully sold — could this just be a generic perturbation effect? Would jamming almost any activation into that spot produce a similar bump? 22 00:08:59,851 --> 00:09:35,176 [Dr. Ada Shannon] They checked. Entity-position patching gives by far the largest, most significant effect; random tokens and beginning-of-sequence are noise, end-of-sequence shows a smaller secondary bump from aggregating pre-generation information. Two more controls: chain-of-thought helps chaining a little but tops out well below self-patching and can hurt intersection; patching in an unrelated fact's representation helps marginally too, probably just from perturbing the forward pass, but lags way behind the correct fact. Not position-agnostic, not a prompting trick, not generic noise — specifically the right fact, at the right position. 23 00:09:35,176 --> 00:09:45,076 [Hal Turing] Convinced. So — an oracle that scans every layer pair per instance isn't shippable. Did they turn this into anything actually deployable? 24 00:09:45,076 --> 00:10:19,626 [Dr. Ada Shannon] They did, by throwing away the per-instance search. Since patches cluster into just those two groups, they fixed two layer pairs per architecture — roughly zero-point-eight-two of total layers feeding zero-point-four-five, and roughly zero-point-one feeding that same target. No scanning, no knowing where a fact lives in advance. It recovers fifty-eight to seventy-five percent of the oracle headroom, across every model and both tasks. That's the move from 'here's why it fails' to something you could bolt onto a fine-tuning pipeline. What's left is instance-specific variation — some facts just don't live where the heuristic expects. 25 00:10:19,626 --> 00:10:52,751 [Hal Turing] Okay, genuinely deployable. But before we close this out, how far does it actually reach? Every model here tops out at eight billion parameters — LLaMA-3.1-8B is the biggest thing tested. Production models people deploy today run thirty, seventy, over a hundred billion, with way more layers and wider residual streams. Does that two-cluster geometry — early-to-mid, late-to-mid — even survive at that depth? Or does the whole picture shift once the residual stream has that much more room to spread information across? 26 00:10:52,751 --> 00:11:15,101 [Dr. Ada Shannon] Honest answer: nobody knows, and the paper doesn't pretend otherwise. Zero experiments past eight billion parameters. And it's not a trivial extrapolation — if a hundred-billion-parameter model has three times the layers, is the effective location still around point-eight-two L and point-one L, or does that fraction drift as depth grows? That's untested. Given how much weight that fractional positioning carries for the whole fixed-heuristic pitch, that's a real hole. 27 00:11:15,101 --> 00:11:39,426 [Hal Turing] Hold on, back up — 'fractional positioning' is bugging me now. Table 7 evaluates the fixed heuristic on the same six models used to eyeball the clustering in Figure 5, right? Where's the held-out model, the architecture they didn't peek at while picking point-eight-two and point-one? Tune your recipe and grade your recipe on the identical model set, and that's not validation — that's circular by construction. 28 00:11:39,426 --> 00:12:10,976 [Dr. Ada Shannon] Guilty as charged — no held-out architecture or task family anywhere. Both fractions come from the same six models graded in Table 7. There's a quieter version of the same issue in section five-two: the permeation heatmaps anchoring the whole 'gradient vanishes, model's stuck' story run on just a hundred sampled chaining tasks over thirty epochs, LoRA only. That's much smaller and shorter than the thousand-sample, fifty-epoch runs behind the oracle numbers in Tables four and seven. Nobody's shown the same stall-before-the-diagonal pattern holds under full fine-tuning at the paper's own headline scale. 29 00:12:10,976 --> 00:12:45,876 [Hal Turing] That scale mismatch connects to something else in the references. Allen-Zhu and Li's 'Physics of Language Models' series, out of Meta AI and Carnegie Mellon, has a capacity-scaling entry — Part three-point-three — cited here. Part three-point-two, 'Knowledge Manipulation,' isn't. That one asks almost this exact question: can a model compose a memorized fact instead of just reciting it? That's the Knowing-Using Gap in different clothes, studied years earlier. Citing the sibling paper on capacity but skipping the one that anticipated your central finding is a strange omission. 30 00:12:45,876 --> 00:13:26,501 [Dr. Ada Shannon] Especially since they do lean on Biran and colleagues' 'Hopping Too Late,' out of Tel Aviv University and UCL, twenty twenty-four, to justify why late-to-late patching is useless. That paper independently showed multi-hop reasoning stalls because by late layers the relevant information has already moved to the final token position — nothing left mid-stream to redirect. That's the exact layer-timing constraint their two-cluster geometry sits on top of. And speaking of citations that go nowhere: CaKE, Yao and colleagues out of Zhejiang University, twenty twenty-five, is a circuit-aware editing method built to solve this same routing problem. It's in related work. It's never benchmarked against the fixed heuristic. That's the missing baseline. 31 00:13:26,501 --> 00:14:11,426 [Hal Turing] Shame, because the honest framing here is genuinely good — this isn't a capacity problem, it's a routing problem. The fact sits in the model, just filed somewhere reasoning can't reach on its own. That matters for ROME and MEMIT too, since both already assume facts live at addressable MLP sites — this is basically saying they're sometimes filed at the wrong address for downstream use. But the fixed heuristic still needs white-box activation access and per-architecture tuning off of L. Most real knowledge injection today happens through closed provider fine-tuning APIs with zero activation access. That's the gap between a nice mechanistic story and something a practitioner can actually run. 32 00:14:11,426 --> 00:14:44,576 [Dr. Ada Shannon] The authors are upfront about a chunk of this themselves. Single anchor position only, so redundant encoding elsewhere could mean the real recoverable headroom is bigger than measured. The oracle's explicitly a diagnostic ceiling, not a deployable method. No early-training signal yet to flag which facts will fail before you've burned the compute. And there's a dual-use note worth saying out loud: the same mechanism that lets you repair a bad knowledge edit could just as easily make it easier to robustly wire in a false or harmful one. 33 00:14:44,576 --> 00:15:10,726 [Hal Turing] Fair, and worth sitting with. Big picture: fine-tuning doesn't fail to teach a model new facts — it fails to route them to where reasoning actually happens, and that routing failure looks at least partly fixable without retraining from scratch. Just hold the claims to what six sub-eight-billion-parameter models on two synthetic KG tasks can actually tell you, not to what the abstract wants you to believe. Thanks for listening, everyone — take care.