A fine-tuned model recalls a new fact perfectly in isolation, then fails the moment it has to reason with it. This companion page traces the "Knowing-Using Gap" — and the layer-level misfiling that causes it.
RAG hands the fact to the model fresh at inference time. Model editing (ROME/MEMIT) overwrites the weights that store it directly. Fine-tuning is the only route meant to fold new knowledge into the model's existing reasoning — which is exactly the route that turns out to misfire.
Direct recall of a newly memorized fact saturates almost immediately. The ability to actually use that fact in a multi-step task shows up much later — and, for chaining tasks, plateaus far short of ceiling. Toggle the task type below.
Self-patching inherits its intervention style from ROME's causal tracing and Google DeepMind's PatchScope, but asks a different question: not where a fact lives, or what a representation decodes to — just whether relocating it flips the answer from wrong to right.
Pick a source layer and a target layer. The middle band (highlighted) is where chaining and intersection reasoning actually happen. Watch which relocations flip a wrong answer to a right one.
Each cell is the accuracy delta from patching the entity representation from a source layer into a target layer. Hover a cell for its exact reading. Step through training with the toggles below.
Full fine-tuning memorizes far faster than LoRA, but that speed doesn't transfer to usability: chaining lands at nearly the same low ceiling either way, and intersection actually gets worse under FFT.
Forcing the single best source/target layer pair per instance — the oracle — lifts chaining accuracy 1.5–6× across every model tested, from a 1.5B Qwen up to an 8B LLaMA. Intersection's smaller gap nearly closes outright.
Dropping the per-instance oracle search for two fixed layer pairs per architecture — no scanning, no knowing where a fact lives in advance — still recovers 58–75% of the oracle's headroom.
Effective patch sources cluster into early layers (~0.10L) and late layers (~0.82L), both feeding the same mid-layer reasoning band (~0.45L). Late-to-late patching does nothing — the information has already moved on.
Entity-position patching dwarfs every control: random tokens, beginning-of-sequence, chain-of-thought prompting, and patching an unrelated fact's representation all lag far behind.