arXiv:2607.08393

The First Fact That Wouldn't Reason

A fine-tuned model recalls a new fact perfectly in isolation, then fails the moment it has to reason with it. This companion page traces the "Knowing-Using Gap" — and the layer-level misfiling that causes it.

Lu Dai, Ziyang Rao, Yili Wang, Hanqing Wang, Hao Liu, Hui Xiong  ·  HKUST(GZ) / HKUST  ·  July 2026

Three routes into a deployed modelWhere fine-tuning sits

RAG hands the fact to the model fresh at inference time. Model editing (ROME/MEMIT) overwrites the weights that store it directly. Fine-tuning is the only route meant to fold new knowledge into the model's existing reasoning — which is exactly the route that turns out to misfire.

The Knowing-Using GapTwo curves, one gap

Direct recall of a newly memorized fact saturates almost immediately. The ability to actually use that fact in a multi-step task shows up much later — and, for chaining tasks, plateaus far short of ceiling. Toggle the task type below.

Method lineageFrom causal tracing to self-patching

Self-patching inherits its intervention style from ROME's causal tracing and Google DeepMind's PatchScope, but asks a different question: not where a fact lives, or what a representation decodes to — just whether relocating it flips the answer from wrong to right.

Interactive demoPatch the entity representation

Pick a source layer and a target layer. The middle band (highlighted) is where chaining and intersection reasoning actually happen. Watch which relocations flip a wrong answer to a right one.

Patched entity representation from L1 into L4: wrong → right.

Source layer × target layerWatching the misfiling happen

Each cell is the accuracy delta from patching the entity representation from a source layer into a target layer. Hover a cell for its exact reading. Step through training with the toggles below.

LoRA vs full fine-tuningEpoch lag & final accuracy

Full fine-tuning memorizes far faster than LoRA, but that speed doesn't transfer to usability: chaining lands at nearly the same low ceiling either way, and intersection actually gets worse under FFT.

LoRA Full Fine-Tuning

Epoch lag (memorization → usable)

Final usable accuracy

Oracle patching, six modelsPre-patch vs best-layer-pair patch

Forcing the single best source/target layer pair per instance — the oracle — lifts chaining accuracy 1.5–6× across every model tested, from a 1.5B Qwen up to an 8B LLaMA. Intersection's smaller gap nearly closes outright.

Pre-patch Oracle-patched

Deployable versionFixed two-layer heuristic recovery

Dropping the per-instance oracle search for two fixed layer pairs per architecture — no scanning, no knowing where a fact lives in advance — still recovers 58–75% of the oracle's headroom.

Where the winning patches come fromTwo clusters, one dead zone

Effective patch sources cluster into early layers (~0.10L) and late layers (~0.82L), both feeding the same mid-layer reasoning band (~0.45L). Late-to-late patching does nothing — the information has already moved on.

Ruling out generic perturbationWhat actually drives the effect

Entity-position patching dwarfs every control: random tokens, beginning-of-sequence, chain-of-thought prompting, and patching an unrelated fact's representation all lag far behind.

References

  1. Towards Mechanistically Understanding Why Memorized Knowledge Fails to Generalize in Large Language Model FinetuningLu Dai, Ziyang Rao, Yili Wang, Hanqing Wang, Hao Liu, Hui Xiong · 2026
  2. Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan Belinkov · 2022
  3. MQuAKE: Assessing Knowledge Editing in Language Models via Multi-Hop QuestionsZexuan Zhong, Zhengxuan Wu, Christopher D. Manning, Christopher Potts, Danqi Chen · 2023
  4. Do Large Language Models Latently Perform Multi-Hop Reasoning?Sohee Yang, Elena Gribovskaya, Nora Kassner, Mor Geva, Sebastian Riedel · 2024
  5. Progress Measures for Grokking via Mechanistic InterpretabilityNeel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith, Jacob Steinhardt · 2023
  6. Physics of Language Models: Part 3.2, Knowledge ManipulationZeyuan Allen-Zhu, Yuanzhi Li · 2023
  7. Hopping too late: Exploring the limitations of large language models on multi-hop queriesEden Biran, Daniela Gottesman, Sohee Yang, Mor Geva, Amir Globerson · 2024
  8. Cake: Circuit-aware editing enables generalizable knowledge learnersYunzhi Yao, Jizhan Fang, Jia-Chen Gu, Ningyu Zhang, Shumin Deng, Huajun Chen, Nanyun Peng · 2025
  9. Model editing at scale leads to gradual and catastrophic forgettingAkshat Gupta, Anurag Rao, Gopala Anumanchipalli · 2024