Pipeline: Knowing vs Saying
The paper separates observation from manipulation: detect a latent count, measure its angle to digit-token rows, then test whether fixing the readout or rerouting hidden states changes behavior.
Mid-layer representations carry a clean scalar quantity signal.
The output head’s digit rows are nearly orthogonal to that signal.
Small output edits help constrained decoding; routing edits help full generation.
Family Snapshot
Mock values follow the transcript’s qualitative pattern: probe recovery generalizes better than the full intervention stack.
Readout Bottleneck
A single hidden direction can be highly informative while still projecting weakly onto the digit basis.
Probe Space vs Output Space
Switch model families to compare two views: a layer-by-digit cosine matrix and a layer-wise probe R² trace. Hover any heatmap cell for the exact alignment estimate.
low alignment
medium
high
Interpretation
The sharp split is the main visual punchline: the model can linearly expose quantity in hidden space while failing to place that signal onto token logits for “0” through “9.”
What the Interventions Actually Fix
Toggle between evaluation modes. The 9-row edit mostly rescues digit selection when the answer space is restricted. LoRA-style Q/V routing changes are what move the open-vocabulary generation curve.
Digit Rank Collapse
The transcript highlights the dramatic logit-lens shift: the correct digit can go from buried in the vocabulary to rank one after routing is fixed.
Parameter Footprint
One intervention repaints the final scoreboard. The other rewires the route into it.
Autoregressive Failure Path
Step through the generation pipeline. In open-vocabulary decoding, the count signal must survive routing, basis mismatch, and competition from every other token before the model can emit a digit.
Token Competition Heatmap
Counting is easy to understand but harsh on the interface: a tiny answer set has to fight a massive vocabulary during normal decoding.
Practical Levers
The paper suggests a narrow but useful lesson: exact numeracy evals should separate representation quality from output formatting and tokenization constraints.
References
Compact source set used in the episode and transcript. arXiv links are included where an ID is explicit in the provided materials.
Why Transformers Fail at Counting • arXiv:2605.03258
Teaching Arithmetic to Small Transformers • McLeish, Irving, Sokota, Black, Chen, et al. (2024)
Language Models Use Trigonometry to Do Addition • McLeish et al. (2024)
Faithfulness of Linear Probes in Transformers • probe-critique literature (2019–2024)
ROME: Locating and Editing Factual Associations in GPT • Meng, Bau, Andonian, Belinkov (2022)
The Geometry of Truth • Gurnee et al. (2023)
A Mathematical Framework for Transformer Circuits • Elhage et al. (2021)
Finding Transformer Circuits with Edge-Level Attribution Patching • Nanda et al. (2023)
Tokenization Counts • Singh, Strouse (2024)
Efficient Numeracy Through Single-Token Number Embeddings • Kreitner et al. (2025)
Arithmetic-Based Pretraining Improving Numeracy • Petrak, Moosavi, Gurevych (2023)
Rethinking Weight Tying • Gu, Aleti, Chen, Zhang (2026)
Latent Causal Probing • Jin (2024)