AI Post Transformers · Episode Companion

Thought Anchors: Which Sentences Really Drive LLM Reasoning

Paul C. Bogdan & Uzay Macar (co-first, coin-flip order) · Neel Nanda & Arthur Conmy (senior, coin-flip order) — MATS / Anthropic
⌁ arXiv:2506.19143 DeepSeek R1‑Distill‑Qwen‑14B · Qwen3‑30B‑A3B · R1‑Distill‑Llama‑8B 3 convergent causal methods thought-anchors.com

Why the sentence is the unit of analysis

Classic mechanistic interpretability dissects a single forward pass. A reasoning model instead unrolls thousands of sequentially-dependent tokens — the old unit of analysis doesn't exist here. The paper's fix: treat each sentence in the chain-of-thought as the atomic unit, and ask which ones are "thought anchors" — sentences with outsized causal weight over everything downstream.

token-level (old toolkit) sentence-level (this paper) thought anchor

Three independent measurements, one target

A black-box behavioral method, an attention-structure method, and a causal-masking method — different failure modes, same converging answer.

Eight-category taxonomy

Adapted from Venhoff et al.'s 2025 steering-vector taxonomy, used to classify every sentence in every trace.

MATH problem 4682 — base-16 66666 → binary bit count

Naive answer: 20 (five hex digits × 4 bits). Correct answer: 19 — a leading zero the naive method misses. Resampling sentence-by-sentence shows accuracy declining through sentences 6–12, then a hard spike at sentence 13.

20
naive (wrong) answer
19
correct answer
#13
pivot sentence
Sentence 13: "Alternatively, maybe I can calculate the value in decimal and then find out how many bits that would require." — one pivot sentence rescues an otherwise-wrong trace.

Sentence-to-sentence causal graph

Masking attention to one sentence and watching the KL-divergence ripple into later logits surfaces a self-correction scaffold.

1 · Counterfactual resampling

Knock out a sentence, regenerate 100 continuations, keep only resamples whose embedding cosine similarity falls below the dataset median (0.8, all‑MiniLM‑L6‑v2) — i.e. genuinely different, not a reworded duplicate. Then compute KL divergence between answer distributions.

2 · Receiver heads

Attention heads whose focus repeatedly narrows onto a small set of earlier sentences (high kurtosis), concentrated in later layers.

Reliability of the effect

Split-half reliability across problem sets, and sentence-attention correlation among the 16 highest-kurtosis heads vs. a random pair.

3 · Attention suppression (causal masking)

Mechanically delete a sentence's attention access, measure the downstream logit shift via KL divergence — a true causal intervention, not correlation.

Category importance — Counterfactual (KL) lens

Same 40-trace dataset, two different importance lenses. Switch the toggle — the ranked categories change substantially.

Cross-model replication

Qwen3‑30B‑A3B on 2,500+ MMLU problems, and R1‑Distill‑Llama‑8B — same category dominance holds.

Receiver-head ablation vs. random ablation

Only at 512 of 1,920 heads (27%) does receiver ablation significantly outperform random ablation at hurting accuracy (t(31)=2.55, p=.02). Long reasoning traces carry heavy redundancy that must be overwhelmed first.

What's still open

Sentence-level triangulation across all three methods is validated deeply on one case study; the resampling-vs-masking correlation across the full dataset is modest (r=.20 overall, r=.34 for nearby sentences). No sweep against a stronger embedding model than MiniLM. Real signal, real open questions.

References