Classic mechanistic interpretability dissects a single forward pass. A reasoning model instead unrolls thousands of sequentially-dependent tokens — the old unit of analysis doesn't exist here. The paper's fix: treat each sentence in the chain-of-thought as the atomic unit, and ask which ones are "thought anchors" — sentences with outsized causal weight over everything downstream.
A black-box behavioral method, an attention-structure method, and a causal-masking method — different failure modes, same converging answer.
Adapted from Venhoff et al.'s 2025 steering-vector taxonomy, used to classify every sentence in every trace.
Naive answer: 20 (five hex digits × 4 bits). Correct answer: 19 — a leading zero the naive method misses. Resampling sentence-by-sentence shows accuracy declining through sentences 6–12, then a hard spike at sentence 13.
Masking attention to one sentence and watching the KL-divergence ripple into later logits surfaces a self-correction scaffold.
Knock out a sentence, regenerate 100 continuations, keep only resamples whose embedding cosine similarity falls below the dataset median (0.8, all‑MiniLM‑L6‑v2) — i.e. genuinely different, not a reworded duplicate. Then compute KL divergence between answer distributions.
Attention heads whose focus repeatedly narrows onto a small set of earlier sentences (high kurtosis), concentrated in later layers.
Split-half reliability across problem sets, and sentence-attention correlation among the 16 highest-kurtosis heads vs. a random pair.
Mechanically delete a sentence's attention access, measure the downstream logit shift via KL divergence — a true causal intervention, not correlation.
Same 40-trace dataset, two different importance lenses. Switch the toggle — the ranked categories change substantially.
Qwen3‑30B‑A3B on 2,500+ MMLU problems, and R1‑Distill‑Llama‑8B — same category dominance holds.
Only at 512 of 1,920 heads (27%) does receiver ablation significantly outperform random ablation at hurting accuracy (t(31)=2.55, p=.02). Long reasoning traces carry heavy redundancy that must be overwhelmed first.
Sentence-level triangulation across all three methods is validated deeply on one case study; the resampling-vs-masking correlation across the full dataset is modest (r=.20 overall, r=.34 for nearby sentences). No sweep against a stronger embedding model than MiniLM. Real signal, real open questions.