1 00:00:01,000 --> 00:00:39,425 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. And Ada, today we're digging into 'Thought Anchors: Which LLM Reasoning Steps Matter?' The co-first authors are Paul C. Bogdan and Uzay Macar — equal contribution, author order literally decided by a coin flip — with Neel Nanda and Arthur Conmy as senior authors, also coin-flipped for order. Four authors total, all out of MATS, the ML Alignment and Theory Scholars program, with Anthropic on the senior-author side. It went up on arXiv back in June 2025. 2 00:00:39,425 --> 00:01:01,200 [Dr. Ada Shannon] Yeah, this one grabbed me for a specific reason — a lot of interpretability papers on reasoning models feel like they're reading tea leaves in the model's output text and calling it a day. This one actually builds theoretical machinery underneath the claims — three separate, independently-derived measurements that are supposed to converge on the same answer. That's a much higher bar than 'the model said Wait so it must be reconsidering.' The grounding is what sold me. 3 00:01:01,200 --> 00:01:50,100 [Hal Turing] Which is a nice bridge, actually, because this is part two of a three-part arc we're running on the cognitive machinery of reasoning models. Last time out we covered 'Cognitive Behaviors Behind Self-Improving Language Model Reasoners' — quick, credited recap: that paper showed whether a base model already has behaviors like verification, backtracking, subgoal setting, and backward chaining determines whether reinforcement learning self-improvement even takes off. I want to be upfront, though — today's paper doesn't cite that one. There's no direct lineage. The connection is our framing, not theirs. That earlier episode asked which behaviors make a reasoner improve. Today we're asking a narrower, more mechanistic question: within one single reasoning trace, which concrete sentences actually carry the reasoning? 4 00:01:50,100 --> 00:02:28,425 [Dr. Ada Shannon] Right, and that reframing matters because standard mechanistic interpretability wasn't built for this. Think about the classic circuits work — Kevin Wang and colleagues out of Redwood Research finding the indirect-object-identification circuit in GPT-2 small back in 2022, or Heimersheim and Janiak's docstring circuit. Those methods dissect one forward pass: activations flow layer to layer, you get an output, done. A reasoning model doesn't do that. It generates a token, feeds that token back in, generates the next one, sometimes for thousands of tokens before it commits to an answer. The unit of analysis the old toolkit assumes — single pass, fixed input — just doesn't exist here. 5 00:02:28,425 --> 00:02:40,451 [Hal Turing] So if you can't just crack open one forward pass, what's the unit you use instead? Tokens feel too granular, and I'd guess a whole paragraph is too coarse to pin down what mattered. 6 00:02:40,451 --> 00:03:19,001 [Dr. Ada Shannon] Exactly the tension the authors land on, and their answer is the sentence. They call the important ones 'thought anchors' — sentences with outsized causal influence on the final answer and on everything that comes after them in the trace. And they don't just eyeball it; they build three converging measurements. First, black-box counterfactual resampling: knock out a sentence, regenerate a hundred times, see how much the answer distribution shifts. Second, receiver heads — attention heads that consistently narrow their focus back onto a small handful of earlier sentences. Third, attention suppression, where you mask attention to one sentence and measure the causal damage to a later one via KL divergence. 7 00:03:19,001 --> 00:03:36,001 [Hal Turing] Wait — oh, sorry, jumping in — the counterfactual importance one has a subtlety, right? It's not just 'does removing this sentence change the answer,' because a resampled replacement could just be a boring rewording of the same sentence, and that would tell you nothing. 8 00:03:36,001 --> 00:04:16,451 [Dr. Ada Shannon] Good catch, and yes — that's exactly why they filter. They only count a resample as informative if its embedding is meaningfully different from the original, cosine similarity below the dataset median of 0.8 using all-MiniLM-L6-v2 embeddings. Only then do they compute the KL divergence between the answer distributions with and without the sentence. That's counterfactual importance. And to sort what kind of sentence tends to matter, they adapt an eight-category taxonomy from Constantin Venhoff and colleagues' 2025 steering-vector paper: Problem Setup, Plan Generation, Fact Retrieval, Active Computation, Uncertainty Management, Result Consolidation, Self Checking, and Final Answer Emission. 9 00:04:16,451 --> 00:04:21,151 [Hal Turing] Give us something concrete to hang this on before we go further. 10 00:04:21,151 --> 00:05:00,226 [Dr. Ada Shannon] Their case study is a great hook. MATH problem 4682: 'when the base-16 number 66666 is written in base 2, how many bits does it have?' Naive answer is 20 — five hex digits, four bits each. Correct answer is 19, because of a leading zero the naive method misses. Running resampling sentence by sentence, accuracy actually declines through sentences six to twelve, then spikes hard at sentence 13, where the model pivots: 'Alternatively, maybe I can calculate the value in decimal and then find out how many bits that would require.' That one pivot sentence is what rescues the trace from the wrong answer. 11 00:05:00,226 --> 00:05:18,551 [Hal Turing] That lines up with something I remember from the s1 paper — Niklas Muennighoff and colleagues, 2025 — where forcing the model to backtrack with 'Wait' tokens at inference time actually boosted accuracy. So is this paper basically confirming that finding mechanistically? 12 00:05:18,551 --> 00:05:34,251 [Dr. Ada Shannon] It's consistent with it, but I'd stop short of 'confirming.' s1 showed a behavioral correlation — force more Wait tokens, get better scores. This paper is trying to show the sentence itself is doing causal work, which is a stronger and different claim. 13 00:05:34,251 --> 00:05:51,576 [Hal Turing] Sure, but honestly, once you've got a sentence whose removal visibly swings the answer distribution like that, doesn't that basically settle it? The chain-of-thought text is telling you what's actually driving the computation — I don't see what's left to argue about faithfulness. 14 00:05:51,576 --> 00:06:11,576 [Dr. Ada Shannon] No, no — that's too fast, Hal. There's a real, unresolved literature here. Miles Turpin and colleagues in 2023 showed language models can produce CoT explanations that sound plausible but don't reflect the actual reasons behind an answer. One case study lighting up at sentence 13 doesn't retire that debate — it's one problem, one trace. 15 00:06:11,576 --> 00:06:29,726 [Hal Turing] Fair — and honestly that's the right instinct, because they do scale this up well past one case study, which we'll get into. Let's hold that thread, because it's exactly the kind of question worth sitting with as we go through how they actually validate this across dozens of traces. 16 00:06:29,726 --> 00:06:56,276 [Hal Turing] They ran DeepSeek R1-Distill-Qwen-14B on the MATH benchmark, targeting problems the model solves correctly twenty-five to seventy-five percent of the time — that's where you actually get variance in the final answer. Twenty problems, one correct and one incorrect trace each, forty traces total. For every sentence in every trace, they generate a hundred resampled continuations and track how the final-answer distribution shifts. 17 00:06:56,276 --> 00:07:22,726 [Dr. Ada Shannon] That buys them something the earlier forced-answer method can't get — that's Lanham and colleagues at Anthropic, 'Measuring Faithfulness in Chain-of-Thought Reasoning,' 2023, where you interrupt the model mid-trace and force an answer. The blind spot: if a sentence matters but the model reliably produces it late regardless of what came before, forcing an early answer makes everything before it look useless — including the planning sentence that set up the whole approach. Resampling forward instead of truncating avoids that trap. 18 00:07:22,726 --> 00:07:39,976 [Hal Turing] Which plays out exactly in the case study — sentence thirteen, the pivot to calculating in decimal, is the one that swings accuracy hardest. Sentence twelve flags the leading-zero concern, but only sentence thirteen's pivot actually moves accuracy. 19 00:07:39,976 --> 00:08:08,151 [Dr. Ada Shannon] And that single case generalizes cleanly across all forty traces. With the real counterfactual importance metric — KL divergence, conditioned on the resample being semantically different — Plan Generation and Uncertainty Management post the highest scores, well above Fact Retrieval or Active Computation. Switch to forced-answer importance, and Active Computation suddenly looks dominant, because arithmetic sentences sit right before the model locks in. Same data, wrong lens, opposite conclusion about what's driving the reasoning. 20 00:08:08,151 --> 00:08:29,777 [Hal Turing] So that settles the black-box side pretty convincingly. What I want to know is whether anything's happening mechanically inside the model that actually lines up with it — because resampling tells you a sentence matters, it doesn't tell you how the model itself is using that sentence downstream. Is there something structural we can point to? 21 00:08:29,777 --> 00:08:49,052 [Dr. Ada Shannon] That's where receiver heads come in. They take every attention head, build a sentence-by-sentence matrix from its token weights, and measure how sharply it narrows attention toward a handful of past sentences — that's kurtosis, basically how spiky the distribution is. A cluster of heads, concentrated more in later layers— 22 00:08:49,052 --> 00:08:59,252 [Hal Turing] Wait — hold on — 'a cluster of heads' isn't a number. Is this a handful of heads doing something interesting, or a robust, repeatable phenomenon? 23 00:08:59,252 --> 00:09:33,802 [Dr. Ada Shannon] Fair — split-half reliability across problems is point-eight-four. A head that's a receiver on one set of problems is reliably a receiver on a completely different set — that's not noise. And among the sixteen highest-kurtosis heads, the correlation in which sentences they attend to is point-five-six on average, versus point-three-five for a random pair. Receiver heads converge on the same sentences as each other. And the punchline: those sentences are overwhelmingly Plan Generation and Uncertainty Management — the exact categories resampling already flagged. Two independent measurements, same answer. 24 00:09:33,802 --> 00:10:02,602 [Hal Turing] They push one layer further with the sentence-to-sentence causal graph — masking attention to one sentence and watching the KL-divergence ripple into later logits. In the case study, three links pop out: sentence twelve into forty-three, where the model verifies and lands on nineteen bits; forty-four into sixty-five, re-checking the arithmetic; and twelve into sixty-six, connecting the leading-zero suspicion to the explanation. Strung together, that's a visible self-correction scaffold. 25 00:10:02,602 --> 00:10:36,477 [Dr. Ada Shannon] They scale that machinery way up too, switching to Qwen3-30b-a3b on twenty-five hundred-plus MMLU problems, since it exposes logits cheaply enough to run thousands of traces. Strong close-range causal links — sentence to the one right next door — track higher accuracy, especially in math and physics. Long-range links, jumping way back in the trace, track lower accuracy and more uncertainty, the model scrambling instead of executing a tight plan. Cross-model replication on R1-Distill-Llama-8B holds up too, with the same category dominance in both resampling and receiver-head scores. 26 00:10:36,477 --> 00:10:57,452 [Hal Turing] Okay, but I want to push on something — I think you're about to tell me this proves receiver heads and resampling are literally the same mechanism, and I don't think convergence gets you there. Two methods agreeing certain sentences matter doesn't mean attention is the causal channel. It could just mean both are picking up the same surface feature. 27 00:10:57,452 --> 00:11:17,577 [Dr. Ada Shannon] I actually disagree with you there, Hal. Attention suppression isn't correlation — you're mechanically deleting the model's access to a sentence and measuring the downstream logit shift. That's a causal intervention, full stop. A black-box method, an attention method, and a causal-masking method, three different failure modes, all land on the same sentences. 28 00:11:17,577 --> 00:11:36,352 [Hal Turing] Sure, but 'three tools converge' and 'we've found the mechanism' aren't the same claim — and that gap is worth sitting with, because it's exactly what determines how much weight this result can actually carry once you push past one case study and a couple thousand traces across two models. 29 00:11:36,352 --> 00:11:53,178 [Dr. Ada Shannon] Fair — and that's honestly the right note to land this stretch on. How far that convergence actually generalizes beyond this pair of models, and where each of the three methods quietly carries its own blind spot, is worth its own close look before we call any of this settled. 30 00:11:53,178 --> 00:12:27,203 [Hal Turing] Alright, so before we get to that close-look Ada promised, there's one thing nagging at me from the appendix. Both the counterfactual importance metric and the receiver-head analysis lean on a cosine similarity cutoff — median 0.8 — computed from a pretty small, generic sentence embedder, all-MiniLM-L6-v2. That threshold is deciding what counts as 'genuinely different' every single time they resample or match a causal link. How much of this whole edifice is riding on that one modest embedding model's judgment call? 31 00:12:27,203 --> 00:13:06,853 [Dr. Ada Shannon] It's a fair worry, and I'll be straight with you — the paper doesn't run the robustness check you'd want, no sweep against a bigger sentence-transformer or an LLM-based embedder, no threshold sensitivity analysis. So yes, 0.8 from MiniLM is somewhat arbitrary on its face. What tempers it for me is that the qualitative story — planning and uncertainty sentences beating active computation — survives two independent stress tests they did run: the additive-smoothing variants in Appendix A, and the full cross-model replication on R1-Distill-Llama-8B. That's not the same as validating the embedder, but it means the finding isn't a knife-edge artifact of one similarity number. 32 00:13:06,853 --> 00:13:31,878 [Hal Turing] Wait, hold on — since we're doing appendix archaeology, I have to bring up Section L, because this one actually worried me more. The receiver-head ablation — they pull heads out of the model to see if it breaks reasoning, and at 256 heads out of 1920, receiver ablation and random ablation are basically indistinguishable. Doesn't that undercut the entire 'these are a specialized mechanism' framing? 33 00:13:31,878 --> 00:14:31,304 [Dr. Ada Shannon] It's the honest part of the paper, and I respect that they published it instead of burying it. At 128 heads: nothing. At 256 heads: 48.8 percent accuracy for receiver ablation versus 52.7 for random — not significantly different. Only at 512 heads, over a quarter of all attention heads, does receiver ablation actually hurt more, 27.7 versus 37.3 percent, t of 31 equals 2.55, p equals .02. So the effect is real and statistically significant, just only detectable with a blunt, large-scale intervention. The authors themselves note they have no prior baseline for what a 'typical' ablation scale even looks like in long-CoT models — this is uncharted territory, not necessarily a failed prediction. My read: long reasoning traces have serious redundancy and error-correction built in, so you have to overwhelm that redundancy before a specialized subsystem's contribution becomes visible. 34 00:14:31,304 --> 00:15:20,829 [Hal Turing] Okay, but step back with me for a second, because I think you're being too generous here, Ada. The flagship claim of this paper is convergence — three independent methods, resampling, receiver heads, causal masking, all pointing at the same sentences. But the deep version of that convergence, the one with the actual named sentence numbers, twelve to forty-three, forty-four to sixty-five — that's ONE problem. Twenty MATH problems and forty traces is the entire quantitative backbone for the 14B model. The MMLU run scales to twenty-five hundred problems, sure, but on a different model, without receiver heads or resampling at all. The three methods were never jointly validated at scale. That's a big gap between 'we found a mechanism' and 'we told a compelling story about one transcript.' 35 00:15:20,829 --> 00:15:53,079 [Dr. Ada Shannon] I actually disagree with you there, Hal — or at least with how strong you're putting it. Calling it 'one anecdote' undersells what's underneath the anecdote. The case study is the illustration, not the evidence. The quantitative claims — plan generation and uncertainty management scoring highest on counterfactual importance, receiver heads attending to those same categories — those come from the full 40-trace statistical comparison, not just problem 4682. And the split-half reliability of .84 on receiver heads is a real psychometric result on real sample size, not narrative color. 36 00:15:53,079 --> 00:16:18,154 [Hal Turing] Sure, the category-level stats are broader than n of one. But the thing being sold as the headline result — three methods converging on the SAME sentences in the SAME trace — that part is exactly the single case study. The category averages are one claim; the 'these methods triangulate on identical anchors' claim is a narrower and much less tested one, and it's the flashier claim in the abstract. 37 00:16:18,154 --> 00:16:52,604 [Dr. Ada Shannon] That's a sharper distinction than I was giving you credit for, and I'll take it — you're right that the paper blurs those two claims together more than it should. Worth noting, too: the paper's own numbers on this are modest. The correlation between resampling-based and masking-based sentence-to-sentence importance is r equals .20 overall, rising to .34 for nearby sentences. That's suggestive, not damning, but it's not a strong convergence number either. So: category-level findings, reasonably solid. Sentence-level triangulation as a general property, still a hypothesis worth taking seriously rather than a settled result. 38 00:16:52,604 --> 00:17:25,354 [Hal Turing] Which actually connects to something bigger — this whole approach is explicitly trying to sidestep the chain-of-thought faithfulness fight, right? Rather than asking 'does the CoT text honestly report what the model computed,' which is the Arcuschin and colleagues line of work, including Neel Nanda, from 2025, showing CoT reasoning in the wild isn't always faithful — this paper says: forget self-report, we can operationalize importance directly from resampling and attention. Does that actually dodge the faithfulness problem, or just relocate it? 39 00:17:25,354 --> 00:18:07,504 [Dr. Ada Shannon] It genuinely dodges part of it — a sentence can be causally load-bearing whether or not its stated reasoning is an accurate narration of the computation, and that's a real methodological contribution. It doesn't dodge all of it, though; Yanda Chen and colleagues at Anthropic showed this year that models don't always say what they think even under pressure to be transparent, so mechanistic importance and honest self-report can still come apart. There's also a nice complement here to Constantin Venhoff and colleagues out of Oxford, whose steering-vector paper this taxonomy is adapted from — they intervene on representations to trigger a reasoning function, this paper intervenes on text and attention to find where that function lives. Two different causal levers converging on the same eight categories is a stronger argument than either alone. 40 00:18:07,504 --> 00:18:43,129 [Dr. Ada Shannon] And there's a real practical payoff if the uncertainty-management anchors hold up: it gives Niklas Muennighoff's s1 budget-forcing trick a mechanistic story. If forcing a 'Wait, let me reconsider' isn't just adding tokens but actually inserting a high-leverage anchor, that suggests smarter test-time compute strategies — forcing anchors selectively rather than just padding length. They've also open-sourced thought-anchors.com, which lets anyone visualize this on new traces, and that's genuinely useful for debugging why a model's reasoning went sideways, independent of how the underlying theory eventually holds up. 41 00:18:43,129 --> 00:19:33,454 [Hal Turing] So where's this actually heading? Adaptive-scale reasoning units instead of fixed sentence boundaries, a proper formal treatment of overdetermination so multiple sufficient causes don't get double-counted, and — we keep coming back to it — somebody needs to run this against a stronger embedding model before the taxonomy gets treated as settled. To wrap the arc: across three episodes we went from which cognitive behaviors let a base model self-improve under RL, to which sentences carry the reasoning within one trace, to today — how much weight that sentence-level story can actually bear. The honest answer is: real signal, real open questions, and a proof-of-concept the authors themselves never oversold as more than that. Thanks for listening to this three-part run with us — we'll see you next time on AI Post Transformers.