1 00:00:01,000 --> 00:00:28,175 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. And today we're closing out a three-part arc, so buckle up. We're covering "Understanding and Steering the Cognitive Behaviors of Reasoning Models at Test-Time," by Zhenyu Zhang et al. — that's ten co-authors total — out of University of Texas at Austin, Together AI, and University of Sydney. This one hit arXiv in December 2025. 2 00:00:28,175 --> 00:01:01,599 [Dr. Ada Shannon] Alright, let's get into how they actually pin these heads down, Hal. They cut the chain-of-thought wherever there's a double newline, so each chunk becomes one discrete reasoning step. Then they re-run the whole trace as a single prefill and grab the hidden state sitting exactly at that delimiter token, for every attention head, every layer — a compact summary of what just happened in that step. Per head, they fit a tiny linear probe: given that one vector, predict linear or non-linear. No deep classifier, just logistic regression, one per head, across the network. The heads that nail that prediction land in the top 10 percent and become the candidates. 3 00:01:01,599 --> 00:01:37,349 [Hal Turing] So it's basically a report card for every head in the model — some heads ace the linear-versus-non-linear quiz, most don't. What I like is how unglamorous the method is. No probing classifier zoo, no contrastive loss, just per-head logistic regression on a single token's hidden state. But doesn't the resulting vector risk carrying a lot of noise? You're deriving 'the' non-linear direction from raw hidden states pulled from a limited set of examples — that seems like it could pick up all sorts of incidental structure riding along with the real signal. 4 00:01:37,349 --> 00:02:17,050 [Dr. Ada Shannon] Good instinct — the paper flags exactly that: the raw head vector is entangled with a noise component alongside the true reasoning direction. Their fix is a shared-subspace PCA. Instead of a separate eigenspace per head, which would leave every head in its own private coordinate universe, they pool activations across all heads in a layer, compute one covariance matrix, and eigen-decompose it. Most of the signal collapses into a tiny number of principal components — barely a hundred dimensions in a space that's thousands wide. Every head's vector then gets projected onto that shared low-rank basis. It's denoising by borrowing statistical strength across heads in the same layer. 5 00:02:17,050 --> 00:02:37,100 [Hal Turing] Oh wait, hold on — before you go further into eigenvectors, I need the payoff. What actually happens when they flip this thing on, live, mid-generation? Because the walkthrough in the paper is kind of wild — they literally pause the model mid-trace and watch it fork onto a completely different path depending on which way they nudge it. 6 00:02:37,100 --> 00:03:19,100 [Dr. Ada Shannon] Right, that's the fun part. Right after the model emits the double-newline ending a step, they rotate that token's hidden state — a full norm-preserving rotation, equation 5 in the paper — zeroing out the component along the steering direction, or amplifying it, without changing the activation's magnitude. That's the whole point: no separate strength knob to hand-tune per model, unlike prior steering work. And the walkthrough example is almost comically mundane — converting the point zero-comma-three into polar coordinates. Pause at step 9, suppress the non-linear direction, and the model skips the 'alternatively, let me reconsider' detour, landing the answer in 12 steps flat. Pause at step 10 instead and amplify that direction, and it spirals into a 45-step hand-wringing session — still correct, just the long way around. 7 00:03:19,100 --> 00:03:45,050 [Hal Turing] Forty-five steps to confirm that a point sitting on the positive y-axis has an angle of pi over two. That's the reasoning equivalent of triple-checking you locked a door you're currently standing in front of. But that example does the paper's job for it — it makes 'redundant reasoning' concrete instead of an abstract efficiency number. So walk me through what that buys across a real benchmark suite, not just one converted coordinate. 8 00:03:45,050 --> 00:04:36,125 [Dr. Ada Shannon] On R1-1.5B and AMC23: vanilla gets 72.5% accuracy burning about 8,951 tokens; CREST gets 90% at 5,584 tokens — a 17.5-point accuracy gain and a 37.6% token cut simultaneously, the headline number in the abstract. Not a one-model fluke. On R1-7B, token savings hold above 30% on both AIME25 and pooled AIME22-24, with accuracy at least matching vanilla. Push to R1-32B and GPT-OSS-20B — that last one's mixture-of-experts, not the dense architecture the method was built around — and gains persist, up to 6.7 points on GPT-OSS on AIME25, with real token reduction. Same story on Qwen3. Four architectures, dense and MoE, and the effect's direction never flips. 9 00:04:36,125 --> 00:05:17,750 [Hal Turing] And here's the part that surprised me most: all of that calibration — the probing, the steering vectors — happens exclusively on MATH500. Five hundred math problems. Then they take those exact vectors, no re-fitting, and throw them at LiveCodeBench for code generation, at GPQA-D for graduate-level science questions, and at Calendar Planning, which isn't math or code at all, it's scheduling logic. The numbers move the right direction across the board — GPQA-D on R1-32B jumps from 32.3% to 40.9%, which isn't a small bump on a benchmark literally designed to resist shortcuts. 10 00:05:17,750 --> 00:05:57,300 [Dr. Ada Shannon] It's a striking transfer result. Though I'll flag something the paper never quite squares away: how many heads they're steering shifts by section. Back in the Section 3.3 demonstration — the polar-coordinates walkthrough — they're steering the top 7% of heads. Section 4.1.1, defining the actual calibration procedure for CREST, keeps the top 10%. But Section 5.3.1's ablation, sweeping the ratio from 4% up past 90%, finds the strongest accuracy-and-efficiency balance around 38%, their 'gold ratio.' Three numbers, three sections, and nowhere does the paper state which ratio actually produced the Table 1 through 3 results we've been quoting. 11 00:05:57,300 --> 00:06:22,175 [Hal Turing] Huh — 7, 10, or 38, and we're supposed to just infer which one is load-bearing for the headline numbers. That's worth sitting with, because it's not cosmetic — it changes what 'the method' even means operationally. I want to come back to that, along with a couple other things nagging at me about how those labels get generated in the first place, right after we get through the rest of what this paper claims. 12 00:06:22,175 --> 00:07:02,750 [Hal Turing] Alright — the finale. Quick recap for anyone just joining: 'Cognitive Behaviors Behind Self-Improving Language Model Reasoners' told us which mental moves — verification, backtracking, subgoal setting, backward chaining — gate whether a model self-improves. 'Thought Anchors: Which Sentences Really Drive LLM Reasoning' showed us where those moves live inside a trace. Neither paper cites the other — that arc is ours. Today's paper is different: it actually cites the cognitive-behaviors work directly and asks whether you can reach into the model and steer that machinery. We spent Part 2 admiring the results, Ada. Time to stress-test them. 13 00:07:02,750 --> 00:07:45,425 [Dr. Ada Shannon] Start with the foundation. 'Non-linear reasoning' here isn't verification or backtracking in any rich sense — it's ten trigger words: Wait, Alternatively, hold on, and so on. That's a keyword shortcut for the richer taxonomy in Gandhi, Chakravarthy, Singh, Lile, and Goodman's 'Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs,' Stanford, 2025 — the paper behind our first episode. CREST borrows that vocabulary but never checks the keyword labels against human judgment. No inter-rater agreement, no precision numbers. So when a head 'predicts non-linear reasoning at 85% accuracy,' what it provably predicts is 'does this chunk contain one of ten strings.' That's a smaller, shakier claim than the paper's framing implies. 14 00:07:45,425 --> 00:08:27,050 [Hal Turing] And that shakiness compounds with the head-ratio mess we flagged at the end of Part 2. Here's the catch: that 38% 'gold ratio' ablation ran on exactly two models — R1-1.5B and R1-7B — on one benchmark, AIME22-24. Then it gets frozen as the universal default across five model families and six benchmarks, including R1-32B, Qwen3-30B, and the mixture-of-experts GPT-OSS-20B, none of which were in the tuning set. The one hyperparameter in a 'training-free, plug-and-play' method was itself tuned on a sliver of the eval suite. 15 00:08:27,050 --> 00:09:03,700 [Dr. Ada Shannon] That assumption gets tested hardest in Table 3, and the paper calls the result 'consistent improvements.' But look at Qwen3-30B on LiveCodeBench: tokens go from 15,307 to 15,317 — that's the wrong direction. GPQA-D on the same model: 70.20 to 70.20, dead flat. Those rows sit right beside double-digit wins, and the text just skips past them. Which is the question I keep circling: is the steering vector really capturing a domain-general reasoning mode, or is it closer to a length-suppression direction that correlates with the keyword set on math and does much less once you leave math? 16 00:09:03,700 --> 00:09:46,400 [Hal Turing] Oh — wait, hold on, that's actually where I wanted to go next, because there's something the paper never checks at all: whether these heads do anything else. Wu, Wang, Xiao, Peng, and Fu's 'Retrieval Head Mechanistically Explains Long-Context Factuality,' out of University of Washington, 2024, identified a distinct class of heads responsible for pulling facts out of long context. CREST never tests overlap between cognitive heads and retrieval heads. If they overlap, suppressing them to shorten a math proof could quietly degrade factual recall in a RAG or agentic setting — and none of these six benchmarks, all single checkable-answer tasks, would ever surface that. 17 00:09:46,400 --> 00:10:29,225 [Dr. Ada Shannon] That's the deeper worry — not just factuality, but self-correction generally. Backtracking and verification should matter most exactly when the model's first instinct is wrong. CREST never breaks results down by first-attempt-correct versus first-attempt-wrong, so a net accuracy gain could hide a redistribution — some previously-right answers flipped wrong because the safety net got suppressed, offset by previously-wrong ones fixed because a redundant loop got cut. And with AIME25 at roughly thirty problems and AMC23 at forty, a swing from 20% to 30% is a handful of questions, sampled at temperature 0.6, with no seeds or error bars reported anywhere. Practically though, I'd still say this is worth trying — the overhead is close to nothing and it doesn't touch weights. 18 00:10:29,225 --> 00:11:21,300 [Hal Turing] Which is exactly why I wouldn't throw this out — I'd just treat the 38% ratio and the MATH500 calibration as starting points, not settled defaults, until someone re-validates them per model and per domain, checks the keyword labels against real human annotation, and runs that retrieval-head overlap check before deploying this near a long-context or RAG pipeline. So here's the arc, start to finish: 'Cognitive Behaviors Behind Self-Improving Language Model Reasoners' told us which mental moves matter, 'Thought Anchors' showed us where they live inside a trace, and this paper showed you can actually reach in and turn those moves up or down at inference time, no retraining required. Real contribution, promising direction, headline numbers I'd want re-checked before trusting them at face value. Thanks for sticking with us through all three parts — catch you next time.