1 00:00:01,000 --> 00:00:44,174 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today we're starting a three-part arc on what actually makes a language model good at self-improving through reinforcement learning. First up: a paper called 'Cognitive Behaviors that Enable Self-Improving Reasoners,' or, and I love this subtitle, 'Four Habits of Highly Effective STaRs.' That's Kanishk Gandhi et al., five co-authors total, out of Stanford University and SynthLabs, posted to arXiv in March 2025 and published as a conference paper at COLM 2025. 2 00:00:44,174 --> 00:01:06,025 [Dr. Ada Shannon] When I first skimmed this one I expected another 'here's a cool RL trick' paper. What actually sold me is how much it leans on real cognitive science and classical AI theory — verification, backtracking, that whole vocabulary — instead of just throwing compute at a model and hoping something emergent shows up. Most reasoning papers are pure empiricism: try it, see if the number goes up. This one has an actual hypothesis for why, and they test it directly. 3 00:01:06,025 --> 00:01:27,725 [Hal Turing] Let's set the stage, because the finding that kicked this off is genuinely surprising. They took two base models of roughly the same size, Qwen-2.5-3B and Llama-3.2-3B, ran identical reinforcement learning on both — same algorithm, same hyperparameters, same task — and got wildly different results. 4 00:01:27,725 --> 00:01:57,725 [Dr. Ada Shannon] This is PPO, Proximal Policy Optimization, Schulman et al. 2017, run for 250 steps on both models, four trajectories per prompt, using the VERL library with the TinyZero implementation. Both models start at roughly the same low baseline. But around step 30, Qwen hits what the authors call a qualitative shift — responses get longer, accuracy climbs. By the end, Qwen's at about 60% accuracy. Llama, same recipe, same steps, plateaus around 30%. Half the score, identical training. 5 00:01:57,725 --> 00:02:01,450 [Hal Turing] So what were they actually training these models to do? 6 00:02:01,450 --> 00:02:33,000 [Dr. Ada Shannon] Countdown — yes, like the British game show. You get a set of numbers and a target, and combine them with plus, minus, times, divide to hit the target exactly. Their example: numbers 25, 30, 3, 4, target 32, solution is thirty minus twenty-five plus three, times four. They picked it deliberately — restricted search space, so it's tractable to analyze, but solving it well requires real planning. And success depends on problem-solving ability rather than memorized math facts, so you're not confusing 'can this model reason' with 'does it just know more arithmetic.' 7 00:02:33,000 --> 00:02:42,949 [Hal Turing] Oh wait wait wait — hold on, we're benchmarking frontier reasoning models on the exact quiz format British retirees watch after lunch? 8 00:02:42,949 --> 00:02:59,699 [Dr. Ada Shannon] Pretty much. And it's a good testbed for exactly that reason — simple enough to verify automatically with zero ambiguity, but the search space is big enough you can't brute-force it, so it forces genuine multi-step reasoning instead of pattern-matching to something memorized. 9 00:02:59,699 --> 00:03:09,125 [Hal Turing] Okay, so same setup, wildly different learning curves. You said the paper isolates four behaviors to explain that gap — what are they? 10 00:03:09,125 --> 00:03:53,375 [Dr. Ada Shannon] Four cognitive behaviors, chosen because they're well-defined and you can actually spot them in a transcript. Verification is systematically checking a result against what you need — catching yourself mid-trace, like noticing a candidate answer isn't right and moving on. Backtracking is explicitly abandoning a failing approach and trying something else — 'let's try a different combination.' Those two show up constantly in Qwen's raw, untrained outputs, and barely at all in Llama's. Then subgoal setting — breaking the big problem into an intermediate target, like 'let's try to get to a multiple of ten' before tackling the whole thing. And backward chaining — starting from the target and reasoning back toward what would produce it: 'working backwards, 24 is 8 times 3.' That one's got real history — it's the inference strategy underlying Prolog and old expert systems from the seventies and eighties, long before anyone trained a transformer. 11 00:03:53,375 --> 00:04:11,474 [Hal Turing] So the claim isn't 'Qwen is just bigger or better.' It's that Qwen's pretraining left it already doing these four things by default, and Llama's didn't — and that gap, present before any RL even starts, is what predicts whether reinforcement learning takes off or stalls. 12 00:04:11,474 --> 00:04:33,074 [Dr. Ada Shannon] Exactly. Not raw scale, not the training algorithm — both models are the same size class, identical RL recipe. It's whether the behavioral scaffolding was already sitting in the base model's weights, ready for RL to grab onto and amplify. Which raises the obvious next question: is that gap fixed by whatever a model happened to be pretrained on, or can you actually go in and install these behaviors after the fact? 13 00:04:33,074 --> 00:04:50,074 [Hal Turing] Okay so before we get to whether you can install these behaviors after the fact, I have to ask the boring-but-necessary question: how do you even measure this at scale? You can't have a grad student read ten thousand reasoning traces and tally up backtracking by hand. 14 00:04:50,074 --> 00:05:13,574 [Dr. Ada Shannon] Right, and that's a real methodological problem they had to solve first. They built a classification pipeline using GPT-4o-mini as the judge, asking it four separate questions per trajectory, one per behavior, and having it count occurrences. They picked 4o-mini specifically to balance inference cost against classification capability, since they're running this over every checkpoint across the full RL trajectory, for multiple models. That's a lot of transcripts to grind through. 15 00:05:13,574 --> 00:05:51,699 [Hal Turing] And walking through the actual numbers, this is where it gets striking. Run this classifier over baseline, unmodified Qwen-2.5-3B, and it's already producing meaningfully high counts of all four behaviors before any RL happens. Llama-3.2-3B, almost none. Then they bring in a third model, Llama-3.1-70B, twenty times the size, and the behaviors do increase, but unevenly — verification and subgoal setting pick up, while backtracking specifically stays low even at 70B. Scale alone doesn't just hand you the missing behavior. 16 00:05:51,699 --> 00:06:26,899 [Dr. Ada Shannon] Which tells you it's not purely a capacity problem, it's something baked into how each family was trained. So the next move is the interventionist one — can you force these behaviors into a model that doesn't have them? They built seven priming datasets on Countdown problems, generated by Claude-3.5-Sonnet, each constrained by its system prompt to show exactly one combination: backtracking alone, backtracking plus verification, backtracking plus subgoal setting, backtracking plus backward chaining, and all four together. Claude was told which behaviors it could use and which were forbidden, so each dataset is a clean isolate. 17 00:06:26,899 --> 00:06:41,424 [Hal Turing] Oh, hold on — that's actually the part I wanted to push on, because my first thought reading this was, isn't this just giving the model more tokens to work with? More thinking room, more compute, that alone could explain a boost. 18 00:06:41,424 --> 00:07:03,824 [Dr. Ada Shannon] That's exactly the confound they built two controls for. One dataset is just an empty chain-of-thought, literally nothing between the think tags. The other is length-matched, same token count as the all-strategies dataset, but filled with placeholder dots instead of reasoning. Both controls land right back at baseline Llama performance, thirty to thirty-five percent. So it's not the extra compute budget doing the work, it's specifically the behavioral content inside it. 19 00:07:03,824 --> 00:07:29,349 [Hal Turing] Which sets up maybe the most surprising result in the whole paper. They also built a version of the all-strategies dataset where the reasoning shows the right behaviors but the final answers are simply wrong — Claude gets the arithmetic wrong but still backtracks, verifies, sets subgoals. Priming Llama on that broken-answer dataset produces essentially the same RL trajectory as priming on the correct-answer version. 20 00:07:29,349 --> 00:07:43,149 [Dr. Ada Shannon] Which is the whole thesis in one experiment. Llama didn't need to see correct math, it needed to see the shape of the search process. Get that shape in, even attached to wrong answers, and RL has something to amplify. 21 00:07:43,149 --> 00:08:11,749 [Hal Turing] So then the question is whether you need Countdown-specific priming at all, or whether you can bake this in earlier, at the pretraining stage. They took Qwen-2.5-32B as a classifier again, this time over OpenWebMath, sorted documents into a behavior-rich set and a behavior-minimized control, then had the same model rewrite each into a clean question-thought-answer format. Eight point three million tokens each, matched for size and math content, differing only in behavioral density. 22 00:08:11,749 --> 00:08:32,099 [Dr. Ada Shannon] And continued pretraining Llama on the behavior-rich set, then running the same RL recipe on Countdown, closes almost the entire gap with Qwen. The control set, same math content, same token count, doesn't. Which matters, because the fix doesn't have to live in task-specific fine-tuning at all — it can live upstream, in what the base model reads before it ever sees Countdown. 23 00:08:32,099 --> 00:08:56,424 [Hal Turing] One more thing worth flagging before we get into how solid all this is: when they checked how common these behaviors actually are in existing math pretraining corpora, OpenWebMath and FineMath, the natural rate was already low. Verification and backtracking barely show up on their own. So it's not that the internet is secretly full of backtracking and nobody noticed — you really do have to go curate for it. 24 00:08:56,424 --> 00:09:34,424 [Hal Turing] So that low natural rate is a good pivot into the harder question, because everything we've walked through — behaviors beating correctness, Qwen's edge, the priming fix, the pretraining fix — all of it rests on one task. Countdown. And one automated metric counting four behavior labels. The paper itself calls its priming method 'domain-specific' and even says future work needs to test whether this holds outside math. They do run a transfer check to GPQA and MATH in the appendix, but I want to be honest about what that actually shows before we call this a general theory of reasoning. 25 00:09:34,424 --> 00:10:12,699 [Dr. Ada Shannon] Right, and the transfer appendix is a good place to start, but let's go one layer deeper first, because the whole edifice depends on GPT-4o-mini being a trustworthy judge of what counts as 'backward chaining' in raw text. They ran an inter-rater reliability check against Claude and two human annotators using ICC3. Verification and subgoal setting land around 0.65 to 0.77 agreement with humans — respectable. Backward chaining is 0.55 and 0.40. That's weak agreement. And the headline claim — Qwen naturally does more backward chaining than Llama — is leaning on the one behavior where the automated count and actual humans disagree the most. 26 00:10:12,699 --> 00:10:54,349 [Hal Turing] Hold on — that's actually a bigger problem than just noisy labels, because look at who generated the data on each side. All the priming trajectories came from Claude-3.5-Sonnet. The pretraining classifier and rewriter is Qwen-2.5-32B — same family as Qwen-2.5-3B, the model that wins. So you've got a Qwen model deciding what counts as good Qwen-style reasoning, then rewriting pretraining text to match. If Qwen-32B has a stylistic bias toward Qwen-flavored phrasing, that could inflate both the 'Qwen naturally has these behaviors' finding and the fidelity of the behavior-enriched pretraining set. 27 00:10:54,349 --> 00:11:31,549 [Dr. Ada Shannon] It's a real confound and they don't rule it out. And it compounds with a second one: this is two model families, at one size, 3B parameters, fully RL-trained. Llama-3.1-70B only shows up as a static behavior count in Figure 4, never RL-trained — so we don't actually know if scale fixes this or if 'family' is standing in for differences in tokenizer, corpus size, or data curation that have nothing to do with cognitive behaviors. And there's a methods choice worth flagging too — they picked PPO over GRPO for stability, calling cross-algorithm performance only 'anecdotally similar,' no data shown. GRPO is what DeepSeek-R1 actually used. 28 00:11:31,549 --> 00:12:15,449 [Hal Turing] And the reward function only checks final-answer correctness plus a small format bonus — nothing in there directly rewards showing your work. So when they say RL 'selectively amplifies' backtracking and verification while suppressing backward chaining and subgoal setting, that could just mean those two behaviors correlate with shorter, cheaper-to-generate correct completions under this specific reward shape, not that they're less useful for reasoning generally. Then there's the transfer result itself, which I think undercuts the thesis more than the paper lets on: behavior-enriched Llama shows way more verification and subgoal setting on GPQA than the control, but accuracy is flat, both around twelve percent. 29 00:12:15,449 --> 00:13:07,149 [Dr. Ada Shannon] They attribute that to 'limited forward inference capability' rather than a knock on the theory, but you can read it the other way — behaviors present with zero accuracy gain outside the RL domain is exactly what you'd expect from Countdown-format overfitting, not a portable cognitive upgrade. Worth grounding this against the lineage too. Countdown itself comes from Gandhi and coauthors' own Stream of Search paper, Stanford, 2024 — this work asks which base models can productively absorb that kind of search training. And Yeo, Tong, Niu, Neubig, and Yue's Demystifying Long Chain-of-Thought paper, 2025, is the prior attempt at this exact question, but it only counted reflective phrases in pretraining text — that's the gap this paper says it's filling, though given the classifier caveats, the upgrade is only partially earned. 30 00:13:07,149 --> 00:14:05,524 [Hal Turing] There's also Dacheng Li and coauthors' 'Structure, not Content' paper out of UC Berkeley's Sky Computing Lab, 2025, showing reasoning transfers from corrupted trajectories if structure survives — this paper explicitly extends that into the RL-priming setting. And Zichen Liu and coauthors at Sea AI Lab published a skeptical pilot study in 2025 arguing R1-style 'aha moments' may already be latent in the base model rather than emergent from RL, which lines up with — and complicates — this paper's own framing that RL just amplifies what Qwen already has. Practically, if you buy the core result even with these caveats, it argues for curating RL data by behavior richness instead of correctness filtering, and treating pretraining composition as a real lever for which base models make good RL starting points. 31 00:14:05,524 --> 00:14:42,574 [Dr. Ada Shannon] The authors are upfront that their four behaviors — grounded loosely in Simon and Newell's 1971 human problem-solving framework — aren't exhaustive. They flag analogy-making and metacognitive awareness as open extensions, and explicitly ask whether the same pattern holds in coding, game-play, or creative writing, where task constraints might amplify totally different behaviors than backtracking and verification. That's the honest state of it: a real, causally-tested effect on one clean puzzle, with a scope-versus-claims gap between the abstract's 'fundamental relationship' language and what two model families on one task actually license. 32 00:14:42,574 --> 00:15:15,000 [Hal Turing] Which is a good place to land this. The core finding held up under real scrutiny — a base model's existing reasoning behaviors, not its scale and not answer correctness, gate whether reinforcement learning turns into self-improvement or a plateau. That's genuinely useful, even bounded to Countdown and two model families for now. Thanks for sticking with us through this one, and thank you both for spending the time with Ada and me on this three-part arc. I'm Hal Turing, alongside Dr. Ada Shannon — until next time, take care of yourselves.