AI Post Transformers · Episode Companion · Part 1 of 3

Cognitive Behaviors Behind Self-Improving Language Model Reasoners

arXiv:2503.01307 Gandhi, Chakravarthy, Singh, Lile, Goodman · 2025 Stanford & SynthLabs · COLM 2025 Qwen-2.5-3B vs Llama-3.2-3B · PPO on Countdown

Same architecture size, same RL algorithm, same hyperparameters — Qwen climbs to ~60% accuracy, Llama plateaus at ~30%. The gap traces to four cognitive behaviors already latent in the base model before any reinforcement learning begins.

The Same Recipe, Two Outcomes

Identical PPO reinforcement learning, run on two ~3B base models against the Countdown number-puzzle game. The only difference is what each model's pretraining left behind.

Accuracy Over RL Training Steps

Both models start near baseline. Around step 30, Qwen undergoes a qualitative shift — longer responses, climbing accuracy. Llama never breaks out.

Four Behaviors, Unevenly Distributed

GPT-4o-mini classified thousands of raw reasoning traces for four behaviors, before any RL. Hover a cell for the reading.

Hover a cell to see the classifier's read on that model × behavior pair.

Priming Llama: Shape Beats Correctness

Seven priming datasets, generated by Claude-3.5-Sonnet under strict behavior constraints, plus two controls. Toggle to compare correct-answer priming against reasoning with the right shape but wrong final answers.

Key result: priming on wrong-answer traces that still verify, backtrack, and set subgoals produces nearly the same RL trajectory as priming on correct traces. The controls (empty chain-of-thought, length-matched placeholder dots) stay at Llama's ~30-35% baseline — so it isn't extra token budget doing the work.

Pushing the Fix Upstream

Instead of task-specific priming, curate pretraining itself: classify OpenWebMath documents by behavior density, rewrite into a clean format, and continue pretraining Llama before ever touching Countdown.

Does It Close the Gap?

Continued pretraining on the behavior-rich set closes most of the gap with Qwen. The token-and-topic-matched control set, low in behavioral density, does not.

How Much Should You Trust the Judge?

Inter-rater reliability (ICC3) between the GPT-4o-mini classifier and Claude / two human annotators, per behavior. Hover a cell for the read.

Hover a cell to see agreement level between the automated classifier and each rater.
Confounds worth flagging:
  • Backward chaining — the behavior that most favors Qwen — is exactly where classifier/human agreement is weakest (ICC as low as 0.40).
  • Claude-3.5-Sonnet generated all priming data; Qwen-2.5-32B classified and rewrote the pretraining corpus — a same-family model judging and producing "Qwen-style" reasoning.
  • Llama-3.1-70B (20x larger) was only measured for static behavior counts, never RL-trained — so scale-vs-family is unresolved.
  • PPO was chosen over GRPO (DeepSeek-R1's algorithm) for stability, with cross-algorithm similarity reported as "anecdotal," no data shown.
  • Reward only checks final-answer correctness plus a format bonus — nothing directly rewards showing work.
  • Transfer to GPQA: behavior-enriched Llama shows far more verification/subgoal setting, but accuracy is flat at ~12% in both conditions.

References