The Same Recipe, Two Outcomes
Identical PPO reinforcement learning, run on two ~3B base models against the Countdown number-puzzle game. The only difference is what each model's pretraining left behind.
Accuracy Over RL Training Steps
Both models start near baseline. Around step 30, Qwen undergoes a qualitative shift — longer responses, climbing accuracy. Llama never breaks out.
Four Behaviors, Unevenly Distributed
GPT-4o-mini classified thousands of raw reasoning traces for four behaviors, before any RL. Hover a cell for the reading.
Priming Llama: Shape Beats Correctness
Seven priming datasets, generated by Claude-3.5-Sonnet under strict behavior constraints, plus two controls. Toggle to compare correct-answer priming against reasoning with the right shape but wrong final answers.
Pushing the Fix Upstream
Instead of task-specific priming, curate pretraining itself: classify OpenWebMath documents by behavior density, rewrite into a clean format, and continue pretraining Llama before ever touching Countdown.
Does It Close the Gap?
Continued pretraining on the behavior-rich set closes most of the gap with Qwen. The token-and-topic-matched control set, low in behavioral density, does not.
How Much Should You Trust the Judge?
Inter-rater reliability (ICC3) between the GPT-4o-mini classifier and Claude / two human annotators, per behavior. Hover a cell for the read.
- Backward chaining — the behavior that most favors Qwen — is exactly where classifier/human agreement is weakest (ICC as low as 0.40).
- Claude-3.5-Sonnet generated all priming data; Qwen-2.5-32B classified and rewrote the pretraining corpus — a same-family model judging and producing "Qwen-style" reasoning.
- Llama-3.1-70B (20x larger) was only measured for static behavior counts, never RL-trained — so scale-vs-family is unresolved.
- PPO was chosen over GRPO (DeepSeek-R1's algorithm) for stability, with cross-algorithm similarity reported as "anecdotal," no data shown.
- Reward only checks final-answer correctness plus a format bonus — nothing directly rewards showing work.
- Transfer to GPQA: behavior-enriched Llama shows far more verification/subgoal setting, but accuracy is flat at ~12% in both conditions.