An 8B student model beats every teacher used to train it on math benchmarks — by turning two weaker models into a proxy teacher via logit subtraction, then distilling on the student's own rollouts. This page visualizes the mechanism, the three contrast recipes, the results, and where the paper's claims run narrower than the abstract suggests.
Where the training signal comes from
Classic distillation trains offline on fixed text. RLVR trains on-policy but with one sparse reward per rollout. On-policy distillation (OPD) combines on-policy rollouts with dense per-token supervision — toggle to compare feedback density.
The logit-space proxy teacher
No model here is stronger than the 8B student. A "teacher" is manufactured by taking the difference between a positive and a negative weak model, scaling it by α, and injecting it onto the student's own base logits.
Three ways to pick the positive/negative pair
Each recipe isolates a different capability direction. Click a recipe to see which pair of models it contrasts.
What each contrast teaches
Episode-classifier taxonomy (Schoenfeld) over which token types each contrast direction loads onto. Hover a cell for the value.
Amplification coefficient α sweet spot
Too small and the injected direction barely registers; too large and the proxy teacher drifts from the student's own distribution.
Plain OPD degrades general ability below the untouched base model. W2S-OPD, trained on the same math-only data, improves it instead.
"Abundant weak models" — the fine print
Logit subtraction requires identical vocab indices. Every experiment is Qwen3-on-Qwen3.
Structurally excludes closed-API models (no raw logits) and cross-family open models (no vocab alignment). "Abundant weak models" quietly means "same-family, local, logit-accessible."
The headline claim vs Table 2
"Student surpasses the teacher" holds for math. For code, the student narrows the gap but does not cross it.
Abstract generalizes a math-specific result into a domain-general claim — the sentence most likely to be cited without checking the table.
Tested scale vs the paper's motivating scenario
The motivation is the frontier — no larger teacher exists at all. Every experiment uses a mid-size 8B student with a 4B→8B gap.