AI Post Transformers · Episode Companion

Weak-to-Strong On-Policy Distillation Beats the Teacher

Paper: Fangxu Yu, Weijia Xu, Michael Xu, Tianyi Zhou, Zinan Lin · 2026 Orgs: University of Maryland · Microsoft Research · MBZUAI
arXiv:2607.26246 →

An 8B student model beats every teacher used to train it on math benchmarks — by turning two weaker models into a proxy teacher via logit subtraction, then distilling on the student's own rollouts. This page visualizes the mechanism, the three contrast recipes, the results, and where the paper's claims run narrower than the abstract suggests.

Where the training signal comes from

Classic distillation trains offline on fixed text. RLVR trains on-policy but with one sparse reward per rollout. On-policy distillation (OPD) combines on-policy rollouts with dense per-token supervision — toggle to compare feedback density.

The logit-space proxy teacher

No model here is stronger than the 8B student. A "teacher" is manufactured by taking the difference between a positive and a negative weak model, scaling it by α, and injecting it onto the student's own base logits.

Three ways to pick the positive/negative pair

Each recipe isolates a different capability direction. Click a recipe to see which pair of models it contrasts.

What each contrast teaches

Episode-classifier taxonomy (Schoenfeld) over which token types each contrast direction loads onto. Hover a cell for the value.

Amplification coefficient α sweet spot

Too small and the injected direction barely registers; too large and the proxy teacher drifts from the student's own distribution.

Benchmark comparison

Qwen3-8B student, weak sources: Qwen3-4B, Qwen3-4B-RL, Qwen3-0.6B. Toggle domain.

Out-of-domain effect (trained on math only)

Plain OPD degrades general ability below the untouched base model. W2S-OPD, trained on the same math-only data, improves it instead.

"Abundant weak models" — the fine print

Logit subtraction requires identical vocab indices. Every experiment is Qwen3-on-Qwen3.

Structurally excludes closed-API models (no raw logits) and cross-family open models (no vocab alignment). "Abundant weak models" quietly means "same-family, local, logit-accessible."

The headline claim vs Table 2

"Student surpasses the teacher" holds for math. For code, the student narrows the gap but does not cross it.

Abstract generalizes a math-specific result into a domain-general claim — the sentence most likely to be cited without checking the table.

Tested scale vs the paper's motivating scenario

The motivation is the frontier — no larger teacher exists at all. Every experiment uses a mid-size 8B student with a 4B→8B gap.

References

  1. [1]Weak-to-Strong On-Policy Distillation — Yu, Xu, Xu, Zhou, Lin, 2026. arXiv:2607.26246
  2. [2]Distilling the Knowledge in a Neural Network — Hinton, Vinyals, Dean, 2015.
  3. [3]A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning (DAgger) — Ross, Gordon, Bagnell, 2011.
  4. [4]On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes (GKD) — Agarwal, Vieillard, et al. (Google DeepMind), 2024.
  5. [5]On-Policy Distillation — Thinking Machines Lab, 2025.
  6. [6]Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision — Burns, Izmailov, et al. (OpenAI Superalignment), 2023.
  7. [7]Weak-to-Strong Generalization beyond Accuracy — Ning et al., 2024.
  8. [8]RLHF / scalable oversight lineage — Christiano, Leike, Brown, Martic, Legg, Amodei, 2017 and related work.
  9. [9]Contrastive Decoding: Open-ended Text Generation as Optimization — Li, Holtzman, Fried, Liang, Eisner, Hashimoto, Zettlemoyer, Lewis, 2023.
  10. [10]Distillation Scaling Laws — Busbridge, Ramapuram, Ablin, et al. (Apple), 2024.
  11. [11]Fast Inference from Transformers via Speculative Decoding — Leviathan, Kalman, Matias, 2023.