Two knobs — prefix source (whose text the student trains on) and KL direction (which way the divergence is measured) — are mathematically independent. Every production pipeline still only ever uses two of the four combinations. Click a cell to see what each quadrant actually is.
The paper's claim: this is two independent forks, not one. Standard SFT is quietly just the top-forward path — forward-KL distillation on teacher-generated text.
Fix a prefix. The gradient of forward KL collapses to plain cross-entropy. The gradient of reverse KL collapses to a REINFORCE-style policy gradient with a dense, per-token reward. Toggle to see each derivation.
Accuracy advantage of reverse KL over forward KL, in Avg@k points, by training horizon.
Mean per-token predictive entropy across training horizon. Forward KL barely moves; reverse KL heads toward zero.
Qwen3-4B teacher, MATH500 accuracy — before and after a GRPO run bolted onto each warm-start checkpoint. The reverse-KL checkpoint starts higher and ends lower.
A convex combination of forward and reverse KL at each prefix. As the mix shifts toward reverse KL, accuracy gain climbs but entropy stability collapses. The paper lands forward-heavy, not fifty-fifty.
Start short. Extend the horizon only while held-out entropy stays above threshold. Stop extending the moment it drops.
Interactive companion visualization: full paper viz