OPSDL trains a long-context model to trust itself: a short-context version of the same weights, reading only an evidence-only excerpt, teaches the long-context version via token-level reverse KL — no reward model, no human preference labels. This page visualizes the pipeline, the token-level correction mechanism, the reported RULER gains, and the methodological gaps the hosts flagged (single model family, missing LongPO baselines at scale, no compute accounting).
Same weights, two conditioning views. Step through the pipeline — each stage lights up in the diagram below.
Positive (warm) = teacher far more confident → student is under-using evidence sitting right there in the short context. Negative (hot/red flag) = student more confident than teacher → it has latched onto irrelevant long-context noise. Click any token to see the underlying probability distributions.
Toggle between the three result views the hosts walked through.
Hover a cell: blue = tested and supported, orange/red = claimed with thin or missing evidence at that scale.
Every experiment in the paper is Qwen2.5-Instruct. Nothing crosses architectures.
| 1 | OPSDL: On-Policy Self-Distillation for Long-Context Language Models — Zhang, Ding, Pan, Yang, Kang, Xiong, Gu | 2026 |
| 2 | LongPO: Long Context Self-Evolution via Short-to-Long Preference Optimization — Chen, Li, Shieh, Bing | 2025 |
| 3 | Self-Distilled Reasoner: On-Policy Self-Distillation for LLMs (OPSD) — Zhao, Xie, Liu, Huang, Pang, Chen, Grover | 2026 |
| 4 | On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes (GKD) — Agarwal, Vieillard, Zhou, Stanczyk, Ramos Garea, Geist, Bachem | 2024 |
| 5 | RULER: What's the Real Context Size of Your Long-Context Language Models? — Hsieh, Sun, Kriman, Acharya, Rekesh, Jia, Zhang, Ginsburg | 2024 |
| 6 | SoloPO: Unlocking Long-Context Capabilities via Short-to-Long Preference Optimization — Sun, Liao, Han, Bai, Gao, Fu, Shen, Wan, Yan, Zhang, et al. | 2025 |