AI Post Transformers — Episode Companion

OPSDL: Teaching Long-Context Models to Trust Their Short-Context Selves

arXiv:2604.17535 Baidu Inc. · 7 authors Qwen2.5-Instruct · 7B / 14B / 32B On-Policy Self-Distillation · Reverse KL

OPSDL trains a long-context model to trust itself: a short-context version of the same weights, reading only an evidence-only excerpt, teaches the long-context version via token-level reverse KL — no reward model, no human preference labels. This page visualizes the pipeline, the token-level correction mechanism, the reported RULER gains, and the methodological gaps the hosts flagged (single model family, missing LongPO baselines at scale, no compute accounting).

How a model becomes its own teacher

Same weights, two conditioning views. Step through the pipeline — each stage lights up in the diagram below.

Click a step above to walk through the pipeline.

Token-level advantage: log(πteacher / πstudent)

Positive (warm) = teacher far more confident → student is under-using evidence sitting right there in the short context. Negative (hot/red flag) = student more confident than teacher → it has latched onto irrelevant long-context noise. Click any token to see the underlying probability distributions.

agreement (~0) ▲ teacher pulls student up (evidence under-used) ▼ student over-confident (hallucination flag)
Click a token above to inspect its teacher vs. student distribution.

Do the numbers hold up?

Toggle between the three result views the hosts walked through.

Evidence coverage vs. claims made

Hover a cell: blue = tested and supported, orange/red = claimed with thin or missing evidence at that scale.

Hover a cell for the detail behind the color.

"Across model families" — same person, three ages

Every experiment in the paper is Qwen2.5-Instruct. Nothing crosses architectures.

References

1OPSDL: On-Policy Self-Distillation for Long-Context Language Models — Zhang, Ding, Pan, Yang, Kang, Xiong, Gu2026
2LongPO: Long Context Self-Evolution via Short-to-Long Preference Optimization — Chen, Li, Shieh, Bing2025
3Self-Distilled Reasoner: On-Policy Self-Distillation for LLMs (OPSD) — Zhao, Xie, Liu, Huang, Pang, Chen, Grover2026
4On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes (GKD) — Agarwal, Vieillard, Zhou, Stanczyk, Ramos Garea, Geist, Bachem2024
5RULER: What's the Real Context Size of Your Long-Context Language Models? — Hsieh, Sun, Kriman, Acharya, Rekesh, Jia, Zhang, Ginsburg2024
6SoloPO: Unlocking Long-Context Capabilities via Short-to-Long Preference Optimization — Sun, Liao, Han, Bai, Gao, Fu, Shen, Wan, Yan, Zhang, et al.2025