Simple Self-Distillation for Better Code Generation
Apple’s 2026 result is visually simple and conceptually sharp: sample code from the model itself, filter only the emptiest stubs, fine-tune on those outputs, then decode again. This page focuses on the geometry of that loop: where probability mass tightens, where exploration survives, and why code benchmarks make the claim unusually testable.
Primary ClaimSelf-sampled SFT can lift code pass rates
Code exposes the tradeoff directly: pass@1 measures first-shot reliability, while pass@k shows whether useful solution branches still exist after sharpening the model.
Self-Distillation Loop
Step through the training loop. The bright path is what Apple keeps; the dimmed bays are the supervision mechanisms intentionally left out.
Lock vs Fork Token Landscape
The heatmap shows mock next-token uncertainty across positions in code solutions. Toggle between baseline and SSD to see confidence sharpen at lock points while branch mass remains at fork points.
Benchmark Surface
Mocked from the episode’s reported pattern: broad gains across model families, with larger jumps on harder problems and unusually strong preservation of pass@5.
Distillation Family Tree
Click a node to trace how targets evolved: teacher logits, sequence outputs, same-architecture teachers, pseudo-labeling, on-policy rollouts, verifier-backed code alignment, and finally Apple’s stripped-down self-loop.