One expensive 24B parent becomes a deployable family of 14B, 8B, and 3B descendants through structured pruning, teacher-student distillation, and late long-context extension to 256k. This page visualizes the recipe as a production pipeline rather than a one-off model launch.
The paper’s contribution is easiest to read as a manufacturing graph: prune, distill, extend context, then use the intermediate child as the next initializer. Click a stage to spotlight where reuse happens.
Each child inherits a teacher that is closer in size and training state. That is the paper’s answer to the classic capacity-gap problem in distillation.
Mock data in this page is schematic. It is designed to show the recipe and tradeoffs described in the episode, not to reproduce a table from the paper.
The transcript keeps returning to one question: how big a teacher jump can a small student absorb? Switch between direct-from-parent and cascade modes to see why intermediate teachers can make the transfer smoother.
Direct 24B→3B distillation is plausible but harsh. The cascade inserts a 14B and 8B bridge so the student is always learning from a teacher that is closer in capacity and already adapted to the family’s data mixture.
The paper argues for competitive descendants, but the episode also stresses the missing systems accounting. Use the toggles to compare quality, efficiency, and long-context practicality as different stories about the same family.
The transcript’s stance: strong engineering signal, weaker proof on end-to-end cost savings versus exact from-scratch controls.
This is not a new model species. It is a familiar decoder-only transformer with grouped-query attention, RoPE, a frozen vision encoder, and a long-context curriculum layered into the family pipeline.
Hover the attention tiles: the page models late long-context extension as broader usable retrieval bands rather than uniform attention everywhere.