A visual walk through the paper’s memory split: what a transformer can cache in weights, what it must recover from the current prompt, and why true induction behavior appears later than cheap global averages.
The task is designed so the model can either lean on dataset-wide averages or notice a local override inside the prompt. Hover the sequence to see where global memory breaks and temporary memory takes over.
Prediction sources across one sequence
Hover tokens and edgesToggle between stationary and local ruleOrange pulse marks where induction matters
Bigram World
This synthetic world makes the gears visible. The left heatmap is the global transition table; the right heatmap is a single sequence’s local override table. High values are likely continuations.
Transition matrices
lowmediumhigh
How an Induction Head Works
A two-layer circuit can chain operations. First attend backward to find the earlier matching token, then use that clue to retrieve what came next. Step through the circuit rather than reading it as a static block diagram.
Progressive circuit view
Birth During Training
The paper’s main temporal claim: output associations sharpen before precise induction targeting does. Toggle models to compare how one-layer and two-layer systems separate early easy learning from later context-sensitive behavior.