AI Post Transformers / Visual Companion

How Induction Heads Emerge in Transformers

A visual walk through the paper’s memory split: what a transformer can cache in weights, what it must recover from the current prompt, and why true induction behavior appears later than cheap global averages.

arXiv 2306.00802 Synthetic bigram world 2-layer success / 1-layer failure Associative memory probes Viz permalink

Weights vs Context

The task is designed so the model can either lean on dataset-wide averages or notice a local override inside the prompt. Hover the sequence to see where global memory breaks and temporary memory takes over.

Prediction sources across one sequence

Hover tokens and edges Toggle between stationary and local rule Orange pulse marks where induction matters

Bigram World

This synthetic world makes the gears visible. The left heatmap is the global transition table; the right heatmap is a single sequence’s local override table. High values are likely continuations.

Transition matrices

low medium high

How an Induction Head Works

A two-layer circuit can chain operations. First attend backward to find the earlier matching token, then use that clue to retrieve what came next. Step through the circuit rather than reading it as a static block diagram.

Progressive circuit view

Birth During Training

The paper’s main temporal claim: output associations sharpen before precise induction targeting does. Toggle models to compare how one-layer and two-layer systems separate early easy learning from later context-sensitive behavior.

Training curves and component emergence

Solid lines: next-token accuracy Dashed lines: induction retrieval score Bars: associative memory probe alignment

References

Birth of a Transformer: A Memory Viewpoint Bietti et al., 2023 arXiv 2306.00802
A Mathematical Framework for Transformer Circuits Elhage et al., 2021 Scholar
In-context Learning and Induction Heads Olsson et al., 2022 Scholar
What needs to go right for an induction head? Singh et al., 2024 Scholar
Data Distributional Properties Drive Emergent ICL Chan et al., 2022 Scholar
Transformer Feed-Forward Layers Are Key-Value Memories Geva et al., 2021 Scholar