Five components, one question
Do loss, accuracy, cost and stability agree? Click a component to jump.
Layer stack (illustrative)
Hover a block. 3 Gated DeltaNet : 1 full attention, n-gram lookup injected at layer 2.
Which number is "the size"?
None alone: accelerator, activated, host.
A ninth of the compute
⅓ activated params × ⅓ tokens. Hover cells.
Headline: 14 pre-training benchmarks vs 397B-A17B
Leads on 8, trails on 6 by at most 2.6 points. Dot positions are schematic; only the counts and the two labelled examples come from the paper.
KV / state memory vs context
Gated DeltaNet edits, it does not append
Same key written 3 times (steps 3–4 repeat key 0). Rows = keys, columns = value dims.
QSA indexer pipeline
Pool keys before RoPE, so rotary phases never average.
Attention cost per layer (log scale)
Context length . Indexer n² becomes the bill; QSA cuts it to n²/4.
Results
Residual stream designs
Hover boxes. Gated Residual replaces pre-norm; the branch-mixing operator is dropped.
Table 5 (560B tokens)
The ratio reverses
Loss gain vs accuracy gain per step (loss ×100 for scale). Static→dynamic: 0.002 loss for 1.98 points.
Hashed n-gram lookup
Hover a token: its bigram and trigram hash into the table (2 heads). Addresses use token IDs only, so host memory can prefetch.
Prefetch overlap
Table 9: vocabulary scaling (20x–200x of 250K)
Loss falls every step. Accuracy jumps once, then flat or noisy.
Table 8: experts removed to fund n-grams
Noise ruler
Spread across Table 7 placements at equal loss vs the 50x→200x change in Table 9. Not a seed study.
Grad-norm spikes per 10k steps
4x optimal LR; one run per arm.
Gate ablation
Single-variable pair; the cleanest stability evidence.
Muon vs AdamW parameter routing
Evidence matrix (this page's reading of the discussion)
Hover cells. Colour = how well each axis is established for each component.