Attention Role Wiring
Tying K/V removes one cached tensor. Tying Q/K does not shrink the cache, but it does push score geometry toward symmetry.
This page turns the episode into a memory map. The center of gravity is not abstract elegance but decode-time state: what gets stored, what gets shared, and why tying K/V is the one collapse that still looks credible when the sweep moves from 300M to 1.2B.
Shared K/V is the interesting corner because it removes live cache bytes. Tying Q/K mostly changes geometry, not resident state.
More compression helps only if the quality hit stays bounded. The mini chart below keeps the core comparison visible before you open the tabs.
Known source paper ID: 2606.04032. Additional transcript matches are extracted below from the transcript text pattern DDDD.DDDDD.
Use the variant switch to move from ordinary QKV to the three tied forms. The left diagram shows the weight-sharing path; the heatmap shows how score geometry gets safer for shared K/V than for shared Q/K.
Tying K/V removes one cached tensor. Tying Q/K does not shrink the cache, but it does push score geometry toward symmetry.
Lower triangle is active under causal masking. Ghost cells above the diagonal show pre-mask symmetry pressure when Q and K are tied.
Move the context slider and switch precision. The bars show total resident cache size; the unit atlas shows what each scheme is storing. Notice that Q=K-V stays almost identical to baseline here because it does not merge K with V.
Scenario: 300M decoder at FP16, context 16K. The ratios come directly from what gets stored, not from a new attention kernel.
Baseline is a full K row plus a full V row. Shared K/V drops that to one row. Grouped or multi-query variants cut how many KV groups survive.
The title sounds universal, but the deepest language-model evidence does not. The 300M sweep checks all tying schemes; the 1.2B continuation mostly asks whether shared K/V still belongs in the same room as GQA and MQA.
The exact episode numbers anchor the frontier: about 50% less cache for about 3.1% worse perplexity, then 87.5% and 96.9% reductions when shared K/V stacks with head sharing.
Heat cells summarize where each variant stands: sweep coverage, continuation, cache help, symmetry pressure, and deployment relevance.
Perplexity and cache arithmetic are the strongest signals. Long-context retrieval, quantized-cache interactions, and edge latency still sit in the open-risk column.
Columns show direct paper signal, deployment relevance, and unresolved systems risk. Hotter cells mean more pressure or more open terrain, not always “better.”
Switch between training and batch-1 decode. Shared K/V matters most in the decode picture because that is where live cache residency and memory traffic dominate.
Primary paper, direct baselines, and adjacent long-context or cache-efficiency work mentioned around the episode.