AI Post Transformers • Visual Companion

Do Transformers Need Three Projections?

This page turns the episode into a memory map. The center of gravity is not abstract elegance but decode-time state: what gets stored, what gets shared, and why tying K/V is the one collapse that still looks credible when the sweep moves from 300M to 1.2B.

Paper Kayyam, Gopal, Lewis • 2026 arXiv 2606.04032 Main Trade about 50% less KV cache for about 3.1% higher perplexity Scale Story full sweep at 300M, only shared K/V path carried to 1.2B
Visual Focus role tying, cache arithmetic, scale survival, evidence gaps Plot Note the exact episode numbers are preserved; other points are illustrative to show the frontier shape

Deployment Snapshot

Shared K/V is the interesting corner because it removes live cache bytes. Tying Q/K mostly changes geometry, not resident state.

50% KV-cache cut for Q-K=V versus standard QKV
87.5% Cache cut when Q-K=V stacks with GQA-4
96.9% Extreme combined reduction for Q-K=V + MQA
8.85B Approximate tokens in the 1.2B continuation noted in the repo

Frontier Sketch

More compression helps only if the quality hit stays bounded. The mini chart below keeps the core comparison visible before you open the tabs.

Transcript ID Scan

Known source paper ID: 2606.04032. Additional transcript matches are extracted below from the transcript text pattern DDDD.DDDDD.

What actually changes when you tie projections?

Use the variant switch to move from ordinary QKV to the three tied forms. The left diagram shows the weight-sharing path; the heatmap shows how score geometry gets safer for shared K/V than for shared Q/K.

Interactive tabs + variant switch + heatmap hover

Attention Role Wiring

Tying K/V removes one cached tensor. Tying Q/K does not shrink the cache, but it does push score geometry toward symmetry.

Score Geometry Heatmap

Lower triangle is active under causal masking. Ghost cells above the diagonal show pre-mask symmetry pressure when Q and K are tied.

blue to red: lower to higher score intensity

KV Cache Math

Move the context slider and switch precision. The bars show total resident cache size; the unit atlas shows what each scheme is storing. Notice that Q=K-V stays almost identical to baseline here because it does not merge K with V.

Interactive context slider + precision toggle + scale toggle

Total Cache Footprint

Scenario: 300M decoder at FP16, context 16K. The ratios come directly from what gets stored, not from a new attention kernel.

Stored Unit Atlas

Baseline is a full K row plus a full V row. Shared K/V drops that to one row. Grouped or multi-query variants cut how many KV groups survive.

The Scaling Frontier Narrows Fast

The title sounds universal, but the deepest language-model evidence does not. The 300M sweep checks all tying schemes; the 1.2B continuation mostly asks whether shared K/V still belongs in the same room as GQA and MQA.

Interactive frontier toggle + hover labels

Memory Reduction vs. Quality Cost

The exact episode numbers anchor the frontier: about 50% less cache for about 3.1% worse perplexity, then 87.5% and 96.9% reductions when shared K/V stacks with head sharing.

Continuation Matrix

Heat cells summarize where each variant stands: sweep coverage, continuation, cache help, symmetry pressure, and deployment relevance.

What the Paper Shows, and What It Does Not

Perplexity and cache arithmetic are the strongest signals. Long-context retrieval, quantized-cache interactions, and edge latency still sit in the open-risk column.

Interactive phase toggle + heatmap hover

Evidence Coverage Heatmap

Columns show direct paper signal, deployment relevance, and unresolved systems risk. Hotter cells mean more pressure or more open terrain, not always “better.”

Where the Bottleneck Moves

Switch between training and batch-1 decode. Shared K/V matters most in the decode picture because that is where live cache residency and memory traffic dominate.

References

Primary paper, direct baselines, and adjacent long-context or cache-efficiency work mentioned around the episode.