Interactive episode companion

Compressed Convolutional Attention in Latent Space

A visual tour of the shift from full-space attention to shared latent-space attention: how MHA evolved into MQA, GQA, MLA, and now CCA/CCGQA, and why pushing the actual attention computation into a smaller workspace changes more than just the KV cache.

arXiv 2510.04476
Paper: Figliolia, Alonso, Iyer, Anthony, Millidge
Theme: latent-space attention economics
Existing viz URL
Compression Pressure Map
low pressure
mid pressure
hot bottleneck
What CCA tries to win at once
↓ attention FLOPs
Attention runs in compressed latent width, not only storage.
↓ KV cache
Shared latent representations shrink decode memory traffic.
↓ parameters
Smaller internal attention structures can free budget elsewhere.
↑ systems payoff
Only real if fused kernels preserve the theoretical gain.
Transcript-derived arXiv IDs

Pattern extracted from transcript: DDDD.DDDDD. Only one explicit ID appears in the episode transcript.

From full-head freedom to shared latent bottlenecks

Each rung targets a different pain point. MQA and GQA mostly attack decode-time KV duplication. MLA compresses storage but often expands for computation. CCA compresses the computation workspace itself.

query diversity retained
shared KV structure
compressed latent workspace

Inside the compressed latent workspace

Use the stepper to move from baseline MHA to MLA to CCA. The left heatmap shows score computation shape. The right flow diagram shows where compression happens: cache-only vs compute-path.

Performance explorer

These charts use realistic illustrative values drawn from the episode narrative: CCA-family methods aim to improve prefill throughput, cut backward cost, and keep quality competitive under equal or stricter cache budgets.

Long-context economics: where the bill lands

Hover the grid to inspect how prompt length shifts pressure between prefill compute and decode memory bandwidth. The bottom stream view shows why methods that only shrink KV storage solve a different slice of the problem than latent-space attention.

Primary paper
Compressed Convolutional Attention: Efficient Attention in a Compressed Latent Space
Tomas Figliolia, Nicholas Alonso, Rishi Iyer, Quentin Anthony, Beren Millidge
Attention ladder context
MQA, GQA, MLA / DeepSeek-V2
Compressed-attention relatives
Linformer, Perceiver, RoFormer