Compressed Convolutional Attention in Latent Space
A visual tour of the shift from full-space attention to shared latent-space attention: how MHA evolved into MQA, GQA, MLA, and now CCA/CCGQA, and why pushing the actual attention computation into a smaller workspace changes more than just the KV cache.
Smaller internal attention structures can free budget elsewhere.
↑ systems payoff
Only real if fused kernels preserve the theoretical gain.
Transcript-derived arXiv IDs
Pattern extracted from transcript: DDDD.DDDDD. Only one explicit ID appears in the episode transcript.
From full-head freedom to shared latent bottlenecks
Each rung targets a different pain point. MQA and GQA mostly attack decode-time KV duplication. MLA compresses storage but often expands for computation. CCA compresses the computation workspace itself.
query diversity retained
shared KV structure
compressed latent workspace
Inside the compressed latent workspace
Use the stepper to move from baseline MHA to MLA to CCA. The left heatmap shows score computation shape. The right flow diagram shows where compression happens: cache-only vs compute-path.
Performance explorer
These charts use realistic illustrative values drawn from the episode narrative: CCA-family methods aim to improve prefill throughput, cut backward cost, and keep quality competitive under equal or stricter cache budgets.
Long-context economics: where the bill lands
Hover the grid to inspect how prompt length shifts pressure between prefill compute and decode memory bandwidth. The bottom stream view shows why methods that only shrink KV storage solve a different slice of the problem than latent-space attention.