A visual companion to the episode about why serving cost is often dominated by memory behavior during autoregressive decoding, not by headline parameter count. The page focuses on KV-cache traffic, attention design, and Attention-FFN Disaggregation.
arXiv 2507.19427One output token touches multiple subsystems. The key visual distinction is that attention spends its life dragging prior-token state across memory, while FFN/MoE spends more of its time on dense compute.
A model can activate fewer parameters and still decode expensively if the resulting token path has poor arithmetic intensity or high KV traffic.
As sequence length grows, attention keeps rereading longer key/value histories. That makes memory bandwidth pressure compound faster than simple parameter counts suggest.
The lower heat strip shows which stages are bandwidth-hot versus compute-hot. In long context, attention shifts deeper into the red zone.
Grouped-query variants reduce cache bytes by sharing KV state. MFA’s pitch is not merely “shrink more” but “preserve more useful attention rank per byte.”
The matrix colors encode which query heads share the same KV pool. More sharing cuts cache footprint but can reduce expressive independence if done too aggressively.
Some compression schemes save bytes but add awkward per-token algebra. The good regime is better quality-per-byte without producing hardware-unfriendly kernels.
Switching modes updates both the head-sharing heatmap and the decoded-cost bars so the architecture-to-systems link stays visible.
The systems argument is that attention and FFN are not the same workload. If their bottlenecks differ, serving them on the same machine profile leaves efficiency on the table.
Disaggregation shifts the problem into routing, transport, scheduling, batching, and heterogeneous hardware management.
Attention nodes can be sized for cache traffic and context scaling; FFN/MoE nodes can chase batch utilization and dense throughput.
Use the step buttons to move from one big block, to phase awareness, to a full attention/FFN service split with transport links.
The episode separates one strong measured Hopper result from broader analytic frontier claims. This tab keeps those evidence types distinct.
Shows throughput at a 50 ms TPOT SLA with the transcript’s Step-3 and DeepSeek-V3 numbers plus illustrative comparators.
Shows a modeled frontier where savings widen as context grows. The exact points are illustrative but reflect the paper’s scaling story from the episode.
A lower-cost curve is compelling only if capability stays matched. The missing isolated quality-retention evidence remains the main caveat.