Interactive Visualization • Decode-Time Economics

Affordable Large-Scale Decoding Through Model-System Co-Design

A visual companion to the episode about why serving cost is often dominated by memory behavior during autoregressive decoding, not by headline parameter count. The page focuses on KV-cache traffic, attention design, and Attention-FFN Disaggregation.

arXiv 2507.19427
Claim cost follows bytes moved per token more than total params
Measured 4,039 tok/s/GPU on Hopper at 50 ms TPOT SLA
Compare DeepSeek-V3 reported at 2,324 tok/s/GPU in same setup
Transcript IDs 2507.19427
Visuals below use transcript-grounded claims plus illustrative mock data where the paper discusses tradeoffs without publishing every exact intermediate value.
Model Scale
321B total / 38B active
The episode’s warning: activated parameters are still an incomplete cost proxy.
Decode Bottleneck
KV cache reads
Attention tends to be bandwidth-bound while FFN/MoE behaves more like compute-bound matrix math.
Co-Design Move
MFA + AFD
Architectural cache savings matter more when the serving topology is built to exploit them.
Skeptical Read
Package beats isolated proof
The strongest evidence is for the full stack, not for perfectly isolated component gains.

Decode Anatomy

One output token touches multiple subsystems. The key visual distinction is that attention spends its life dragging prior-token state across memory, while FFN/MoE spends more of its time on dense compute.

Why params mislead

A model can activate fewer parameters and still decode expensively if the resulting token path has poor arithmetic intensity or high KV traffic.

Traffic grows with context

As sequence length grows, attention keeps rereading longer key/value histories. That makes memory bandwidth pressure compound faster than simple parameter counts suggest.

Hover the bands

The lower heat strip shows which stages are bandwidth-hot versus compute-hot. In long context, attention shifts deeper into the red zone.

cool / low pressure moderate hot / dominant

Attention Tradeoffs

Grouped-query variants reduce cache bytes by sharing KV state. MFA’s pitch is not merely “shrink more” but “preserve more useful attention rank per byte.”

Head sharing

The matrix colors encode which query heads share the same KV pool. More sharing cuts cache footprint but can reduce expressive independence if done too aggressively.

Arithmetic intensity matters

Some compression schemes save bytes but add awkward per-token algebra. The good regime is better quality-per-byte without producing hardware-unfriendly kernels.

Toggle modes

Switching modes updates both the head-sharing heatmap and the decoded-cost bars so the architecture-to-systems link stays visible.

KV groups
Bytes/token
Expressive rank
Intensity
Mock quality

Attention-FFN Disaggregation

The systems argument is that attention and FFN are not the same workload. If their bottlenecks differ, serving them on the same machine profile leaves efficiency on the table.

Not free plumbing

Disaggregation shifts the problem into routing, transport, scheduling, batching, and heterogeneous hardware management.

Why the split helps

Attention nodes can be sized for cache traffic and context scaling; FFN/MoE nodes can chase batch utilization and dense throughput.

Progressive diagram

Use the step buttons to move from one big block, to phase awareness, to a full attention/FFN service split with transport links.

Cost Frontier

The episode separates one strong measured Hopper result from broader analytic frontier claims. This tab keeps those evidence types distinct.

Measured panel

Shows throughput at a 50 ms TPOT SLA with the transcript’s Step-3 and DeepSeek-V3 numbers plus illustrative comparators.

Analytic panel

Shows a modeled frontier where savings widen as context grows. The exact points are illustrative but reflect the paper’s scaling story from the episode.

Skeptical lens

A lower-cost curve is compelling only if capability stays matched. The missing isolated quality-retention evidence remains the main caveat.

References

Step-3 is Large yet Affordable: Model-system Co-design for Cost-effective Decoding StepFun et al., 2025. Core paper behind the episode’s decode-time cost argument. arXiv 2507.19427
Fast Transformer Decoding: One Write-Head is All You Need Noam Shazeer, 2019. The multi-query attention lineage. Scholar link
GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints Ainslie et al., 2023. Grouped-query tradeoffs between flexibility and cache savings. Scholar link
Multi-matrix Factorization Attention Hu et al., 2024. MFA’s architectural argument for higher effective rank per byte. Scholar link
Splitwise Patel et al., 2023. Phase splitting for generative inference. Scholar link
P/D-Serve and related disaggregation work Serving systems that split bottlenecks across hardware roles. P/D-Serve