Interactive Visualization AI Infrastructure Podcast Companion arXiv:2601.11577 ↗

Computation‑Bandwidth‑Memory Trade‑offs for AI Infrastructure

A visual map of the paper’s “AI Trinity”: compute, interconnect bandwidth, and memory as one coupled design space. Explore where modern AI workloads move pressure around the triangle: all‑reduce traffic, memory disaggregation, checkpointing, FlashAttention‑style IO optimization, and KV‑cache serving.

Primary paper
2601.11577
Main framing
AI Trinity
Mode
Joint resource trade‑offs

Overview: the coupled triangle

Each edge is a substitution path: spend one resource to relieve pressure on another. Hover the workload markers.

flow diagram

Workload pressure snapshot

heatmap

Visual reading

Hot cells mean the workload is constrained by that resource. Training skews toward bandwidth and memory; long‑context inference shifts hard toward memory via KV cache.

Connected references

Roofline, FlashAttention, ZeRO, checkpointing, vDNN, split computing, and KV‑cache compression are all examples of moving along this triangle rather than solving a single bottleneck in isolation.

Trade‑off explorer

Switch the exchange direction. Bubble size = relative deployment relevance; arrow color = resource being spent.

Stage‑by‑stage resource flow

stepper

Mock system outcomes

Illustrative comparisons grounded in the episode’s examples. Toggle the scenario to see how the bottleneck shifts.

Roofline‑style operating points

line chart

Matrix views of what gets moved or kept

Heatmaps show representative tensor traffic and retained state: dense gradient sync, tiled attention IO, and compressed KV caches.

Cache residency curve

area chart

Why this matters

Not all memory is equal. HBM, CPU RAM, SSD page cache, disaggregated pools, and network buffers all carry different latency and bandwidth costs.

Inference reality

Long contexts increasingly look memory‑bound before they look compute‑bound. That is why KV‑cache eviction, compression, retrieval heads, and streaming caches matter so much.

References

  1. Fan, Weng, Li (2025). Computation‑Bandwidth‑Memory Trade‑offs: A Unified Paradigm for AI Infrastructure.
  2. Rajbhandari et al. (2020). ZeRO: Memory Optimizations Toward Training Trillion Parameter Models.
  3. Chen et al. (2016). Training Deep Nets with Sublinear Memory Cost.
  4. Rhu et al. (2016). vDNN: Virtualized Deep Neural Networks for Scalable, Memory‑Efficient Neural Network Design.
  5. Shoeybi et al. (2019). Megatron‑LM: Training Multi‑Billion Parameter Language Models Using Model Parallelism.
  6. Dao et al. (2022). FlashAttention: Fast and Memory‑Efficient Exact Attention with IO‑Awareness.
  7. Lin et al. (2018). Deep Gradient Compression.
  8. Kang et al. (2017/2020 survey line). Neurosurgeon / Split Computing for Mobile Deep Inference.
  9. Recent inference memory papers: RazorAttention, StreamKV, Not All Heads Matter; recent in‑network aggregation systems: GRID, InArt, and related work.
Additional arXiv IDs extracted from transcript text pattern: 2601.11577