A visual map of the paper’s “AI Trinity”: compute, interconnect bandwidth, and memory as one coupled design space. Explore where modern AI workloads move pressure around the triangle: all‑reduce traffic, memory disaggregation, checkpointing, FlashAttention‑style IO optimization, and KV‑cache serving.
Each edge is a substitution path: spend one resource to relieve pressure on another. Hover the workload markers.
Hot cells mean the workload is constrained by that resource. Training skews toward bandwidth and memory; long‑context inference shifts hard toward memory via KV cache.
Roofline, FlashAttention, ZeRO, checkpointing, vDNN, split computing, and KV‑cache compression are all examples of moving along this triangle rather than solving a single bottleneck in isolation.
Switch the exchange direction. Bubble size = relative deployment relevance; arrow color = resource being spent.
Illustrative comparisons grounded in the episode’s examples. Toggle the scenario to see how the bottleneck shifts.
Heatmaps show representative tensor traffic and retained state: dense gradient sync, tiled attention IO, and compressed KV caches.
Not all memory is equal. HBM, CPU RAM, SSD page cache, disaggregated pools, and network buffers all carry different latency and bandwidth costs.
Long contexts increasingly look memory‑bound before they look compute‑bound. That is why KV‑cache eviction, compression, retrieval heads, and streaming caches matter so much.