AI Post Transformers•Episode Companion Viz

Understanding Inference Scaling: Prefill, Decode, and Reasoning Bottlenecks

Reasoning traces past ten thousand tokens flip LLM serving from a compute problem into a memory-capacity problem. This page turns Micron and Argonne's node-scale study — 8B to 671B parameters on an 8×H200 NVLink node — into interactive diagrams for the prefill/decode split, the KV-cache capacity trap, the DP/TP/PP parallelism fight, and dense-vs-MoE cache pressure.

Arif, Maurya, Vazhkudai, Nicolae • Micron + Argonne, 2026 Hardware: 8× H200, NVLink4/NVSwitch, 900GB/s GPU↔GPU Scale tested: 8B → 671B params (dense + MoE) arXiv:2605.19775

References

[1] arXiv:2605.19775, 2026 Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles — Arif, Maurya, Vazhkudai, Nicolae (Micron / Argonne). arxiv.org/abs/2605.19775
[2] PyTorch Distributed, 2020 Experiences on Accelerating Data Parallel Training — Li, Zhao, Varma, et al. (Meta AI). Scholar
[3] ZeRO, 2020 Memory Optimizations Toward Training Trillion Parameter Models — Rajbhandari, Rasley, Ruwase, He (Microsoft). Scholar
[4] vLLM / PagedAttention, 2023 Efficient Memory Management for LLM Serving with PagedAttention — Kwon, Li, Zhuang, et al. (UC Berkeley). Scholar
[5] Orca, 2022 A Distributed Serving System for Transformer-Based Generative Models — Yu, Jeong, Kim, Kim, Chun (Seoul National Univ. / FriendliAI). Scholar
[6] DistServe, 2024 — uncited by [1] Disaggregating Prefill and Decoding for Goodput-optimized LLM Serving — Zhong, Liu, Chen, et al. Scholar
[7] Splitwise, 2024 — uncited by [1] Efficient Generative LLM Inference Using Phase Splitting — Patel, Choukse, Zhang, et al. (Microsoft Azure). Scholar
[8] Llumnix, 2024 Dynamic Scheduling for Large Language Model Serving — Sun, Huang, Zhao, et al. Scholar
[9] vAttention, 2024 Dynamic Memory Management for Serving LLMs without PagedAttention — Prabhu, Nayak, Mohan, et al. Scholar
[10] KIVI, 2024 A Tuning-Free Asymmetric 2-bit Quantization for KV Cache — Liu, Yuan, Jin, et al. Scholar

Chart values are realistic mock data illustrating the episode's discussion and the cited paper's framing, not a reproduction of the paper's raw measurements.