[1] arXiv:2605.19775, 2026
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles — Arif, Maurya, Vazhkudai, Nicolae (Micron / Argonne).
arxiv.org/abs/2605.19775
[2] PyTorch Distributed, 2020
Experiences on Accelerating Data Parallel Training — Li, Zhao, Varma, et al. (Meta AI).
Scholar
[3] ZeRO, 2020
Memory Optimizations Toward Training Trillion Parameter Models — Rajbhandari, Rasley, Ruwase, He (Microsoft).
Scholar
[4] vLLM / PagedAttention, 2023
Efficient Memory Management for LLM Serving with PagedAttention — Kwon, Li, Zhuang, et al. (UC Berkeley).
Scholar
[5] Orca, 2022
A Distributed Serving System for Transformer-Based Generative Models — Yu, Jeong, Kim, Kim, Chun (Seoul National Univ. / FriendliAI).
Scholar
[6] DistServe, 2024 — uncited by [1]
Disaggregating Prefill and Decoding for Goodput-optimized LLM Serving — Zhong, Liu, Chen, et al.
Scholar
[7] Splitwise, 2024 — uncited by [1]
Efficient Generative LLM Inference Using Phase Splitting — Patel, Choukse, Zhang, et al. (Microsoft Azure).
Scholar
[8] Llumnix, 2024
Dynamic Scheduling for Large Language Model Serving — Sun, Huang, Zhao, et al.
Scholar
[9] vAttention, 2024
Dynamic Memory Management for Serving LLMs without PagedAttention — Prabhu, Nayak, Mohan, et al.
Scholar
[10] KIVI, 2024
A Tuning-Free Asymmetric 2-bit Quantization for KV Cache — Liu, Yuan, Jin, et al.
Scholar
Chart values are realistic mock data illustrating the episode's discussion and the cited paper's framing, not a reproduction of the paper's raw measurements.