LLM Inference Pipeline Overview
Visualizing the two-phase transformer inference workflow on GPUs: Prefill (parallel prompt processing) and Decode (autoregressive token generation).
Prefill vs Decode: Roofline Model
Roofline model positioning shows Prefill as compute-bound and Decode as memory-bound, with batch size shifting decode's intensity.
Small (1) / Large (64)
Performance & Energy Comparison
Throughput and energy breakdown for prefill and decode phases on A100 and H100 GPUs.
Throughput / Energy
Mixture-of-Experts Memory Impact
Heatmap of memory bandwidth efficiency degradation due to sparse expert activation and non-contiguous weight loading in MoE decode.
GPU Warp Stall Analysis
Breakdown of dominant warp stall reasons during prefill and decode phases, highlighting memory pipe saturation vs instruction dependency.
References
- Wang et al., 2025 - arXiv:2512.01644
- Kwon et al., 2023 - Efficient Memory Management for Large Language Model Serving with PagedAttention
- Yu et al., 2022 - Orca: A Distributed Serving System for Transformer-Based Generative Models
- Li et al., 2022 - AlpaServe: Statistical Multiplexing with Model Parallelism for Deep Learning Serving
- Dettmers et al., 2022 - LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale
- Patel et al., 2024 - Splitwise: Efficient Generative LLM Inference Using Phase Splitting
- Zhong et al., 2024 - DistServe: Disaggregating Prefill and Decoding for Goodput-Optimized Large Language Model Serving
- Agrawal et al., 2024 - Sarathi-Serve: Efficient LLM Inference by Piggybacking Decodes on Chunked Prefills
- Williams, Waterman & Patterson, 2009 - Roofline: An Insightful Visual Performance Model for Multicore Architectures