Systematic Characterization of LLM Inference on GPUs

Authors: Haonan Wang et al. arXiv:2512.01644

LLM Inference Pipeline Overview

Visualizing the two-phase transformer inference workflow on GPUs: Prefill (parallel prompt processing) and Decode (autoregressive token generation).

References