AI Post Transformers • Podcast Companion

Accelerating LLM Cold Starts with Programmable Page Cache

A visualization-first walkthrough of why model loading is dominated by storage I/O, how PPC + MAIO exploit deterministic weight access, and where the compatibility story tightens once framework drift, offloading, and dynamic GPU placement enter the picture.

FAST'26 Paper PDF USENIX FAST 2026 arXiv Search Topic: storage-aware inference Claimed speedup: 2-4× cold starts
79%
reported latency reduction vs stock kernel page cache
36%
higher elastic deployment throughput
~1000×
HBM vs SSD bandwidth gap motivating the design

Cold Start Pipeline

The slow path is not token generation. It is moving checkpoint bytes from NVMe, through host memory, toward accelerator memory while the kernel page cache behaves like the workload is generic and unpredictable.

Interactive flow + access heatmap
cold/unused pages active stream window hot sequential burst readahead mismatch

Programmable Cache Control Plane

PPC inserts a thin routing file system in-kernel and pushes policy logic into userspace. MAIO then replays a profiled I/O template to prefetch aggressively, place pages near the target XPU, and evict with Burn-after-Reading.

Step-through architecture

Results Explorer

Mock data tracks the paper’s reported direction: MAIO wins by turning small, conservative reads into wide, sequential streams and by refusing to keep one-time weight pages resident after startup.

Toggle baselines + environment

Compatibility Surface

The page cache is programmable, but the workload assumptions are not free. Stable I/O templates, static affinity, and one-shot loading are strong assumptions; this matrix shows where they hold and where friction appears.

Hover matrix + risk bands

References

Primary source first, then adjacent systems discussed in the episode. Several entries are linked via arXiv search because exact identifiers were not supplied in the prompt.

2. Orca, 2022 • Scholar • arXiv search
3. AlpaServe, 2023 • Scholar • arXiv search
4. FlexGen, 2023 • Scholar • arXiv search
5. ServerlessLLM, 2024 • Scholar • arXiv search
6. DeepSpeed-Inference, 2022 • Scholar • arXiv search
7. ZeRO-Offload, 2021 • Scholar • arXiv search
8. Safetensors, 2022 • Scholar • arXiv search