This episode explores a USENIX FAST'26 paper that addresses the infrastructure bottleneck of loading massive language model weights from storage into accelerator memory during inference deployments. The authors present a programmable page cache framework that achieves 2-4× faster cold start times by exploiting predictable sequential access patterns and XPU affinity, while maintaining full compatibility with existing model formats, inference frameworks, and hardware—unlike prior approaches such as ServerlessLLM and BlitzScale that require custom formats or specific interconnects. The discussion examines why the standard kernel page cache underutilizes modern SSD bandwidth through conservative prefetching and inappropriate LRU eviction policies designed for general workloads, and how a userspace-programmable caching layer can optimize for the specific characteristics of model loading without intrusive kernel modifications. Listeners interested in production ML infrastructure, storage systems optimization, or the operational challenges of deploying large models at scale will find concrete insights into how I/O dominates cold start latency and emerging solutions that bridge the three-orders-of-magnitude gap between SSD and GPU memory bandwidth.

Sources:
1. Accelerating LLM Cold Starts with Programmable Page Cache
https://www.usenix.org/system/files/fast26-liu-yubo.pdf
2. Orca: A Distributed Serving System for Transformer-Based Generative Models — Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim, Byung-Gon Chun, 2022
https://scholar.google.com/scholar?q=Orca%3A+A+Distributed+Serving+System+for+Transformer-Based+Generative+Models
3. AlpaServe: Statistical Multiplexing with Model Parallelism for Deep Learning Serving — Zhuohan Li, Lianmin Zheng, Yinmin Zhong, Vincent Liu, Ying Sheng, Xin Jin, Yanping Huang, Zhifeng Chen, Hao Zhang, Joseph E. Gonzalez, Ion Stoica, 2023
https://scholar.google.com/scholar?q=AlpaServe%3A+Statistical+Multiplexing+with+Model+Parallelism+for+Deep+Learning+Serving
4. FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU — Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, Daniel Y. Fu, Zhiqiang Xie, Beidi Chen, Clark Barrett, Joseph E. Gonzalez, Percy Liang, Christopher Ré, Ion Stoica, Ce Zhang, 2023
https://scholar.google.com/scholar?q=FlexGen%3A+High-Throughput+Generative+Inference+of+Large+Language+Models+with+a+Single+GPU
5. ServerlessLLM: Locality-Enhanced Serverless Inference for Large Language Models — Yao Fu, Leyang Xue, Yeqi Huang, Andrei-Octavian Brabete, Dmitrii Ustiugov, Yuvraj Patel, Luo Mai, 2024
https://scholar.google.com/scholar?q=ServerlessLLM%3A+Locality-Enhanced+Serverless+Inference+for+Large+Language+Models
6. DeepSpeed-Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale — Aminabadi et al., 2022
https://scholar.google.com/scholar?q=DeepSpeed-Inference%3A+Enabling+Efficient+Inference+of+Transformer+Models+at+Unprecedented+Scale
7. ZeRO-Offload: Democratizing Billion-Scale Model Training — Ren et al., 2021
https://scholar.google.com/scholar?q=ZeRO-Offload%3A+Democratizing+Billion-Scale+Model+Training
8. Safetensors: Simple, safe way to store and distribute tensors — HuggingFace, 2022
https://scholar.google.com/scholar?q=Safetensors%3A+Simple%2C+safe+way+to+store+and+distribute+tensors
9. AI Post Transformers: LLM Cold Starts: Fixing Linux Page Cache for Model Loading — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-17-llm-cold-starts-fixing-linux-page-cache-a9f9a9.mp3
10. AI Post Transformers: SolidAttention: Efficient SSD-based KV Cache Offloading for Long-Context LLMs — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-17-solidattention-efficient-ssd-based-kv-ca-336b79.mp3
11. AI Post Transformers: Bidaw: Computation-Storage Aware KV Caching for LLMs — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-16-bidaw-computation-storage-aware-kv-cachi-9d89fb.mp3
12. AI Post Transformers: xLLM: Co-Locating Online and Offline LLM Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-16-xllm-co-locating-online-and-offline-llm-10bb81.mp3
Interactive Visualization: Accelerating LLM Cold Starts with Programmable Page Cache

This episode explores SolidAttention, a system that enables large language models to run on memory-constrained consumer PCs by offloading the KV cache to SSD storage. The paper addresses a fundamental mismatch: sparse attention patterns create random I/O access that kills SSD performance, while previous offloading solutions like FlexGen only work well with high request concurrency unavailable on local machines. The researchers co-designed sparse attention algorithms with SSD storage management to enable coarse-grained sequential reads instead of fine-grained random access, achieving practical local LLM inference on systems with just 8-16GB of RAM. The discussion covers why KV caches consume four times the memory of model weights, the trade-offs of quantization versus offloading, and why treating attention sparsity and storage optimization as separate problems fails on consumer hardware.

Sources:
1. SolidAttention: Co-Designing Sparse Attention and SSD I/O
https://www.usenix.org/system/files/fast26-zheng.pdf
2. FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU — Sheng et al., 2023
https://scholar.google.com/scholar?q=FlexGen%3A+High-Throughput+Generative+Inference+of+Large+Language+Models+with+a+Single+GPU
3. Efficient Streaming Language Models with Attention Sinks — Xiao et al., 2024
https://scholar.google.com/scholar?q=Efficient+Streaming+Language+Models+with+Attention+Sinks
4. H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models — Zhang et al., 2023
https://scholar.google.com/scholar?q=H2O%3A+Heavy-Hitter+Oracle+for+Efficient+Generative+Inference+of+Large+Language+Models
5. SSD I/O Characteristics: Impacts of Request Size, Access Pattern, and Parallelism — Chen et al., 2016
https://scholar.google.com/scholar?q=SSD+I%2FO+Characteristics%3A+Impacts+of+Request+Size%2C+Access+Pattern%2C+and+Parallelism
6. vLLM: Efficient Memory Management for Large Language Model Serving with PagedAttention — Kwon et al., 2023
https://scholar.google.com/scholar?q=vLLM%3A+Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention
7. AI Post Transformers: SolidAttention: Efficient SSD-based KV Cache Offloading for Long-Context LLMs — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-17-solidattention-efficient-ssd-based-kv-ca-336b79.mp3
8. AI Post Transformers: SolidAttention: Fast SSD-Based Serving on Memory-Constrained PCs — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-16-solidattention-fast-ssd-based-serving-on-1c305d.mp3
9. AI Post Transformers: SolidAttention: Low-Latency SSD-based Serving on Memory-Constrained PCs — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-16-solidattention-low-latency-ssd-based-ser-e22a0d.mp3
10. AI Post Transformers: Bidaw: Bidirectional Awareness for Interactive LLM KV Caching — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-17-bidaw-bidirectional-awareness-for-intera-87c311.mp3
11. AI Post Transformers: Bidaw: Reducing LLM KV Cache Latency with Two-Tier Storage — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-16-bidaw-reducing-llm-kv-cache-latency-with-15dd25.mp3
12. AI Post Transformers: Bidaw: Computation-Storage Aware KV Caching for LLMs — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-16-bidaw-computation-storage-aware-kv-cachi-9d89fb.mp3
13. AI Post Transformers: CacheSlide: Unlocking Cross Position-Aware KV Cache Reuse for LLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-17-cacheslide-unlocking-cross-position-awar-487b2b.mp3
14. AI Post Transformers: Efficient KV Cache Reuse in Dynamic Agent Workflows — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-16-efficient-kv-cache-reuse-in-dynamic-agen-558f19.mp3
15. AI Post Transformers: Accelerating LLM Cold Starts with Programmable Page Cache — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-17-accelerating-llm-cold-starts-with-progra-0912d1.mp3
16. AI Post Transformers: LLM Cold Starts: Fixing Linux Page Cache for Model Loading — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-pacific.com/episodes/2026-03-17-llm-cold-starts-fixing-linux-page-cache-a9f9a9.mp3
Interactive Visualization: SolidAttention: Co-Designing Sparse Attention and SSD I/O

This episode explores a 2026 USENIX FAST paper that proposes replacing hand-written file system code with LLM-generated implementations derived from formal specifications. The authors demonstrate SYSSPEC, a system that uses three types of formal specifications—Hoare logic for functionality, rely-guarantee conditions for modularity, and explicit concurrency protocols—to guide code generation while using validation agents to catch hallucinations and ensure correctness. Analysis of Ext4's commit history reveals that 82.4% of changes are bug fixes and maintenance, suggesting traditional file system development wastes enormous effort on code upkeep rather than innovation. The researchers show that their approach can generate a working file system (SPECFS) and evolve it by patching specifications rather than code, potentially transforming how systems software is developed and maintained.

Sources:
1. Generative File Systems: Replacing Code with Formal Specifications
https://www.usenix.org/system/files/fast26-liu-qingyuan.pdf
2. Yggdrasil: An Optimized System for Training Deep Decision Trees at Scale — Fabrice Popineau, Artem Vysogorets, et al., 2020
https://scholar.google.com/scholar?q=Yggdrasil%3A+An+Optimized+System+for+Training+Deep+Decision+Trees+at+Scale
3. Hyperkernel: Push-Button Verification of an OS Kernel — Luke Nelson, Helgi Sigurbjarnarson, Kaiyuan Zhang, et al., 2017
https://scholar.google.com/scholar?q=Hyperkernel%3A+Push-Button+Verification+of+an+OS+Kernel
4. Program Synthesis from Natural Language Using Recurrent Neural Networks — Xi Victoria Lin, Chenglong Wang, Luke Zettlemoyer, Michael D. Ernst, 2017
https://scholar.google.com/scholar?q=Program+Synthesis+from+Natural+Language+Using+Recurrent+Neural+Networks
5. Crash Hoare Logic — Tej Chajed, Frans Kaashoek, Butler Lampson, Nickolai Zeldovich, 2018
https://scholar.google.com/scholar?q=Crash+Hoare+Logic
6. FSCQ: A Verified File System — Haogang Chen et al., 2015
https://scholar.google.com/scholar?q=FSCQ%3A+A+Verified+File+System
7. Yxv6: An Educational File System with Formal Specifications — Helgi Sigurbjarnarson et al., 2016
https://scholar.google.com/scholar?q=Yxv6%3A+An+Educational+File+System+with+Formal+Specifications
8. Crash Consistency in Database Systems — Goetz Graefe, 2009
https://scholar.google.com/scholar?q=Crash+Consistency+in+Database+Systems
9. Using Crash Hoare Logic for Certifying the FSCQ File System — Haogang Chen et al., 2015
https://scholar.google.com/scholar?q=Using+Crash+Hoare+Logic+for+Certifying+the+FSCQ+File+System
10. Jitk: A Trustworthy In-Kernel Interpreter Infrastructure — Xi Wang et al., 2014
https://scholar.google.com/scholar?q=Jitk%3A+A+Trustworthy+In-Kernel+Interpreter+Infrastructure
11. AI Post Transformers: LLM Agents Reason About Code Without Running It — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-15-llm-agents-reason-about-code-without-run-2a1876.mp3
12. AI Post Transformers: SYSSPEC: LLM-Generated File Systems from Formal Specifications — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-17-sysspec-llm-generated-file-systems-from-02f5a9.mp3
13. AI Post Transformers: Generative File Systems from Formal Specifications with SysSpec — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-16-generative-file-systems-from-formal-spec-ff240b.mp3
14. AI Post Transformers: Sharpen the Spec, Cut the Code: LLM-Generated File Systems — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-16-sharpen-the-spec-cut-the-code-llm-genera-8eb6b1.mp3
Interactive Visualization: Generative File Systems: Replacing Code with Formal Specifications

This episode explores Xerxes, a new open-source simulator designed to model CXL 3.0 features before the hardware exists. The hosts explain how CXL adds cache coherence to PCIe to solve memory access bottlenecks in AI and HPC workloads, then dive into the two major architectural changes in CXL 3.0: Port-Based Routing, which enables arbitrary fabric topologies beyond rigid trees, and Device-Managed Coherence, which lets devices handle coherence protocols peer-to-peer without routing every transaction through the host CPU. The discussion highlights why this simulator matters for designing next-generation rack-scale memory pools and accelerator fabrics, addressing the chicken-and-egg problem of validating designs before physical hardware ships. The hosts question how validation works without reference hardware and preview a deeper look at Xerxes' architecture and methodology.

Sources:
1. Xerxes: CXL 3.0 Simulation for Scalable Memory Systems
https://www.usenix.org/system/files/fast26-an.pdf
2. CXL Memory Disaggregation: Opportunities and Challenges — Guz et al. (Intel), 2023
https://scholar.google.com/scholar?q=CXL+Memory+Disaggregation%3A+Opportunities+and+Challenges
3. Pond: CXL-Based Memory Pooling Systems for Cloud Platforms — Li et al., 2023
https://scholar.google.com/scholar?q=Pond%3A+CXL-Based+Memory+Pooling+Systems+for+Cloud+Platforms
4. TPP: Transparent Page Placement for CXL-Enabled Tiered Memory — Maruf et al., 2023
https://scholar.google.com/scholar?q=TPP%3A+Transparent+Page+Placement+for+CXL-Enabled+Tiered+Memory
5. The CXL Memory Expander: Performance and Cost Analysis — Gouk et al. (SK hynix), 2023
https://scholar.google.com/scholar?q=The+CXL+Memory+Expander%3A+Performance+and+Cost+Analysis
6. Exploring CXL 3.0 Port-Based Routing for Scalable Memory Systems — Pan et al., 2024
https://scholar.google.com/scholar?q=Exploring+CXL+3.0+Port-Based+Routing+for+Scalable+Memory+Systems
7. SMART: Scalable Memory Architecture with Port-Based Routing — Kim et al., 2024
https://scholar.google.com/scholar?q=SMART%3A+Scalable+Memory+Architecture+with+Port-Based+Routing
8. Deadlock-Free Routing for CXL Fabrics — Zhang et al., 2024
https://scholar.google.com/scholar?q=Deadlock-Free+Routing+for+CXL+Fabrics
9. DMC: Distributed Cache Coherence for CXL Memory Systems — Lee et al., 2024
https://scholar.google.com/scholar?q=DMC%3A+Distributed+Cache+Coherence+for+CXL+Memory+Systems
10. Scaling Cache Coherence to Thousands of Devices with CXL DMC — Wang et al., 2024
https://scholar.google.com/scholar?q=Scaling+Cache+Coherence+to+Thousands+of+Devices+with+CXL+DMC
11. Coherence Protocol Verification for CXL Device-Managed Coherence — Chen et al., 2024
https://scholar.google.com/scholar?q=Coherence+Protocol+Verification+for+CXL+Device-Managed+Coherence
12. gem5: A Multiple-ISA Full-System Simulator — Binkert et al., 2011
https://scholar.google.com/scholar?q=gem5%3A+A+Multiple-ISA+Full-System+Simulator
13. The ZSim Simulator: Fast and Accurate Multicore Simulation — Sanchez and Kozyrakis, 2013
https://scholar.google.com/scholar?q=The+ZSim+Simulator%3A+Fast+and+Accurate+Multicore+Simulation
14. Simulating Multi-Core Systems with Shared Memory Coherence — Martin et al. (Wisconsin Multifacet group), 2005
https://scholar.google.com/scholar?q=Simulating+Multi-Core+Systems+with+Shared+Memory+Coherence
15. PARADE: A Cycle-Accurate Full-System Simulation Platform for Accelerator-Rich Architectures — Fuchs et al., 2020
https://scholar.google.com/scholar?q=PARADE%3A+A+Cycle-Accurate+Full-System+Simulation+Platform+for+Accelerator-Rich+Architectures
16. A Primer on Memory Consistency and Cache Coherence — Sorin, Hill, and Wood, 2011
https://scholar.google.com/scholar?q=A+Primer+on+Memory+Consistency+and+Cache+Coherence
17. Coherence and Consistency Models in Shared-Memory Multiprocessors — Adve and Gharachorloo, 1996
https://scholar.google.com/scholar?q=Coherence+and+Consistency+Models+in+Shared-Memory+Multiprocessors
18. DASH: A Scalable Directory-Based Multiprocessor — Lenoski et al. (Stanford DASH project), 1992
https://scholar.google.com/scholar?q=DASH%3A+A+Scalable+Directory-Based+Multiprocessor
19. Directory-Based Cache Coherence in Large-Scale Multiprocessors — Chaiken et al. (Alewife project), 1991
https://scholar.google.com/scholar?q=Directory-Based+Cache+Coherence+in+Large-Scale+Multiprocessors
20. Enabling Rack-Scale Confidential Computing using Heterogeneous Trusted Execution Environment — Jianping Zhu, Hang Yin, Yuekai Jia, Wenhao Wang, Chunhui Li, Jiashuo Liang, Shoumeng Yan, Zhengyu He, Qingkui Liu, Alex X. Liu, 2024
https://scholar.google.com/scholar?q=Enabling+Rack-Scale+Confidential+Computing+using+Heterogeneous+Trusted+Execution+Environment
21. Understanding the Overheads of Hardware Memory Coherence — Lena E. Olson, Joseph Izraelevitz, Mark D. Hill, 2015
https://scholar.google.com/scholar?q=Understanding+the+Overheads+of+Hardware+Memory+Coherence
22. AI Post Transformers: SolidAttention: Co-Designing Sparse Attention and SSD I/O — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-18-solidattention-co-designing-sparse-atten-5a8622.mp3
23. AI Post Transformers: Accelerating LLM Cold Starts with Programmable Page Cache — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-17-accelerating-llm-cold-starts-with-progra-0912d1.mp3
24. AI Post Transformers: xLLM: Co-Locating Online and Offline LLM Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-16-xllm-co-locating-online-and-offline-llm-10bb81.mp3
Interactive Visualization: Xerxes: CXL 3.0 Simulation for Scalable Memory Systems

This episode examines the statistical foundations of Mixture of Block Attention (MoBA), a sparse attention mechanism that divides key-value sequences into blocks and routes queries only to the most relevant ones. The paper derives a signal-to-noise ratio showing that retrieval accuracy depends on the square root of head dimension divided by block size, revealing why smaller blocks improve a router's ability to distinguish relevant from irrelevant content despite increasing computational overhead. The authors introduce FlashMoBA, a hardware-optimized CUDA kernel that makes small block sizes practical on GPUs, and demonstrate how depthwise convolutions on keys can cluster related signals to further boost routing performance. The work provides theoretical grounding for why routing-based sparse attention succeeds at reducing quadratic attention costs to near-linear scaling in long-context language models.

Sources:
1. Optimizing Mixture of Block Attention — Guangxuan Xiao, Junxian Guo, Kasra Mazaheri, Song Han, 2025
http://arxiv.org/abs/2511.11571v2
2. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness — Dao et al., 2022
https://scholar.google.com/scholar?q=FlashAttention%3A+Fast+and+Memory-Efficient+Exact+Attention+with+IO-Awareness
3. Mixture of Experts: A Survey — Various (MoE literature), 2020-2024
https://scholar.google.com/scholar?q=Mixture+of+Experts%3A+A+Survey
4. Sparse Attention Mechanisms (Zaheer et al., Guo et al., Xu et al.) — Cited in paper, 2020-2025
https://scholar.google.com/scholar?q=Sparse+Attention+Mechanisms+%28Zaheer+et+al.%2C+Guo+et+al.%2C+Xu+et+al.%29
5. AI Post Transformers: Optimizing Mixture of Block Attention for Long-Context Transformers — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-17-optimizing-mixture-of-block-attention-fo-ea4612.mp3
6. AI Post Transformers: SolidAttention: Co-Designing Sparse Attention and SSD I/O — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-18-solidattention-co-designing-sparse-atten-5a8622.mp3
7. AI Post Transformers: Bidaw: Bidirectional Awareness for Interactive LLM KV Caching — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-17-bidaw-bidirectional-awareness-for-intera-87c311.mp3
Interactive Visualization: Optimizing Mixture of Block Attention Through Statistical Theory

This episode examines CacheSlide from USENIX FAST26, a system that enables LLMs to reuse cached key-value pairs across shifting prompt positions in agentic workflows. The paper introduces chunked contextual position encoding and priority-based eviction to solve the position mismatch problem that prevents KV cache reuse when prompt segments shift in multi-turn agent conversations.

This episode explores a new system called Bidaw that dramatically improves the performance of long, multi-turn AI chatbot conversations by solving a critical caching problem. The paper reveals that existing approaches waste over 93% of computation redundantly recalculating conversation history, and that naive two-tier storage systems (using both RAM and SSD) increase latency by 3.8x because the GPU scheduler and storage system don't coordinate. Bidaw introduces "bidirectional awareness" where the scheduler prioritizes requests whose data is already in fast memory while background-loading slower SSD data, and the storage system uses conversation flow patterns to predict which cached data to keep hot. Listeners interested in LLM infrastructure, production ML systems, or the practical challenges of deploying interactive AI services will learn how clever coordination between compute and storage layers can unlock major performance gains without requiring more expensive hardware.

Sources:
1. Bidaw: Computation-Storage Aware KV Caching for LLMs
https://www.usenix.org/system/files/fast26-hu-shipeng.pdf
2. FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU — Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, Daniel Y. Fu, Zhiqiang Xie, Beidi Chen, Clark Barrett, Joseph E. Gonzalez, Percy Liang, Christopher Ré, Ion Stoica, Ce Zhang, 2023
https://scholar.google.com/scholar?q=FlexGen%3A+High-Throughput+Generative+Inference+of+Large+Language+Models+with+a+Single+GPU
3. vLLM: Easy, Fast, and Cheap LLM Serving with PagedAttention — Woosuk Kwon, Zhuohan Li, Siddharth Devadas, Ion Stoica, Joseph E. Gonzalez, 2023
https://scholar.google.com/scholar?q=vLLM%3A+Easy%2C+Fast%2C+and+Cheap+LLM+Serving+with+PagedAttention
4. PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU — Yixin Song, Zeyu Mi, Haotong Xie, Haibo Chen, 2023
https://scholar.google.com/scholar?q=PowerInfer%3A+Fast+Large+Language+Model+Serving+with+a+Consumer-grade+GPU
5. DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving — Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, Hao Zhang, 2024
https://scholar.google.com/scholar?q=DistServe%3A+Disaggregating+Prefill+and+Decoding+for+Goodput-optimized+Large+Language+Model+Serving
6. LineFS: Efficient SmartNIC Offload of a Distributed File System with Pipeline Parallelism — Jongyul Kim, Sangwoo Kang, Juhyeong Ryu, Jaehyeong Im, Seongyeop Jeong, Jin-Soo Kim, 2021
https://scholar.google.com/scholar?q=LineFS%3A+Efficient+SmartNIC+Offload+of+a+Distributed+File+System+with+Pipeline+Parallelism
7. Flashield: a Hybrid Key-value Cache that Controls Flash Write Amplification — Yiwen Zhang, Xin Chen, Zhuo Chang, Huanchen Zhang, 2019
https://scholar.google.com/scholar?q=Flashield%3A+a+Hybrid+Key-value+Cache+that+Controls+Flash+Write+Amplification
8. Nexus: A GPU Cluster Engine for Accelerating DNN-Based Video Analysis — Haichen Shen, Lequn Chen, Yuchen Jin, Liangyu Zhao, Bingyu Kong, Matthai Philipose, Arvind Krishnamurthy, Ravi Sundaram, 2019
https://scholar.google.com/scholar?q=Nexus%3A+A+GPU+Cluster+Engine+for+Accelerating+DNN-Based+Video+Analysis
9. Clockwork: A Scheduler for GPU-Accelerated Deep Learning Serving — Arpan Gujarati, Reza Karimi, Safya Alzayat, Wei Hao, Antoine Kaufmann, Ymir Vigfusson, Jonathan Mace, 2020
https://scholar.google.com/scholar?q=Clockwork%3A+A+Scheduler+for+GPU-Accelerated+Deep+Learning+Serving
10. Learning to Cache: Neural Adaptive Caching Policies — Giuseppe DeCandia, Deniz Hastorun, Madan Jampani, Gunavardhan Kakulapati, Avinash Lakshman, Alex Pilchin, Swaminathan Sivasubramanian, Peter Vosshall, Werner Vogels, 2018
https://scholar.google.com/scholar?q=Learning+to+Cache%3A+Neural+Adaptive+Caching+Policies
11. Semantic Caching for Large Language Models — Zheng Gao, Peiyuan Liu, Junwei Cao, Xin Li, 2023
https://scholar.google.com/scholar?q=Semantic+Caching+for+Large+Language+Models
12. Predicting User Behavior in Multi-Turn Dialogue Systems — Yun-Nung Chen, Dilek Hakkani-Tür, Gökhan Tür, Jianfeng Gao, Li Deng, 2016
https://scholar.google.com/scholar?q=Predicting+User+Behavior+in+Multi-Turn+Dialogue+Systems
13. Machine Learning for Storage Systems: A Comprehensive Survey — Jianliang Zhang, Zeke Wang, Tong Zhang, 2023
https://scholar.google.com/scholar?q=Machine+Learning+for+Storage+Systems%3A+A+Comprehensive+Survey
14. PagedAttention: Efficient Memory Management for LLM Serving — Kwon et al. (vLLM), 2023
https://scholar.google.com/scholar?q=PagedAttention%3A+Efficient+Memory+Management+for+LLM+Serving
15. Adaptive Replacement Cache (ARC) — Megiddo and Modha, 2003
https://scholar.google.com/scholar?q=Adaptive+Replacement+Cache+%28ARC%29
16. Learned Cache Replacement Policies — Vietri et al., 2020
https://scholar.google.com/scholar?q=Learned+Cache+Replacement+Policies
17. LoRA: Low-Rank Adaptation of Large Language Models — Hu et al., 2021
https://scholar.google.com/scholar?q=LoRA%3A+Low-Rank+Adaptation+of+Large+Language+Models
18. No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization — MiKV authors, 2024-2025
https://scholar.google.com/scholar?q=No+Token+Left+Behind%3A+Reliable+KV+Cache+Compression+via+Importance-Aware+Mixed+Precision+Quantization
19. CommVQ: Commutative Vector Quantization for KV Cache Compression — CommVQ authors, 2024-2025
https://scholar.google.com/scholar?q=CommVQ%3A+Commutative+Vector+Quantization+for+KV+Cache+Compression
20. KVLink: Accelerating Large Language Models via Efficient KV Cache Reuse — KVLink authors, 2024-2025
https://scholar.google.com/scholar?q=KVLink%3A+Accelerating+Large+Language+Models+via+Efficient+KV+Cache+Reuse
21. Compute or Load KV Cache? Why Not Both? — Unknown, 2024-2025
https://scholar.google.com/scholar?q=Compute+or+Load+KV+Cache%3F+Why+Not+Both%3F
22. MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention — MInference authors, 2024-2025
https://scholar.google.com/scholar?q=MInference+1.0%3A+Accelerating+Pre-filling+for+Long-Context+LLMs+via+Dynamic+Sparse+Attention
23. KVCache Cache in the Wild: Characterizing and Optimizing KVCache at a Large Cloud Provider — Cloud provider study authors, 2024-2025
https://scholar.google.com/scholar?q=KVCache+Cache+in+the+Wild%3A+Characterizing+and+Optimizing+KVCache+at+a+Large+Cloud+Provider
24. Efficient KV Cache Reuse in Dynamic Agent Workflows
https://podcast.do-not-panic.com/episodes/2026-03-16-efficient-kv-cache-reuse-in-dynamic-agen-558f19.mp3
25. 50x KV Cache Compression in Seconds via Attention Matching
https://podcast.do-not-panic.com/episodes/2026-03-09-50x-kv-cache-compression-in-seconds-via-9402c1.mp3
26. Statistical Routing Theory in CARTRIDGE Block Attention
https://podcast.do-not-panic.com/episodes/2026-03-16-statistical-routing-theory-in-cartridge-2083f4.mp3
27. xLLM: Co-Locating Online and Offline LLM Inference
https://podcast.do-not-panic.com/episodes/2026-03-16-xllm-co-locating-online-and-offline-llm-10bb81.mp3
Interactive Visualization: Bidaw: Computation-Storage Aware KV Caching for LLMs

Hal Turing and Dr. Ada Shannon dig into "DualPath: Breaking the Storage Bandwidth Bottleneck in Agentic LLM Inference," a February 2026 paper from a thirteen-author team spanning Peking University, Tsinghua University, and DeepSeek-AI. The episode opens with a striking observation from production: on disaggregated inference clusters running agentic workloads, prefill machines saturate their storage network interfaces at 100% utilization while the equivalent hardware on decode machines sits nearly idle. The hosts use this asymmetry as a lens into a counterintuitive reality — H100 GPUs throttled to 40% compute utilization not by arithmetic limits, but by a storage NIC. Ada explains the structural reason agentic workloads are uniquely hostile to existing infrastructure: the short-append pattern. Unlike standard multi-turn chat, agentic sessions accumulate dozens to hundreds of turns where each round appends only a small number of tokens — a tool result, a stack trace, a code output — onto a context that may already span tens of thousands of tokens. Because that prior context never changes, its KV-Cache was computed once and stored. DeepSeek's production traces show KV-Cache hit rates of 95% or higher, meaning the dominant cost shifts from GPU computation to storage I/O: loading gigabytes of persistent key-value state from external NVMe-backed distributed storage, layer by layer, into prefill engines via RDMA. Hal presses on that 95% figure specifically, establishing that it is grounded in real production traffic rather than idealized assumptions — a distinction that determines whether storage bandwidth or GPU compute is the correct optimization target. The episode frames DualPath's core insight against this background: the storage NICs on decode engines represent idle bandwidth that could absorb KV-Cache load traffic currently overwhelming prefill-side storage interfaces. By routing that traffic through decode-side hardware and transferring it to prefill engines over RDMA, DualPath breaks the single-path bottleneck without adding new hardware. The hosts connect this to the broader memory wall argument — that as context lengths grow and agentic sessions deepen, the architectural shift toward disaggregated inference is not optional, and the constraints driving system design are increasingly about data movement rather than floating-point throughput. DualPath's reported throughput improvement of up to 1.96x is presented as evidence that exploiting idle hardware asymmetries, rather than scaling compute, is where near-term agentic inference gains will be found.

Sources:
1. DualPath: Breaking the Storage Bandwidth Bottleneck in Agentic LLM Inference — Yongtong Wu, Shaoyuan Chen, Yinmin Zhong, Rilin Huang, Yixuan Tan, Wentao Zhang, Liyue Zhang, Shangyan Zhou, Yuxuan Liu, Shunfeng Zhou, Mingxing Zhang, Xin Jin, Panpan Huang, 2026
http://arxiv.org/abs/2602.21548
2. Efficient Memory Management for Large Language Model Serving with PagedAttention — Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, Ion Stoica, 2023
https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention
3. SGLang: Efficient Execution of Structured Language Model Programs — Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Jeff Huang, Chuyue Sun, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph Gonzalez, Clark Barrett, Ying Sheng, 2024
https://scholar.google.com/scholar?q=SGLang%3A+Efficient+Execution+of+Structured+Language+Model+Programs
4. Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving — Ruoyu Qin, Zheming Li, Weiran He, Mingxing Zhang, Yongwei Wu, Weimin Zheng, Xinran Xu, 2024
https://scholar.google.com/scholar?q=Mooncake%3A+A+KVCache-centric+Disaggregated+Architecture+for+LLM+Serving
5. CacheGen: Fast Context Loading for Language Model Applications — Yuhan Liu, Hanchen Li, Yihua Cheng, Siddhant Ray, Yuyang Huang, Qizheng Zhang, Kuntai Du, Jiayi Yao, Shan Lu, Ganesh Ananthanarayanan, Michael Maire, Henry Hoffmann, Ari Holtzman, Junchen Jiang, 2024
https://scholar.google.com/scholar?q=CacheGen%3A+Fast+Context+Loading+for+Language+Model+Applications
6. DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving — Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, Hao Zhang, 2024
https://scholar.google.com/scholar?q=DistServe%3A+Disaggregating+Prefill+and+Decoding+for+Goodput-optimized+Large+Language+Model+Serving
7. Splitwise: Efficient Generative LLM Inference Using Phase Splitting — Pratyush Patel, Esha Choukse, Chaojie Zhang, Aashaka Shah, Igo Goiri, Saeed Maleki, Ricardo Bianchini, 2024
https://scholar.google.com/scholar?q=Splitwise%3A+Efficient+Generative+LLM+Inference+Using+Phase+Splitting
8. Sarathi-Serve: Chunked Prefills for Efficient LLM Inference — Amey Agrawal, Nitin Kedia, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav Gulavani, Alexey Tumanov, Ramachandran Ramjee, 2024
https://scholar.google.com/scholar?q=Sarathi-Serve%3A+Chunked+Prefills+for+Efficient+LLM+Inference
9. Llumnix: Dynamic Scheduling for Large Language Model Serving — Biao Sun, Ziming Huang, Hanyu Zhao, Wencong Xiao, Xinyi Zhang, Yong Li, Wei Lin, 2024
https://scholar.google.com/scholar?q=Llumnix%3A+Dynamic+Scheduling+for+Large+Language+Model+Serving
10. FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU — Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, Beidi Chen, Percy Liang, Christopher Re, Ion Stoica, Ce Zhang, 2023
https://scholar.google.com/scholar?q=FlexGen%3A+High-Throughput+Generative+Inference+of+Large+Language+Models+with+a+Single+GPU
11. InfiniGen: Efficient Generative Inference of Large Language Models with Dynamic KV Cache Management — Wonbeom Lee, Jungi Lee, Junghwan Seo, Jaewoong Sim, 2024
https://scholar.google.com/scholar?q=InfiniGen%3A+Efficient+Generative+Inference+of+Large+Language+Models+with+Dynamic+KV+Cache+Management
12. Parrot: Efficient Servicing of LLM-based Applications with Semantic Variable — Chaofan Lin, Zhenhua Han, Chengruidong Zhang, Yuqing Yang, Fan Yang, Chen Chen, Lili Qiu, 2024
https://scholar.google.com/scholar?q=Parrot%3A+Efficient+Servicing+of+LLM-based+Applications+with+Semantic+Variable
13. AgentBench: Evaluating LLMs as Agents — Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, Shudan Zhang, Xiang Deng, Aohan Zeng, Zhiyuan Liu, Yuxiao Dong, Jie Tang, 2023
https://scholar.google.com/scholar?q=AgentBench%3A+Evaluating+LLMs+as+Agents
14. MemGPT: Towards LLMs as Operating Systems — Charles Packer, Sarah Wooders, Kevin Lin, Vivian Fang, Shishir G. Patil, Ion Stoica, Joseph Gonzalez, 2023
https://scholar.google.com/scholar?q=MemGPT%3A+Towards+LLMs+as+Operating+Systems
15. Sarathi-Serve: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills — Amey Agrawal, Nikhil Mohan, Sudhanshu Panwar, et al., 2024
https://scholar.google.com/scholar?q=Sarathi-Serve%3A+Efficient+LLM+Inference+by+Piggybacking+Decodes+with+Chunked+Prefills
16. CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowledge Fusion — Jiayi Yao, Hanchen Li, Yuhan Liu, et al., 2024
https://scholar.google.com/scholar?q=CacheBlend%3A+Fast+Large+Language+Model+Serving+for+RAG+with+Cached+Knowledge+Fusion
17. CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving — Liu et al., 2024
https://scholar.google.com/scholar?q=CacheGen%3A+KV+Cache+Compression+and+Streaming+for+Fast+Large+Language+Model+Serving
18. Compute or Load KV Cache? Why Not Both? — Unknown (snippet only), 2024-2025
https://scholar.google.com/scholar?q=Compute+or+Load+KV+Cache%3F+Why+Not+Both%3F
19. Semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage — Unknown (snippet only), 2024-2025
https://scholar.google.com/scholar?q=Semi-PD%3A+Towards+Efficient+LLM+Serving+via+Phase-Wise+Disaggregated+Computation+and+Unified+Storage
20. Prefill-Decode Aggregation or Disaggregation? Unifying Both for Goodput-Optimized LLM Serving (TaiChi) — Unknown (snippet only), 2024-2025
https://scholar.google.com/scholar?q=Prefill-Decode+Aggregation+or+Disaggregation%3F+Unifying+Both+for+Goodput-Optimized+LLM+Serving+%28TaiChi%29
21. Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live — Unknown (snippet only), 2024-2025
https://scholar.google.com/scholar?q=Continuum%3A+Efficient+and+Robust+Multi-Turn+LLM+Agent+Scheduling+with+KV+Cache+Time-to-Live
22. SideQuest: Model-Driven KV Cache Management for Long-Horizon Agentic Reasoning — Unknown (snippet only), 2024-2025
https://scholar.google.com/scholar?q=SideQuest%3A+Model-Driven+KV+Cache+Management+for+Long-Horizon+Agentic+Reasoning
23. RDMA Point-to-Point Communication for LLM Systems — Unknown (snippet only), 2024-2025
https://scholar.google.com/scholar?q=RDMA+Point-to-Point+Communication+for+LLM+Systems
24. AI Post Transformers: FAST26: Bidaw: Enhancing Key-Value Caching for Interactive LLM Serving via Bidirectional — Hal Turing & Dr. Ada Shannon, Wed,
https://podcasters.spotify.com/pod/show/12146088098/episodes/FAST26-Bidaw-Enhancing-Key-Value-Caching-for-Interactive-LLM-Serving-via-Bidirectional-e3fjgkh
25. AI Post Transformers: SYMPHONY: Memory Management for LLM Multi-Turn Inference — Hal Turing & Dr. Ada Shannon, Mon,
https://podcasters.spotify.com/pod/show/12146088098/episodes/SYMPHONY-Memory-Management-for-LLM-Multi-Turn-Inference-e3ap0pf
26. AI Post Transformers: CXL-SpecKV: Bridging the LLM Memory Wall with Speculative FPGA Disaggregation — Hal Turing & Dr. Ada Shannon, Sat,
https://podcasters.spotify.com/pod/show/12146088098/episodes/CXL-SpecKV-Bridging-the-LLM-Memory-Wall-with-Speculative-FPGA-Disaggregation-e3foad0
27. AI Post Transformers: AI and the Memory Wall: Overcoming Bottlenecks — Hal Turing & Dr. Ada Shannon, Fri,
https://podcasters.spotify.com/pod/show/12146088098/episodes/AI-and-the-Memory-Wall-Overcoming-Bottlenecks-e36ki0u
28. AI Post Transformers: Tempo: SLO-Aware LLM Serving Maximizing Service Gain — Hal Turing & Dr. Ada Shannon, Mon,
https://podcasters.spotify.com/pod/show/12146088098/episodes/Tempo-SLO-Aware-LLM-Serving-Maximizing-Service-Gain-e3ap12a
Interactive Visualization: DualPath Breaks Storage Bandwidth Bottleneck in Agentic Inference