This episode explores ForkKV, a systems paper on serving multiple LoRA-based agents from one base language model without duplicating massive KV caches for shared context. It explains why ordinary prefix caching breaks once different LoRA adapters change the activations, then walks through the paper’s core idea: split cache state into a large shared base and a small adapter-specific residual, using an operating-system-style copy-on-write model for agent branches. The discussion connects that design to prior work on LoRA, prefix caching, PagedAttention, and disaggregated memory, making the argument that the real win is practical GPU memory efficiency for coding assistants and tool-using agent workflows. Listeners would find it interesting because it frames transformer serving as a memory-management problem and shows how borrowing ideas from Unix process forking could make multi-agent LLM systems far more scalable.
Sources:
1. ForkKV: Scaling Multi-LoRA Agent Serving via Copy-on-Write Disaggregated KV Cache — Shao Wang, Rui Ren, Lin Gui, 2026
http://arxiv.org/abs/2604.063702. The UNIX Time-Sharing System — Dennis M. Ritchie and Ken Thompson, 1974
https://scholar.google.com/scholar?q=The+UNIX+Time-Sharing+System3. Efficient Memory Management for Large Language Model Serving with PagedAttention — Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica, 2023
https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention4. vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention — Ramya Prabhu, Ajay Nayak, Jayashree Mohan, Ramachandran Ramjee, and Ashish Panwar, 2024
https://scholar.google.com/scholar?q=vAttention%3A+Dynamic+Memory+Management+for+Serving+LLMs+without+PagedAttention5. ForkKV: Scaling Multi-LoRA Agent Serving via Copy-on-Write Disaggregated KV Cache — Shao Wang, Rui Ren, and Lin Gui, 2026
https://scholar.google.com/scholar?q=ForkKV%3A+Scaling+Multi-LoRA+Agent+Serving+via+Copy-on-Write+Disaggregated+KV+Cache6. Efficiently Programming Large Language Models using SGLang — Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Jeff Huang, Chuyue Sun, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonzalez, Clark Barrett, and Ying Sheng, 2023
https://scholar.google.com/scholar?q=Efficiently+Programming+Large+Language+Models+using+SGLang7. DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving — Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, and Hao Zhang, 2024
https://scholar.google.com/scholar?q=DistServe%3A+Disaggregating+Prefill+and+Decoding+for+Goodput-optimized+Large+Language+Model+Serving8. MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool — Cunchen Hu, Heyang Huang, Junhao Hu, Jiang Xu, Xusheng Chen, Tao Xie, Chenxi Wang, Sa Wang, Yungang Bao, Ninghui Sun, and Yizhou Shan, 2024
https://scholar.google.com/scholar?q=MemServe%3A+Context+Caching+for+Disaggregated+LLM+Serving+with+Elastic+Memory+Pool9. Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving — Ruoyu Qin, Zheming Li, Weiran He, Mingxing Zhang, Yongwei Wu, Weimin Zheng, and Xinran Xu, 2024
https://scholar.google.com/scholar?q=Mooncake%3A+A+KVCache-centric+Disaggregated+Architecture+for+LLM+Serving10. LRAgent: efficient kv cache sharing for multi-lora llm agents — H. Jeon, H. Ha, and J. Kim, 2026
https://scholar.google.com/scholar?q=LRAgent%3A+efficient+kv+cache+sharing+for+multi-lora+llm+agents11. Punica: multi-tenant lora serving — L. Chen, Z. Ye, Y. Wu, D. Zhuo, L. Ceze, and A. Krishnamurthy, 2024
https://scholar.google.com/scholar?q=Punica%3A+multi-tenant+lora+serving12. S-LoRA: serving thousands of concurrent lora adapters — Y. Sheng, S. Cao, D. Li, C. Hooper, N. Lee, S. Yang, C. Chou, B. Zhu, L. Zheng, K. Keutzer, J. E. Gonzalez, and I. Stoica, 2023
https://scholar.google.com/scholar?q=S-LoRA%3A+serving+thousands+of+concurrent+lora+adapters13. DLoRA: dynamically orchestrating requests and adapters for LoRA LLM serving — B. Wu, R. Zhu, Z. Zhang, P. Sun, X. Liu, and X. Jin, 2024
https://scholar.google.com/scholar?q=DLoRA%3A+dynamically+orchestrating+requests+and+adapters+for+LoRA+LLM+serving14. Tokencake: A KV-Cache-centric Serving Framework for LLM-based Multi-Agent Applications — Z. Bian, F. Wu, T. Ma, and Y. Zhuo, 2025
https://scholar.google.com/scholar?q=Tokencake%3A+A+KV-Cache-centric+Serving+Framework+for+LLM-based+Multi-Agent+Applications15. KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows — Z. Pan, A. Patel, Z. Hu, Y. Shen, Y. Guan, W. Li, L. Qin, Y. Wang, and Y. Ding, 2025
https://scholar.google.com/scholar?q=KVFlow%3A+Efficient+Prefix+Caching+for+Accelerating+LLM-Based+Multi-Agent+Workflows16. MobiLoRA: Accelerating LoRA-based LLM Inference on Mobile Devices via Context-aware KV Cache Optimization — Borui Li et al., 2025
https://scholar.google.com/scholar?q=MobiLoRA%3A+Accelerating+LoRA-based+LLM+Inference+on+Mobile+Devices+via+Context-aware+KV+Cache+Optimization17. Efficient Multi-Adapter LLM Serving via Cross-Model KV-Cache Reuse with Activated LoRA — Allison Li, Kristjan Greenewald, Thomas Parnell, Navid Azizan, 2025
https://scholar.google.com/scholar?q=Efficient+Multi-Adapter+LLM+Serving+via+Cross-Model+KV-Cache+Reuse+with+Activated+LoRA18. KVLink: Accelerating Large Language Models via Efficient KV Cache Reuse — Jingbo Yang et al., 2025
https://scholar.google.com/scholar?q=KVLink%3A+Accelerating+Large+Language+Models+via+Efficient+KV+Cache+Reuse19. AdaFuse: Accelerating Dynamic Adapter Inference via Token-Level Pre-Gating and Fused Kernel Optimization — Qiyang Li et al., 2026
https://scholar.google.com/scholar?q=AdaFuse%3A+Accelerating+Dynamic+Adapter+Inference+via+Token-Level+Pre-Gating+and+Fused+Kernel+Optimization20. ChunkKV: Semantic-Preserving KV Cache Compression for Efficient Long-Context LLM Inference — Xiang Liu et al., 2025
https://scholar.google.com/scholar?q=ChunkKV%3A+Semantic-Preserving+KV+Cache+Compression+for+Efficient+Long-Context+LLM+Inference21. SCBench: A KV Cache-Centric Analysis of Long-Context Methods — Yucheng Li et al., 2025
https://scholar.google.com/scholar?q=SCBench%3A+A+KV+Cache-Centric+Analysis+of+Long-Context+Methods22. ELORA: Efficient LoRA and KV Cache Management for Multi-LoRA LLM Serving — Jiuchen Shi et al., 2026
https://scholar.google.com/scholar?q=ELORA%3A+Efficient+LoRA+and+KV+Cache+Management+for+Multi-LoRA+LLM+Serving23. AI Post Transformers: Efficient KV Cache Sharing for Multi-LoRA Agents — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-22-efficient-kv-cache-sharing-for-multi-lor-afda05.mp324. AI Post Transformers: TokenDance for Multi-Agent KV Cache Sharing — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-22-tokendance-for-multi-agent-kv-cache-shar-aa9b99.mp325. AI Post Transformers: ContiguousKV for Faster LLM Prefill KV Reuse — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-20-contiguouskv-for-faster-llm-prefill-kv-r-59f545.mp326. AI Post Transformers: Speculative Decoding in Real vLLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-speculative-decoding-in-real-vllm-servin-6f4e2b.mp327. AI Post Transformers: Why LLM Serving Needs Mathematical Optimization — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-why-llm-serving-needs-mathematical-optim-647fc6.mp3Interactive Visualization: ForkKV for Multi-LoRA Agent Serving