← All episodes SolidAttention: Co-Designing Sparse Attention and SSD I/O

SolidAttention: Co-Designing Sparse Attention and SSD I/O

Mar 19, 2026
This episode explores SolidAttention, a system that enables large language models to run on memory-constrained consumer PCs by offloading the KV cache to SSD storage. The paper addresses a fundamental mismatch: sparse attention patterns create random I/O access that kills SSD performance, while previous offloading solutions like FlexGen only work well with high request concurrency unavailable on local machines. The researchers co-designed sparse attention algorithms with SSD storage management to enable coarse-grained sequential reads instead of fine-grained random access, achieving practical local LLM inference on systems with just 8-16GB of RAM. The discussion covers why KV caches consume four times the memory of model weights, the trade-offs of quantization versus offloading, and why treating attention sparsity and storage optimization as separate problems fails on consumer hardware.
Sources:
1. SolidAttention: Co-Designing Sparse Attention and SSD I/O
https://www.usenix.org/system/files/fast26-zheng.pdf
2. FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU — Sheng et al., 2023
https://scholar.google.com/scholar?q=FlexGen%3A+High-Throughput+Generative+Inference+of+Large+Language+Models+with+a+Single+GPU
3. Efficient Streaming Language Models with Attention Sinks — Xiao et al., 2024
https://scholar.google.com/scholar?q=Efficient+Streaming+Language+Models+with+Attention+Sinks
4. H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models — Zhang et al., 2023
https://scholar.google.com/scholar?q=H2O%3A+Heavy-Hitter+Oracle+for+Efficient+Generative+Inference+of+Large+Language+Models
5. SSD I/O Characteristics: Impacts of Request Size, Access Pattern, and Parallelism — Chen et al., 2016
https://scholar.google.com/scholar?q=SSD+I%2FO+Characteristics%3A+Impacts+of+Request+Size%2C+Access+Pattern%2C+and+Parallelism
6. vLLM: Efficient Memory Management for Large Language Model Serving with PagedAttention — Kwon et al., 2023
https://scholar.google.com/scholar?q=vLLM%3A+Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention
7. AI Post Transformers: SolidAttention: Efficient SSD-based KV Cache Offloading for Long-Context LLMs — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-17-solidattention-efficient-ssd-based-kv-ca-336b79.mp3
8. AI Post Transformers: SolidAttention: Fast SSD-Based Serving on Memory-Constrained PCs — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-16-solidattention-fast-ssd-based-serving-on-1c305d.mp3
9. AI Post Transformers: SolidAttention: Low-Latency SSD-based Serving on Memory-Constrained PCs — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-16-solidattention-low-latency-ssd-based-ser-e22a0d.mp3
10. AI Post Transformers: Bidaw: Bidirectional Awareness for Interactive LLM KV Caching — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-17-bidaw-bidirectional-awareness-for-intera-87c311.mp3
11. AI Post Transformers: Bidaw: Reducing LLM KV Cache Latency with Two-Tier Storage — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-16-bidaw-reducing-llm-kv-cache-latency-with-15dd25.mp3
12. AI Post Transformers: Bidaw: Computation-Storage Aware KV Caching for LLMs — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-16-bidaw-computation-storage-aware-kv-cachi-9d89fb.mp3
13. AI Post Transformers: CacheSlide: Unlocking Cross Position-Aware KV Cache Reuse for LLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-17-cacheslide-unlocking-cross-position-awar-487b2b.mp3
14. AI Post Transformers: Efficient KV Cache Reuse in Dynamic Agent Workflows — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-16-efficient-kv-cache-reuse-in-dynamic-agen-558f19.mp3
15. AI Post Transformers: Accelerating LLM Cold Starts with Programmable Page Cache — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-17-accelerating-llm-cold-starts-with-progra-0912d1.mp3
16. AI Post Transformers: LLM Cold Starts: Fixing Linux Page Cache for Model Loading — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-pacific.com/episodes/2026-03-17-llm-cold-starts-fixing-linux-page-cache-a9f9a9.mp3
Interactive Visualization: SolidAttention: Co-Designing Sparse Attention and SSD I/O