AI Post Transformers•FAST '26 Companion Viz

SolidAttention: Co-Designing Sparse Attention and SSD I/O

A visual-first walkthrough of why sparse attention alone is not enough on consumer PCs: SSDs want large sequential reads, while naive KV offloading produces tiny random fetches. This page turns the paper’s ideas into interactive diagrams, heatmaps, timelines, and tradeoff charts.

Paper: Zheng et al. • USENIX FAST 2026 Target hardware: 8–16GB DRAM / consumer NVMe Core claim: 3.1× faster vs naive SSD offload Paper PDF arXiv Search

References

[1] FAST 2026 SolidAttention paper by Zheng et al. usenix.org PDF
[2] FlexGen, 2023 High-throughput generative inference on a single GPU. Scholar
[3] Attention Sinks, 2024 Efficient streaming language models with long contexts. Scholar
[4] H2O, 2023 Heavy-hitter oracle for efficient generative inference. Scholar
[5] SSD I/O Characteristics, 2016 Request size, access pattern, and parallelism matter. Scholar
[6] vLLM, 2023 PagedAttention for efficient memory management. Scholar
[7–16] Podcast episodes Companion discussions in the AI Post Transformers series. podcast.do-not-panic.com

Visualization values are realistic mock data derived from the episode discussion and cited paper framing; they illustrate behavior, not an official reproduction dataset.