A visualization-first guide to KVNAND and the systems question behind it: when decode-time inference is bottlenecked by memory movement, can compute-enabled 3D NAND hold both weights and the KV cache—or does “DRAM-free” overstate what’s really removed?
Visualizing why the paper exists: decode is often limited less by arithmetic throughput than by repeated movement of weights and KV state.
Hover the matrix. MHA, GQA, and MQA radically change KV storage growth, which changes how impressive a flash-resident KV design looks.
Step through the method: flash-resident KV, two hardware variants, head-group parallelism, and page-level layout tuned to flash page granularity.
Charts use paper-inspired mock data anchored on transcript numbers: throughput speedups, energy gains, memory cost reduction, and the gap between “fits” and “feels fast.”
Compact references behind the visuals: primary paper, neighboring in-flash systems, KV management work, and the memory-movement lineage.