AI Post Transformers • Interactive Visual Companion

DRAM-Free In-Flash Computing for LLM Inference

A visualization-first guide to KVNAND and the systems question behind it: when decode-time inference is bottlenecked by memory movement, can compute-enabled 3D NAND hold both weights and the KV cache—or does “DRAM-free” overstate what’s really removed?

Paper
KVNAND (2025)
arXiv
Core Claim
Weights + KV in flash
Skeptical Lens
Capacity ≠ fully flash-only execution

1. Decode Bottleneck Map

Visualizing why the paper exists: decode is often limited less by arithmetic throughput than by repeated movement of weights and KV state.

compute-near-data
weight traffic
KV traffic
hot path / bottleneck
Pressure Point
Decode
Token-by-token generation repeatedly touches large operands; hardware spends time waiting on data movement.
New Problem
KV Cache
Long context makes the attention state itself a systems bottleneck, not just model weights.
Question
“DRAM-free”?
The compelling version is “no external DRAM on the main operand path,” not necessarily “flash is the only memory anywhere.”

2. KV Cache Growth + Attention Variant Heatmap

Hover the matrix. MHA, GQA, and MQA radically change KV storage growth, which changes how impressive a flash-resident KV design looks.

cool → hot memory pressure
Interpretation
MHA > GQA > MQA
Sharing K/V across more query heads lowers cache growth and memory traffic.
Why It Matters
Representativeness
A result on MHA can be a valid stress-test, but it can also exaggerate benefits relative to deployment-realistic GQA/MQA models.
Edge Deployment
Context-sensitive
The more persistent and messy the context, the more the KV path competes with weights as the dominant memory problem.

3. KVNAND Architecture Explorer

Step through the method: flash-resident KV, two hardware variants, head-group parallelism, and page-level layout tuned to flash page granularity.

Discrete
Longer Contexts
Separate flash die groups for weights and KV reduce interference and allow more overlap on KV-heavy workloads.
Compact
Fixed Budget
Shared arrays can preserve stronger weight parallelism when context is smaller and dedicated KV resources would be underutilized.
Co-design
Layout + Schedule
This is not just “store KV in flash”; the mapping, buffering, and execution order are designed together.

4. Reported Gains vs. What They Mean

Charts use paper-inspired mock data anchored on transcript numbers: throughput speedups, energy gains, memory cost reduction, and the gap between “fits” and “feels fast.”

Headline
~2×
Transcript-reported geomean speedups: 1.98× at 128 tokens, 1.94× at 1K, 2.05× at 10K versus DRAM-equipped IFC baselines.
Efficiency
1.17×–1.32×
Energy improvements were cited at 10K and 30K contexts; enough to matter, but smaller than the speedup headline.
Skeptic’s Filter
100K = capacity first
The transcript suggests “supports 100K” is strongest as an out-of-memory avoidance claim, not yet a full end-to-end UX proof.

5. Source Constellation

Compact references behind the visuals: primary paper, neighboring in-flash systems, KV management work, and the memory-movement lineage.

KVNAND — Efficient On-Device Large Language Model Inference Using DRAM-Free In-Flash Computing (2025). arXiv:2512.03608
FlashAttention — kernel-side evidence that memory movement, not just FLOPs, dominates practical attention performance.
PagedAttention / vLLM — KV cache management as a first-class systems problem for serving and long context.
Cambricon-LLM — compute-enabled flash for memory-efficient LLM inference; important predecessor in flash-centric design space.
Lincoln — long-context inference with in-flash computing; adjacent hardware-software co-design for flash-resident operands.
Processing-in-Memory survey — broader framing for “compute where the data lives.”
Computational Storage — storage devices evolving from passive media into execution substrates.
MQA / GQA transformer variants — attention designs that reduce KV growth and complicate how universal a KV-centric hardware fix really is.
PowerInfer / LLM in a Flash / FlexGen — memory-hierarchy engineering as the practical path to local inference on constrained hardware.
Recent KV eviction / compression work — RazorAttention, AhaKV, G-KV, and related systems challenge exact full-KV retention as the default target.
Extracted arXiv IDs from transcript/source text: 2512.03608