Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference

DUAL-BLADE — Bodon Jeong, Hongsu Byun, Youngjae Kim, Weikuan Yu, Kyungkeun Lee, Jihoon Yang, Sungyong Park · Sogang University / Florida State University / Samsung Electronics · arXiv, April 29, 2026
arXiv:2604.26557 ↗ Systems / Storage-for-ML Edge LLM Inference KV-Cache Offloading

Why mmap-and-let-the-OS-page-cache Breaks Down

On a memory-constrained edge box, the KV-cache eventually outgrows DRAM and has to spill to NVMe. The default move — taken by FlexLLMGen, itself descended from Stanford's 2023 FlexGen — is to memory-map a file and let the kernel's generic LRU policy decide what stays resident. That policy has no idea it is serving an autoregressive cache, and three failures compound on top of each other.

Diagram generated from the paper's diagnosis of the FlexLLMGen baseline — hover any box for detail.

KV Placement Unit: Routing Tensors Around the Bottlenecks

At init, every K/V tensor per layer is wrapped into a KV Placement Unit (KPU). A budgeter reads cgroup memory stats, subtracts DMA buffers pinned for GPU transfers, and derives a usable page-cache budget B_pc. Algorithm 1 walks layers in order, filling Group 1 until that budget is spent — everything past the cutoff goes to Group 2, the NVMe-direct path.

4 KiB Alignment & the OPT-13B Parity Fix

Group 2 tensors are bound in access order, so each tensor's start LBA is just the previous tensor's end. Algorithm 2 chunks transfers to the device's max transfer size (256 KiB on SSD A), keeping every chunk a clean multiple of the 4 KiB LBA size. OPT-13B's 10 KiB per-token tensor doesn't divide cleanly on its own — the fix forces an even batch size so two tensors together land on a 20 KiB / 5-block boundary.

This is a parity fix specific to a 10 KiB tensor, not a general solver — the transcript flags this as untested on smaller, GQA-shrunk KV tensors.

Overlap Strategy: Self-Tuned Per Decode Phase

DUAL-BLADE trials one iteration of each strategy every decode phase and locks in the faster one for the rest of generation.

Latency Reduction vs. Baseline mmap

Across the memory sweep, DUAL-BLADE cuts prefill latency by up to 33.1% and decode latency by up to 42.4% on SSD A (PCIe Gen5). SSD B (older Gen4) shows the same qualitative pattern.

Baseline normalized to 100%. SSD B figures are illustrative of the paper's reported "same pattern" — exact SSD B percentages are not tabulated in the transcript.

Page-Cache Hit Ratio vs. Memory Budget

Baseline collapses into the thrashing zone. CachePolicy-Only recovers via posix_fadvise-driven eviction (still a memcpy through the cache). DUAL-BLADE recovers further because Group 2 never touches the page cache at all.

Per-Tensor Write Latency (256 KiB decode write, SSD A)

Device busy time pinned at 100% in both cases — same hardware work, ~98% less software tax.

Data-Wrangling Workloads (OPT-6.7B, 4 GB limit)

Hospital (error detection) is a near-wash: its 1.58 GB KV-cache already fits in page cache.

Evaluation Gap Matrix

Every experiment in the paper runs on OPT-6.7B/13B — dense multi-head attention, no GQA. The paper's own framing points to Llama 3 / Mistral / Qwen-class models as what's dominant on edge hardware today, and none of those are tested. Closest competitors (KVSwap, InfiniGen) and PagedAttention-style dynamic allocators are cited in related work but never benchmarked.

The static-footprint precondition — the whole KV layout must be knowable before execution — conflicts with how production stacks (vLLM, llama.cpp, LMCache) allocate KV blocks dynamically. The paper calls its design "framework-agnostic" but never integrates or benchmarks against a PagedAttention-based server.

References

  1. DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference — Jeong, Byun, Kim, Yu, Lee, Yang, Park (2026)
  2. Efficient Memory Management for LLM Serving with PagedAttention — Kwon et al. (SOSP 2023)
  3. FlexGen: High-Throughput Generative Inference with a Single GPU — Sheng et al. (2023)
  4. LLM in a Flash: Efficient LLM Inference with Limited Memory — Alizadeh et al., Apple (2024)
  5. PowerInfer: Fast LLM Serving with a Consumer-grade GPU — Song et al., SJTU IPADS (2023)
  6. KVSwap: Disk-aware KV Cache Offloading for Long-Context On-device Inference — Zhang, Xia, Wang (2025)
  7. InfiniGen: Dynamic KV Cache Management — Lee, Lee, Seo, Sim (OSDI 2024)
  8. CachedAttention (AttentionStore) — Gao, He, Sharma, Kang, Jevdjic et al. (USENIX ATC 2024)
  9. LMCache: An Efficient KV Cache Layer for Enterprise-scale LLM Inference — Liu, Cheng, Yao, An, Chen et al. (2025)
  10. BaM: GPU-initiated Storage Access — Qureshi, Mailthody, Gelado, Min et al. (ASPLOS 2023)