Why mmap-and-let-the-OS-page-cache Breaks Down
On a memory-constrained edge box, the KV-cache eventually outgrows DRAM and has to spill to NVMe. The default move — taken by FlexLLMGen, itself descended from Stanford's 2023 FlexGen — is to memory-map a file and let the kernel's generic LRU policy decide what stays resident. That policy has no idea it is serving an autoregressive cache, and three failures compound on top of each other.
KV Placement Unit: Routing Tensors Around the Bottlenecks
At init, every K/V tensor per layer is wrapped into a KV Placement Unit (KPU). A budgeter reads cgroup memory stats, subtracts DMA buffers pinned for GPU transfers, and derives a usable page-cache budget B_pc. Algorithm 1 walks layers in order, filling Group 1 until that budget is spent — everything past the cutoff goes to Group 2, the NVMe-direct path.
4 KiB Alignment & the OPT-13B Parity Fix
Group 2 tensors are bound in access order, so each tensor's start LBA is just the previous tensor's end. Algorithm 2 chunks transfers to the device's max transfer size (256 KiB on SSD A), keeping every chunk a clean multiple of the 4 KiB LBA size. OPT-13B's 10 KiB per-token tensor doesn't divide cleanly on its own — the fix forces an even batch size so two tensors together land on a 20 KiB / 5-block boundary.
Overlap Strategy: Self-Tuned Per Decode Phase
DUAL-BLADE trials one iteration of each strategy every decode phase and locks in the faster one for the rest of generation.
Latency Reduction vs. Baseline mmap
Across the memory sweep, DUAL-BLADE cuts prefill latency by up to 33.1% and decode latency by up to 42.4% on SSD A (PCIe Gen5). SSD B (older Gen4) shows the same qualitative pattern.
Page-Cache Hit Ratio vs. Memory Budget
Baseline collapses into the thrashing zone. CachePolicy-Only
recovers via posix_fadvise-driven eviction (still a memcpy
through the cache). DUAL-BLADE recovers further because Group 2 never
touches the page cache at all.
Per-Tensor Write Latency (256 KiB decode write, SSD A)
Data-Wrangling Workloads (OPT-6.7B, 4 GB limit)
Evaluation Gap Matrix
Every experiment in the paper runs on OPT-6.7B/13B — dense multi-head attention, no GQA. The paper's own framing points to Llama 3 / Mistral / Qwen-class models as what's dominant on edge hardware today, and none of those are tested. Closest competitors (KVSwap, InfiniGen) and PagedAttention-style dynamic allocators are cited in related work but never benchmarked.