FLINT: Sharing High Bandwidth Flash for Scalable LLM Inference

Oliveira, Tavakkol, Zhu, Yüzügüler, Arulchelvan, Cavigelli, Andri, Sadrosadati, Xinglei, Mutlu, Ke, Bergman, Zhang — Huawei Zurich Research Lab, ETH Zürich, HUST — Aug 2026
arXiv:2608.25062 Trace-Driven Simulation — Not Fabricated Silicon Emerging Memory Tier: HBF

The Capacity Gap FLINT Is Trying To Close

DeepSeek-V3 needs ~1.3 TB of BF16 weights. Today's top accelerator packages ship 80–192 GB of on-package HBM — roughly an order of magnitude short, before KV cache or activations are even loaded.

Where HBF Sits In The Package

High Bandwidth Flash stacks 3D NAND dies with through-silicon vias, physically beside HBM in the accelerator package. HBM keeps fast-changing state; HBF holds read-only weights.

Compute (GPU/NPU die) HBM — KV cache / activations HBF — read-only weights

NAND Flash Hierarchy: Die → Plane → Block → Page

Hover a level to highlight it. A block is the smallest erasable unit; a page (one wordline) is the smallest readable/programmable unit.

The Punishing Timing Asymmetry (log scale, microseconds)

A read takes ~1.5 µs. Refresh — read, ECC-correct, rewrite elsewhere — costs roughly five orders of magnitude more than a single read.

Prefetch: Static Hints (H3) vs Demand-Based Coalescing (FLINT)

The burst-buffer controller on the HBF base die watches real cache-line reads, groups requests landing on the same page, and issues a plane-parallel burst — reusing HBF's own page/cache buffers instead of dedicated SRAM.

Phantom-Plane Refresh

One extra physical plane per die is hidden from the accelerator at all times. Refresh writes only ever land on the phantom, never on a plane serving live reads.

Step 1 / 4

Read-Only FTL: Metadata Footprint

Weights are written once, never updated — so FLINT skips garbage collection and wear leveling entirely, tracking one burst-to-block-page table.

FLINT vs Baselines

GPU Packages Needed to Hit 50ms TPOT SLO

FLINT needs 3.1× fewer packages than HBM-only on average — up to 8× on some MoE models. Dense Llama-3.1-405B is the exception (no savings).

Projected Lifetime vs P/E Cycle Assumption

No vendor has published an HBF endurance figure, so the paper sweeps assumed cycles — 29 days to 8 years of continuous decode.

Area Overhead

Burst-buffer controller + phantom-plane logic + read-only FTL cost 3.1% extra HBF die area — 3.9 mm² at 7nm.

Refresh Skew: Physical Wear Didn't Disappear, It Relocated

Phantom-plane refresh moves rewrite cost out of the foreground latency path, but the most popular blocks still refresh 1.1–7.9× more often than average. Hover a cell.

Per-Model Result: Five Wins, One Loss

Five of six evaluated models are MoE, where FLINT's expert-locality bet pays off. Dense Llama-3.1-405B reads every weight every token — nothing to coalesce.

Simulated Projection, Not Measured Silicon

ComponentStatus
Burst-buffer controller, phantom-plane logic, read-only FTLNever fabricated
HBM-side timingCalibrated vs Ramulator 2.0
Area modelCACTI-derived, projected 22nm → 7nm
HBF device physics (read/program/erase/endurance)Literature-derived assumptions
HBF's own citationsSanDisk marketing blog (2025) + 2 unpublished KAIST slide decks
Hardware baseline (H3)FLINT team's own reimplementation, not original silicon/code
Sophisticated software baselines (LLM in a Flash, FlexGen)Cited in related work, never run as a baseline
1,205× over HBM+SSD is the headline number — but the baseline is one NVMe drive over PCIe with no software cleverness applied, not the flash-aware software systems already in the citation list.

References