The Capacity Gap FLINT Is Trying To Close
DeepSeek-V3 needs ~1.3 TB of BF16 weights. Today's top accelerator packages ship 80–192 GB of on-package HBM — roughly an order of magnitude short, before KV cache or activations are even loaded.
Where HBF Sits In The Package
High Bandwidth Flash stacks 3D NAND dies with through-silicon vias, physically beside HBM in the accelerator package. HBM keeps fast-changing state; HBF holds read-only weights.
NAND Flash Hierarchy: Die → Plane → Block → Page
Hover a level to highlight it. A block is the smallest erasable unit; a page (one wordline) is the smallest readable/programmable unit.
The Punishing Timing Asymmetry (log scale, microseconds)
A read takes ~1.5 µs. Refresh — read, ECC-correct, rewrite elsewhere — costs roughly five orders of magnitude more than a single read.
Prefetch: Static Hints (H3) vs Demand-Based Coalescing (FLINT)
The burst-buffer controller on the HBF base die watches real cache-line reads, groups requests landing on the same page, and issues a plane-parallel burst — reusing HBF's own page/cache buffers instead of dedicated SRAM.
Phantom-Plane Refresh
One extra physical plane per die is hidden from the accelerator at all times. Refresh writes only ever land on the phantom, never on a plane serving live reads.
Read-Only FTL: Metadata Footprint
Weights are written once, never updated — so FLINT skips garbage collection and wear leveling entirely, tracking one burst-to-block-page table.
FLINT vs Baselines
GPU Packages Needed to Hit 50ms TPOT SLO
FLINT needs 3.1× fewer packages than HBM-only on average — up to 8× on some MoE models. Dense Llama-3.1-405B is the exception (no savings).
Projected Lifetime vs P/E Cycle Assumption
No vendor has published an HBF endurance figure, so the paper sweeps assumed cycles — 29 days to 8 years of continuous decode.
Area Overhead
Burst-buffer controller + phantom-plane logic + read-only FTL cost 3.1% extra HBF die area — 3.9 mm² at 7nm.
Refresh Skew: Physical Wear Didn't Disappear, It Relocated
Phantom-plane refresh moves rewrite cost out of the foreground latency path, but the most popular blocks still refresh 1.1–7.9× more often than average. Hover a cell.
Per-Model Result: Five Wins, One Loss
Five of six evaluated models are MoE, where FLINT's expert-locality bet pays off. Dense Llama-3.1-405B reads every weight every token — nothing to coalesce.
Simulated Projection, Not Measured Silicon
| Component | Status |
|---|---|
| Burst-buffer controller, phantom-plane logic, read-only FTL | Never fabricated |
| HBM-side timing | Calibrated vs Ramulator 2.0 |
| Area model | CACTI-derived, projected 22nm → 7nm |
| HBF device physics (read/program/erase/endurance) | Literature-derived assumptions |
| HBF's own citations | SanDisk marketing blog (2025) + 2 unpublished KAIST slide decks |
| Hardware baseline (H3) | FLINT team's own reimplementation, not original silicon/code |
| Sophisticated software baselines (LLM in a Flash, FlexGen) | Cited in related work, never run as a baseline |