1 00:00:01,000 --> 00:00:50,299 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today we're digging into a paper called FLINT: Efficiently Leveraging High Bandwidth Flash for Capacity-Scalable LLM Inference Acceleration. Co-primary authors Geraldo F. Oliveira and Arash Tavakkol, with eleven more co-authors for thirteen total, out of Huawei Technologies Switzerland AG, Huawei Technologies Co., Ltd., ETH Zürich, and HUST. This one hit arXiv on August 25th, 2026. And Ada, the number that stopped me cold reading this: DeepSeek-V3 needs about 1.3 terabytes of weights in BF16, and a top-end GPU package today ships with, what, 80 to 192 gigabytes of on-package memory? 2 00:00:50,299 --> 00:01:23,275 [Dr. Ada Shannon] Right, so you're off by roughly an order of magnitude before you've even loaded the KV cache or activations. And that gap is the entire premise of this paper. The industry's default fix has been brute force — bolt on more GPU packages until you've got enough HBM stacks to hold the model, even when a single package already has plenty of compute for your batch size. You're buying capacity and getting compute you didn't need, whether you wanted it or not. FLINT's pitch is that there's a cheaper way to buy that capacity: a new memory tier called high bandwidth flash, sitting right next to HBM instead of instead of it. 3 00:01:23,275 --> 00:01:37,500 [Hal Turing] Okay, so let's actually define that, because I don't think we've ever talked about flash memory on this show in any real depth. High bandwidth flash, HBF — walk me through it like I've only ever thought about DRAM. 4 00:01:37,500 --> 00:02:16,550 [Dr. Ada Shannon] Sure. Take ordinary 3D NAND flash — the stuff inside every SSD and USB stick you own — and instead of putting it behind a PCIe bus and a filesystem, you stack the dies vertically, wire them up with through-silicon vias the same way HBM stacks DRAM dies, and drop that stack physically next to the accelerator in the same package. So you keep flash's two best properties, dirt-cheap cost per bit and huge density, multi-terabyte in a small footprint, but you give it a much fatter, more direct electrical path to the compute than a SATA or NVMe drive ever gets. It's flash wired up like memory instead of wired up like storage. 5 00:02:16,550 --> 00:02:29,350 [Hal Turing] And the idea is HBM holds the stuff that changes constantly — KV cache, activations — while HBF holds the weights, which basically never change once the model's loaded. 6 00:02:29,350 --> 00:02:52,400 [Dr. Ada Shannon] Exactly, and that read-only framing matters a lot for what's coming later. To be clear though — this is not a shipping product. HBF is, in the paper's own words, emerging. There's no JEDEC-standardized part like there is for HBM3E. This is Huawei's Zurich research lab, with ETH Zurich and HUST, proposing the substrate a system should use once that hardware exists, sort of like how HBM memory controllers got designed before HBM was common on the shelf. 7 00:02:52,400 --> 00:03:03,175 [Hal Turing] Oh wait, wait — before we go further, I need the NAND vocabulary, because I know this trips people up. Die, plane, block, page — what's the actual hierarchy? 8 00:03:03,175 --> 00:03:45,025 [Dr. Ada Shannon] Fair, let's get it out of the way now so nothing later is a mystery. A flash die is divided into planes. Each plane holds hundreds of blocks, plus a page buffer and a cache buffer that stage data in and out. A block is the smallest unit you can erase, built from strings of NAND cells. A page — one wordline within a block — is the smallest unit you can read or program. And the timing is wildly asymmetric: a page read is roughly one to two microseconds, a program is around fifty microseconds, and an erase is a few milliseconds. Reads also physically disturb neighboring cells, so periodically you have to refresh a block — read it, correct it with error-correcting codes, and rewrite it elsewhere — and that refresh costs roughly five orders of magnitude longer than a single read. 9 00:03:45,025 --> 00:03:54,175 [Hal Turing] Five orders of magnitude. So refresh is basically the thing you desperately don't want happening while the accelerator is mid-request. 10 00:03:54,175 --> 00:04:50,300 [Dr. Ada Shannon] Which is precisely where prior HBF designs fall apart. There's really one prior design worth naming here, called H3, and the paper identifies three failure modes in that lineage. First, they lean on static, compiler-emitted prefetch hints to hide flash read latency — decided ahead of time, not adapted to what the accelerator is actually asking for. Second, refresh runs directly on the accelerator-visible critical path, so a maintenance operation can stall a live inference request. Third, they carry over SSD-class flash-translation-layer machinery built for arbitrary, unpredictable writes, even though LLM weights are written once and read forever. FLINT's whole architecture is built to fix exactly those three things, and it names its fixes a burst-buffer controller, a phantom-plane refresh mechanism, and a read-only FTL — what each of those actually does is where we're headed next. 11 00:04:50,300 --> 00:05:08,700 [Hal Turing] Okay so that's the setup — static hints picked ahead of time, refresh sitting in the foreground, an FTL built for writes nobody needs. So how does FLINT actually fix the prefetch problem? 'Stop guessing at compile time' sounds nice, but the hardware still has to know what to fetch next. 12 00:05:08,700 --> 00:05:46,750 [Dr. Ada Shannon] That's the burst-buffer controller, and it lives right on the HBF base die. Instead of a compiler deciding ahead of time which weight block to stage, it watches the accelerator's actual fine-grained cache-line reads as they arrive, groups the ones landing on the same physical page, and issues a coordinated plane-parallel read once enough requests pile up on a target. Because it reconstructs bursts from real demand instead of a predicted layer order, it doesn't need a dedicated SRAM staging buffer at all — HBF already has a page buffer and a cache buffer in every plane, and those absorb two bursts in flight, one draining out while the next senses. 13 00:05:46,750 --> 00:06:00,275 [Hal Turing] So the flash's own buffering doubles as the latency-hiding structure. Refresh feels harder to dodge though — you can't reorder physics, a block still has to get rewritten eventually. What's the trick there? 14 00:06:00,275 --> 00:06:39,225 [Dr. Ada Shannon] This is phantom-plane refresh, and it's a resource-duplication trick. Each die gets one extra physical plane beyond what the logical address space needs. At any moment that spare plane is invisible to the accelerator — the phantom — while every other plane serves reads at full bandwidth. When a block's read counter trips the disturb threshold, the controller ECC-corrects the data as it's already flowing past on a normal read, and copies it into the phantom plane using that plane's own independent program circuitry. So the actual rewriting — the slow part — never touches a plane currently serving decode traffic. Once a plane fills, it becomes the next phantom and rotation continues round-robin. 15 00:06:39,225 --> 00:06:54,175 [Hal Turing] Wait, hold on — so refresh literally never competes with a foreground read on the same plane, because it's always happening one plane over? That's a genuinely clean way to sidestep the problem instead of just scheduling around it. 16 00:06:54,175 --> 00:07:32,400 [Dr. Ada Shannon] Exactly — spatial separation instead of temporal scheduling. The third piece follows the same read-only philosophy: a read-only FTL. Since weights are written once at deployment and never touched again, you skip out-of-place updates, garbage collection, wear leveling. FLINT keeps one burst translation table mapping logical burst index to a physical block-and-page coordinate — at 2-megabyte bursts on a 512-gigabyte stack that's 256,000 entries, about a megabyte of SRAM on the base die, versus roughly a gigabyte of DRAM a page-level SSD FTL needs for the same capacity. 17 00:07:32,400 --> 00:07:41,450 [Hal Turing] All three lean on the same insight — weights don't change, so throw away everything built for writes. What did they actually test this on? 18 00:07:41,450 --> 00:08:46,700 [Dr. Ada Shannon] Six production models — five MoE, DeepSeek-V3, DeepSeek-V4-Pro, Qwen3-235B-A22B, Llama-4 Maverick, Kimi K2, plus dense Llama-3.1-405B — through an in-house trace-driven simulator, not real silicon, which we'll get back to. Three baselines: HBM-plus-SSD spilling weights to NVMe off one GPU, HBM-only sharding across enough packages to hold everything in aggregate HBM, and H3, their reimplementation of that static-prefetch design. FLINT recovers 90 to 97 percent of fetched HBF traffic as actually useful, versus H3 re-fetching 80 to 95 percent of what it pulls on the MoE models. That's 1,205x decode throughput over HBM-plus-SSD, 2.2x over HBM-only, 6.2x over H3, with energy dropping 408x, 1.1x, and 6.8x across those same three baselines respectively. To hit a 50-millisecond time-per-output-token SLO, FLINT needs 3.1x fewer GPU packages than HBM-only — up to 8x on some models. 19 00:08:46,700 --> 00:08:51,475 [Hal Turing] And the refresh overhead, plus how long these things actually last? 20 00:08:51,475 --> 00:09:28,075 [Dr. Ada Shannon] Phantom-plane refresh adds no measurable throughput hit — FLINT with refresh enabled matches FLINT without it in every configuration tested, against roughly 20x average slowdown for naive in-place refresh. Lifetime is where they're upfront about uncertainty: since no vendor has published an HBF endurance figure, they sweep P/E cycle assumptions from 10-to-the-5th to 10-to-the-7th, landing anywhere from 29 days to 8 years of continuous decode depending which number you believe. All of this costs 3.1 percent extra area on the HBF die, 3.9 square millimeters at 7 nanometers for the base die additions. 21 00:09:28,075 --> 00:09:50,150 [Hal Turing] ...since no vendor's published an HBF endurance figure yet — so that whole sweep is really just saying 'we don't know, here's the entire plausible range.' And that loops into something bigger that's been nagging me this whole episode, Ada: none of this — the burst-buffer controller, phantom-plane refresh, the FTL — actually exists in silicon, does it? 22 00:09:50,150 --> 00:10:35,850 [Dr. Ada Shannon] Correct, and it's worth being blunt about it. Every number here — throughput, energy, that 3.9 square-millimeter area figure — comes from an in-house trace-driven simulator calibrated against Ramulator 2.0 for the HBM side, plus a CACTI-derived area model projected from 22 down to 7 nanometers. No FTL, burst-buffer controller, or phantom-plane mechanism was ever fabricated. And the paper's own citations for HBF itself are a 2025 SanDisk marketing blog post and two unpublished KAIST lab presentations, not peer-reviewed silicon. So the read latency, program and erase timing, even the endurance range are literature-derived assumptions about a device category that doesn't ship. A headline of 1,205 times the SSD baseline, sitting on unmeasured device physics, is a projection, not a demonstrated result. 23 00:10:35,850 --> 00:11:21,475 [Hal Turing] And that 1,205x is the number everyone screenshots. But look what it's actually beating — HBM+SSD, one NVMe drive over PCIe Gen5, with no apparent software cleverness applied. Meanwhile their own related-work section cites LLM in a Flash — Alizadeh, Mirzadeh, and colleagues out of Apple, ACL 2024 — which already serves flash-resident LLMs in software under memory constraints. And FlexGen, Sheng and Zheng et al., ICML 2023, single-GPU offloading to disk and CPU. Neither shows up as an actual baseline. Why grade against the naive path when the sophisticated software path was sitting right there in the citation list? 24 00:11:21,475 --> 00:12:02,775 [Dr. Ada Shannon] Because the naive path makes the better headline, honestly. Their one hardware baseline is H3 — Ha, Kim, and Kim, IEEE Computer Architecture Letters, 2026 — and even that rests on FLINT's own team reimplementing H3 rather than running the original code or hardware. So 'static prefetch wastes 86 to 96 percent of traffic' is FLINT grading its own reimplementation of a rival design. Interesting thread underneath, though — co-author Ahmet Caner Yüzügüler previously worked on PRESERVE, a software weight-and-KV-cache prefetcher with Zhuang and Cavigelli, arXiv 2025. You can read FLINT as that same team deciding prefetch hints in software weren't enough and moving the logic into hardware instead. 25 00:12:02,775 --> 00:12:26,075 [Hal Turing] Oh wait wait — hold on, something's bugging me about the traces. They captured MoE routing from a real GPU deployment, but serving MMLU prompts — short, closed-form, multiple-choice. The headline numbers are all reported at 128K-token context. How do you stretch benchmark routing patterns out to long-context generation without just guessing? 26 00:12:26,075 --> 00:13:04,550 [Dr. Ada Shannon] They don't really spell out the extrapolation, which is a real gap — routing under long, open-ended generation or tool calls could look nothing like MMLU, and expert-selection locality is exactly what the burst-buffer controller is betting on. There's a second representativeness problem sitting in the results too: dense Llama-3.1-405B is the one model where FLINT actually loses — 1.64 times the energy of HBM-only, zero GPU-package savings at batch 64, because every token reads every weight with nothing to coalesce. Five of the six evaluated models are MoE, so the headline averages lean hard on the workload class where FLINT's trick happens to shine. 27 00:13:04,550 --> 00:13:49,900 [Hal Turing] There's a structural blind spot too, right? The read-only FTL — no garbage collection, no out-of-place writes — assumes HBF only ever holds immutable base weights. But the industry's heading toward per-tenant LoRA adapters and cartridges sitting near the accelerator for capacity reasons, written far more often than base weights ever are. If any of that mutable state needs HBF's capacity, the read-only premise breaks. And on power loss — they treat every power-up as a clean install, about 6.7 minutes to erase and reload. So the full weight set has to come from somewhere external after any outage, quietly reintroducing the off-package dependency HBF was supposed to remove. 28 00:13:49,900 --> 00:14:32,800 [Dr. Ada Shannon] Even the read-only wear story isn't fully clean — their own data shows the most popular block refreshes 1.1 to 7.9 times more often than average. Phantom-plane refresh moves that cost out of foreground latency, but the physical wear skew across cells is still there, just relocated from the latency domain into the endurance domain. Worth placing this in the wider landscape, too — MemExplorer, Wu and Cao et al., 2026, treats HBF as one tier in a broader heterogeneous-memory design space without proposing new control mechanisms, and Lincoln, Sun and Gao et al., HPCA 2025, bets on compute-inside-flash over an LPDDR interface rather than keeping compute on the GPU. FLINT is a narrow, specific bet inside a much wider design space. 29 00:14:32,800 --> 00:15:15,050 [Hal Turing] So if the mechanisms genuinely hold up in real silicon, the upside is real — single-package serving of frontier MoE models, meaningfully fewer GPUs to hit latency SLOs, energy savings from skipping cross-package synchronization. Worth someone building. But what exists today is a well-argued architecture paper whose numbers are simulated projections stacked on vendor-unconfirmed device assumptions, benchmarked mostly against a self-reimplemented rival and a weak SSD baseline. The engineering ideas are genuinely clean. The confidence attached to the numbers should be a lot lower than the abstract implies. 30 00:15:15,050 --> 00:15:35,300 [Dr. Ada Shannon] What's needed next is obvious — real HBF silicon, vendor-published endurance and timing figures, and a rerun against LLM in a Flash or FlexGen-class software baselines instead of naive SSD spillover. Until then, FLINT is a solid hardware proposal wrapped around a parameter sweep dressed up as a reliability result. 31 00:15:35,300 --> 00:15:50,926 [Hal Turing] Good place to land it — sound architecture, thin evidentiary floor. FLINT is three clever, well-motivated mechanisms for an HBF tier that doesn't physically exist yet. Thanks for listening, everyone — we'll catch you next time.