Interactive SVG companion

FengHuang for Rack-Scale LLM Inference Memory

A visualization-first companion to the podcast episode on Microsoft Research’s 2025 FengHuang vision paper: shifting LLM inference from GPU-local-HBM-centered design toward rack-scale disaggregated tensor memory with TAB, active tensor paging, and predictive prefetching.
Paper: FengHuang: Next-Generation Memory Orchestration for AI Inferencing arXiv:2511.10753 Category: systems / memory hierarchy / inference serving viz permalink
93%
claimed local-memory reduction
under modeled conditions
50%
fewer GPUs in some targets
by decoupling memory from compute
16–70×
communication speedup headline
for shared-memory motifs
HBM + TAB
hot local tier + larger rack-scale
remote memory tier

Rack-scale tensor memory, not GPU-local memory as the only universe

Toggle between a conventional GPU-centric layout and the FengHuang/TAB-style disaggregated rack. The visuals emphasize where bytes live, where traffic concentrates, and why “buy more GPUs for memory” can become structurally wasteful.

Architecture flow diagram

hot / local HBM TAB / shared memory fabric remote LPDDR pool / colder tensors heavy communication path

Where the pressure moves

Mock allocation profile for a large serving deployment. In the baseline, memory demand forces extra GPUs; in FengHuang, capacity shifts toward the shared rack tier while compute stays closer to actual need.

Communication motif

Shared-memory communication can reduce repeated sender-side traffic for one-to-many patterns, but only if contention stays controlled.

Active tensor paging and the prefetcher

Step through the execution model. The key question is whether tensors arrive in local HBM before use, rather than after a miss stalls generation. The mini heatmap shows tier residency over decode steps for weights, KV pages, and communication buffers.

Execution timeline

Step 1 — decode starts with hot tensors in HBM and colder pages resident in the remote tier.

Tier residency heatmap

Blue = remote/cold, orange = warming, red = hot/local at the next moment of use.

Prefetch miss sensitivity

The promise depends on prediction quality. Tail behavior can degrade quickly as miss rate or remote contention rises.

Performance and skepticism dashboard

These are realistic illustrative numbers, not reproduced measurements. Use the toggles to compare idealized simulated conditions versus a messier production-style setting with bursty traffic and lower prefetch accuracy.

Latency / throughput / GPU count

Cost surface: memory need vs compute need

What changes across baselines

The podcast’s framing: compute still matters. The issue is the economic mismatch of using GPU scale-out to solve memory capacity problems.

Memory pressure maps across serving regimes

Hover the heatmap. Rows represent workload families; columns represent architectural pain points. The right-side matrix shows how different mitigations attack different bottlenecks: software-only, compression, or rack-scale memory disaggregation.

Pressure heatmap

Bottleneck composition radar-strip

Shared-memory operation motifs

A compact sketch of the five communication styles discussed in the episode: read, write, multicast-style fanout, gather/reduce-style collection, and synchronization/coordination.

Selected references

FengHuang: Next-Generation Memory Orchestration for AI Inferencing Li et al., 2025 · arXiv:2511.10753
Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism Shoeybi et al., 2019
vLLM: Easy, Fast, and Cheap LLM Serving with PagedAttention Kwon et al., 2023
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness Dao et al., 2022
ZeRO-Infinity: Breaking the GPU Memory Wall for Extreme Scale Deep Learning Rajbhandari et al., 2021
Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving Ying et al., 2024
GPUDirect Storage NVIDIA, 2024
Podcast episode and related memory-wall episodes podcast.do-not-panic.com
Additional arXiv IDs found in transcript: none beyond 2511.10753.