Toggle between a conventional GPU-centric layout and the FengHuang/TAB-style disaggregated rack. The visuals emphasize where bytes live, where traffic concentrates, and why “buy more GPUs for memory” can become structurally wasteful.
Step through the execution model. The key question is whether tensors arrive in local HBM before use, rather than after a miss stalls generation. The mini heatmap shows tier residency over decode steps for weights, KV pages, and communication buffers.
These are realistic illustrative numbers, not reproduced measurements. Use the toggles to compare idealized simulated conditions versus a messier production-style setting with bursty traffic and lower prefetch accuracy.
Hover the heatmap. Rows represent workload families; columns represent architectural pain points. The right-side matrix shows how different mitigations attack different bottlenecks: software-only, compression, or rack-scale memory disaggregation.