Nvidia keeps a small, blazing-fast dedicated memory pool next to compute. Apple gives the whole chip one large, slower, shared pool. Everything downstream — the offload cliff, the quantization pressure, the efficiency gap — follows from this one architectural fork.
RTX 5090 talks to its own GDDR7 pool at ~1800 GB/s — but that pool is fixed at 32GB in silicon. Once weights exceed it, the only path to more capacity is across a PCIe bus running 32–64 GB/s: 25–55x slower.
CPU and GPU share one LPDDR pool, up to 512GB on the largest configs. No VRAM/system-RAM split means no PCIe bottleneck to fall into — at the cost of lower per-byte bandwidth than dedicated GDDR.
Hardware doesn't decide the outcome alone — the runtime sitting on top determines whether that hardware gets used properly. And on Nvidia, picking the wrong runtime silently erases your quantization gains.
Nvidia. Ahead-of-time compilation into a fixed engine per GPU. v1.1.0 added NVFP4 — 4-bit float tuned for Blackwell — but split into two backends that behave very differently: the "Backend Dichotomy."
Apple. No compilation step. Lazy evaluation builds the computation graph at runtime, with native 4-bit quantization directly on the unified pool — no VRAM spike while quantizing.
Cross-platform fallback. Offloads whatever doesn't fit in VRAM to CPU + system RAM. The mechanism behind Nvidia's collapse once a model outgrows 32GB.
Qwen3-8B, NVFP4-quantized, same RTX 5090. The 1.6x NVFP4 speedup only exists on the PyTorch backend — pick the legacy C++ engine and you paid the 4-bit accuracy cost for zero throughput gain over plain BF16.
Time-to-first-token, Qwen3-8B. The brand-new 5090 loses to its own predecessor by 2.2x — not a hardware regression, but Blackwell's TensorRT-LLM prefill path still catching up on day one.
Qwen2.5-1.5B fits comfortably everywhere. This is what each platform looks like with the VRAM Wall removed entirely — and it inverts depending on what you optimize for.
RTX 5090 leads raw throughput by 70% — 265 vs 155 tok/s. Flip the metric and the M3 Ultra delivers up to 23x the tokens-per-joule of the 5090.
Past 30B parameters the question stops being "which chip is fastest" and becomes "which chip still works at all." Table 5 of the paper is the whole argument in one grid.
RTX 5090 with aggressive quantization (IQ3_XXS / Q2_K_XL) stays healthy. Keep higher-fidelity Q4_K_M and offload ~25% of layers over PCIe, and throughput craters >90%. Apple runs the same 70–80B models natively, no offload, landing in between.
On dense models every token touches every parameter, so bandwidth wins outright: M3 Ultra's ~800 GB/s crushes the M4 Pro's 273 GB/s. On MoE models only a slice of parameters activate per token — bandwidth stops dominating, and the two MoE results split in opposite directions by single digits. That's noise dressed as a trend, not a law of MoE inference.