AI Post Transformers · Episode Companion

Silicon Showdown: GPU vs Apple Silicon LLM Inference Limits

Based on "Silicon Showdown: Performance, Efficiency, and Ecosystem Barriers in Consumer-Grade LLM Inference" — Abdurrahman Javat, Allan Kazakov (Bahçeşehir University, 2026)
arXiv:2605.00519 1.5B → 80B params tested 11 chips · 2 ecosystems May 2026

Two Memory Philosophies

Nvidia keeps a small, blazing-fast dedicated memory pool next to compute. Apple gives the whole chip one large, slower, shared pool. Everything downstream — the offload cliff, the quantization pressure, the efficiency gap — follows from this one architectural fork.

Nvidia: Discrete GPU + PCIe hop

RTX 5090 talks to its own GDDR7 pool at ~1800 GB/s — but that pool is fixed at 32GB in silicon. Once weights exceed it, the only path to more capacity is across a PCIe bus running 32–64 GB/s: 25–55x slower.

Apple: Unified Memory Architecture

CPU and GPU share one LPDDR pool, up to 512GB on the largest configs. No VRAM/system-RAM split means no PCIe bottleneck to fall into — at the cost of lower per-byte bandwidth than dedicated GDDR.

Three Ecosystems, Three Trade-offs

Hardware doesn't decide the outcome alone — the runtime sitting on top determines whether that hardware gets used properly. And on Nvidia, picking the wrong runtime silently erases your quantization gains.

TensorRT-LLM

Nvidia. Ahead-of-time compilation into a fixed engine per GPU. v1.1.0 added NVFP4 — 4-bit float tuned for Blackwell — but split into two backends that behave very differently: the "Backend Dichotomy."

MLX

Apple. No compilation step. Lazy evaluation builds the computation graph at runtime, with native 4-bit quantization directly on the unified pool — no VRAM spike while quantizing.

GGUF (llama.cpp)

Cross-platform fallback. Offloads whatever doesn't fit in VRAM to CPU + system RAM. The mechanism behind Nvidia's collapse once a model outgrows 32GB.

Qwen3-8B, NVFP4-quantized, same RTX 5090. The 1.6x NVFP4 speedup only exists on the PyTorch backend — pick the legacy C++ engine and you paid the 4-bit accuracy cost for zero throughput gain over plain BF16.

Time-to-first-token, Qwen3-8B. The brand-new 5090 loses to its own predecessor by 2.2x — not a hardware regression, but Blackwell's TensorRT-LLM prefill path still catching up on day one.

No Memory Pressure: Raw Character

Qwen2.5-1.5B fits comfortably everywhere. This is what each platform looks like with the VRAM Wall removed entirely — and it inverts depending on what you optimize for.

RTX 5090 leads raw throughput by 70% — 265 vs 155 tok/s. Flip the metric and the M3 Ultra delivers up to 23x the tokens-per-joule of the 5090.

Where Capacity Actually Matters

Past 30B parameters the question stops being "which chip is fastest" and becomes "which chip still works at all." Table 5 of the paper is the whole argument in one grid.

collapsed (2–4 tok/s) usable (13–49 tok/s) healthy (46–76 tok/s)

RTX 5090 with aggressive quantization (IQ3_XXS / Q2_K_XL) stays healthy. Keep higher-fidelity Q4_K_M and offload ~25% of layers over PCIe, and throughput craters >90%. Apple runs the same 70–80B models natively, no offload, landing in between.

On dense models every token touches every parameter, so bandwidth wins outright: M3 Ultra's ~800 GB/s crushes the M4 Pro's 273 GB/s. On MoE models only a slice of parameters activate per token — bandwidth stops dominating, and the two MoE results split in opposite directions by single digits. That's noise dressed as a trend, not a law of MoE inference.

References

  1. Javat, Kazakov. Silicon Showdown: Performance, Efficiency, and Ecosystem Barriers in Consumer-Grade LLM Inference. 2026. arXiv:2605.00519
  2. Frantar, Ashkboos, Hoefler, Alistarh. GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers. 2022. Scholar
  3. Lin, Tang, Tang, Yang, Dang, Han. AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration. 2023. Scholar
  4. Dettmers, Lewis, Belkada, Zettlemoyer. LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale. 2022. Scholar
  5. Jiang et al. (Mistral AI). Mixtral of Experts. 2024. Scholar
  6. Leviathan, Kalman, Matias. Fast Inference from Transformers via Speculative Decoding. 2023. Scholar
  7. Kwon et al. Efficient Memory Management for Large Language Model Serving with PagedAttention. 2023. Scholar