Interactive podcast companion

Nemotron 3 Super
Hybrid Mamba‑Transformer MoE

A visualization-first guide to NVIDIA’s open 120.6B-parameter model with only 12.7B active per token, combining LatentMoE sparsity, Mamba-style state-space layers, periodic attention, NVFP4 pretraining, native multi-token prediction, and 1M-token context support.
arXiv: 2604.12374
Authors: NVIDIA et al.
Date: 2026
Architecture: Hybrid Mamba + Attention + LatentMoE
Use case: Agentic reasoning

Claim surface

system bundle
Efficiency Capability Deployment story

Architecture flow

88 layers · 512 experts · top‑22 active
Dense text is minimized here: the point is to show where compute is sparse, where sequence processing is compressed into state, and where exact retrieval is restored with attention.

Expert routing and retrieval behavior

hover cells and experts
Three views share the same space: top‑k expert selection, sparse retrieval over long context, and compressed recurrent state flow. Hover over SVG elements to inspect token, expert, or segment behavior.

Speed–quality tradeoffs

mock interactive plots from reported values
Nemotron Super NVFP4 Nemotron Super BF16 GPT‑OSS‑120B Qwen3.5‑122B

What deserves credit?

ablation gap map
This section turns the podcast’s skepticism into pictures: architecture, precision, post-training, token budget, and serving stack are entangled. The most persuasive result is a system point, not a clean causal decomposition.

References

compact source map
[1]
Nemotron 3 Super — NVIDIA et al., 2026.
arXiv:2604.12374
[2]
LatentMoE — Elango et al., 2026.
Scholar link
[5]
GShard — Lepikhin et al., 2020.
Scholar link
[11–13]
Multi-token prediction references.
Planning via MTP
[14–16]
FP4 / precision scaling references.
Native FP4 training
[19–25]
Related podcast episodes on Mamba, MoE decoding, KV cache, sparse attention.
Podcast archive
Additional arXiv IDs detected in transcript: 2604.12374 only.