NVIDIA Nemotron 3 Hybrid SSM Transformer Architecture

White Paper published Dec 24, 2025 — arXiv:2512.20856
arXiv:2512.20856
Input Tokens Mamba-2 SSM + MoE Blocks (×7) Mamba-2 MoE Mamba-2 MoE Mamba-2 MoE Mamba-2 Sparse Self-Attention Layers (×1) Attention Output Tokens Token Embeddings Mamba-2 + MoE Blocks Sparse Attention Final Output

References

  1. NVIDIA Nemotron 3 White Paper (arXiv:2512.20856)
  2. Transformers are SSMs (Dao & Gu, 2024)
  3. Better & Faster Large Language Models via Multi-Token Prediction (Gloeckle et al., 2024)
  4. Thinking Tokens for Language Modeling (Herel & Mikolov, 2023)
  5. Latent Prototype Routing (Approximate, 2024)
  6. Attn-QAT: 4-Bit Attention With Quantization-Aware Training (Approximate, 2024)
  7. FP4 All the Way: Fully Quantized Training of LLMs (Approximate, 2024)