References
- NVIDIA Nemotron 3 White Paper (arXiv:2512.20856)
- Transformers are SSMs (Dao & Gu, 2024)
- Better & Faster Large Language Models via Multi-Token Prediction (Gloeckle et al., 2024)
- Thinking Tokens for Language Modeling (Herel & Mikolov, 2023)
- Latent Prototype Routing (Approximate, 2024)
- Attn-QAT: 4-Bit Attention With Quantization-Aware Training (Approximate, 2024)
- FP4 All the Way: Fully Quantized Training of LLMs (Approximate, 2024)