896 Routed Experts, 16 Active Per Token
Sparsity Ratio: 56×
3:1 Hybrid Attention Stack
Full vs. Block Attention Residuals
Cross-Layer Link Count at True Scale (93 Layers)
Expert Parallelism: Routing Becomes a Networking Problem
Kimi K2 → K3 Parameter Growth
Stable LatentMoE, At 896 Experts
The routed path chains ~4 matmuls, producing exploding activations — fixed with RMSNorm before the up-projection and a soft-capped SiTU-GLU activation in place of unbounded SwiGLU. Load balancing across ~900 experts breaks the usual bias-update rule — replaced with Quantile Balancing: each expert's bias is set from its score quantile against target load, estimated cheaply via histograms at global-batch scale.
Validation Loss Scaling: K2 vs K3
Benchmark Standing
Evidence Confidence by Component
References
- 1Kimi K3 Technical Report — Kimi Team, Moonshot AI
- 2Deep Residual Learning for Image Recognition — He, Zhang, Ren, Sun, 2015 (CVPR 2016)
- 3DenseFormer — Pagliardini, Mohtashami, Fleuret, Jaggi, 2024
- 4Value Residual Learning — Zhou, Wu, Jiang, Lan et al., 2024
- 5Hyper-Connections — Zhu, Huang, Huang, Zeng, Mao, Wu, Min, Zhou (ByteDance Seed), 2024
- 6Attention Residuals — Kimi Team, 2026
- 7Kimi Linear — Kimi Team et al., 2025 (arXiv:2510.26692)
- 8DeepSeek-V2 — DeepSeek-AI, 2024
- 9ReplaySSM: Cache SSM Inputs, Not State — Dao AI Lab, 2026 (blog)