AI Post Transformers · Episode Companion

Kimi K3 Goes Trillion-Scale With Sparse MoE Efficiency

2.8 trillion total parameters, 104 billion active per token, a million-token context window, and native multimodal training from the ground up. This page visualizes the architecture — Mixture-of-Experts width, the hybrid KDA/MLA attention stack, Attention Residuals depth-routing, expert parallelism — and the evidence gaps behind the claimed 2.5× scaling-efficiency gain over Kimi K2.

896 Routed Experts, 16 Active Per Token

Hover any cell — glowing borders mark the 16 experts routed for this mock token. A dense model would light up every cell, every time.
low load
mid load
high load
routed this token

Sparsity Ratio: 56×

16 of 896 routed experts fire per token — the mechanism traces to Shazeer et al.'s Sparsely-Gated MoE (2017): total capacity scales almost independently of per-token compute.

3:1 Hybrid Attention Stack

Every block: three Kimi Delta Attention (KDA) layers, then one Gated MLA layer — repeated across the 93-layer backbone, with an extra MLA layer at the very end.

Full vs. Block Attention Residuals

Depth axis: can layer 50 selectively retrieve what layer 12 computed, instead of only seeing the accumulated residual sum?

Cross-Layer Link Count at True Scale (93 Layers)

All-pairs connections, log scale. This is the operational cost Block AttnRes buys down — from every layer to every block.

Expert Parallelism: Routing Becomes a Networking Problem

896 experts sharded across 8 GPUs (~112 each). A token's top-16 experts can land on any shard — the activation hops the datacenter fabric to get there.

Kimi K2 → K3 Parameter Growth

Table 1, each metric on its own scale.

Stable LatentMoE, At 896 Experts

Two failure modes appear at this sparsity, and two fixes:

The routed path chains ~4 matmuls, producing exploding activations — fixed with RMSNorm before the up-projection and a soft-capped SiTU-GLU activation in place of unbounded SwiGLU. Load balancing across ~900 experts breaks the usual bias-update rule — replaced with Quantile Balancing: each expert's bias is set from its score quantile against target load, estimated cheaply via histograms at global-batch scale.

Validation Loss Scaling: K2 vs K3

Figure 7's headline number — 2.5× efficiency gain — bundles KDA, Attention Residuals, Stable LatentMoE, and the data recipe together. No per-component ablation isolates what any single piece contributes.

Benchmark Standing

K3 trails Claude Fable 5 and GPT-5.6 Sol, leads every other model tested. Toggle category:

Evidence Confidence by Component

Blue = well-grounded, red = resting on the inventors' word alone. KDA and Gated MLA carry published lineages (Kimi Linear 2025; DeepSeek-V2 2024), and KDA's speculative-decoding trick was independently reproduced by Dao AI Lab's ReplaySSM. Attention Residuals traces almost entirely to a same-team, same-year preprint — including the "N≈8" claim and the Full-vs-Block comparison — and its own stated affordability condition, L < 100, is already strained at K3's 93 layers.
high confidence
mixed
self-cited only

References