1 00:00:01,000 --> 00:00:37,025 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today's paper is 'Kimi K3: Open Frontier Intelligence' — a technical report credited simply to 'Kimi Team,' out of Moonshot AI, no individual byline. And the numbers are the hook: 2.8 trillion total parameters, only 104 billion active per token, a million-token context window, native vision from the ground up, and — notably — they released the actual weights, open, at frontier scale. That combination barely existed a year ago. 2 00:00:37,025 --> 00:01:13,875 [Dr. Ada Shannon] Here's why it's worth your time: most 'open' frontier attempts either underperform their own claims or just scale by buying more GPUs. What sold me on K3 is that the gains are grounded in actual architectural theory, not brute force. The report frames progress along two axes — pretraining scale, the expensive one, where open labs have historically lagged, and test-time or RL scale, squeezing more reasoning out of a fixed model through longer thinking and agentic rollouts, where open models have mostly caught up. K3 is Moonshot's attempt to close that gap without just buying more H100s. 3 00:01:13,875 --> 00:01:35,850 [Hal Turing] Before we go further, let's ground the MoE claim, because '2.8 trillion parameters, only 104 billion active' sounds almost like a trick — like owning a huge library but only ever reading four books off the shelf every visit. Is that a fair comparison, Ada, or does it undersell what's actually happening architecturally? 4 00:01:35,850 --> 00:02:19,700 [Dr. Ada Shannon] It's a decent first approximation. A dense transformer runs every parameter on every token. An MoE network instead replaces the big feedforward blocks with a bank of smaller, specialized 'experts,' plus a router that sends each token to just a handful of them. So a token only pays for the experts it's routed to. The modern version traces back to the Sparsely-Gated Mixture-of-Experts paper, Noam Shazeer and collaborators at Google Brain, 2017, reviving an old ensemble-learning idea for deep learning. The payoff: total capacity grows almost independently of per-token compute cost, since you only pay for the active slice. That's how K3 gets 2.8 trillion parameters of capacity while spending roughly what a 104-billion-parameter dense model costs per forward pass. 5 00:02:19,700 --> 00:02:35,975 [Hal Turing] That's the width story. But the report also talks about a depth axis — how information moves down through dozens of stacked layers. What's the baseline there? What actually happens today when a signal from an early layer needs to reach a very late one? 6 00:02:35,975 --> 00:03:09,725 [Dr. Ada Shannon] The near-universal mechanism is the residual connection: each layer computes output equals input plus that layer's contribution, so you get one running stream that accumulates additively through every layer. That's from Deep Residual Learning for Image Recognition, Kaiming He and coauthors at Microsoft Research, 2015 — the ResNet paper, arguably the idea that made very deep networks trainable. But notice the limit: layer fifty can't reach back and specifically ask what layer twelve computed. It only ever sees the accumulated sum, already blended with everything added since. 7 00:03:09,725 --> 00:03:32,375 [Hal Turing] Okay, filed away for later. Let's talk scale in the other direction — sequence length. A million-token context window is enormous, practically a whole codebase sitting in working memory. And 'Expert Parallelism' keeps showing up in the infrastructure sections. Are those two things actually related, or separate engineering problems bolted onto the same model? 8 00:03:32,375 --> 00:04:09,900 [Dr. Ada Shannon] Related only in that both are 'scale' problems solved at different layers of the stack. A long-context LLM is engineered to reason coherently over that whole million-token input without the usual degradation — plenty of models accept long inputs but effectively forget most of what's past the first chunk. That's a modeling problem, mostly attention and position-encoding design. Expert Parallelism is purely infrastructural: once you've got hundreds of experts in an MoE model, they can't all fit on one GPU, so you shard them across many GPUs and route each token to whichever machine holds the expert it needs. 9 00:04:09,900 --> 00:04:26,850 [Hal Turing] Wait, wait, hold on — so every single token might have to hop across the network to a totally different physical machine just to get processed? That's not a software problem anymore, that's a datacenter networking problem wearing an ML paper's clothes. 10 00:04:26,850 --> 00:04:58,350 [Dr. Ada Shannon] Pretty much — it becomes as much a systems-engineering problem as a machine-learning one, and we'll get into how K3 handles that imbalance later. One more background piece first: vision. K3 is natively multimodal — text, images, and video share one token stream and one backbone from the start of training, rather than bolting a separately pretrained encoder like CLIP or SigLIP onto a text-only model afterward. Post-hoc bolt-ons ship faster; native training tends to be more stable, since nothing has to reconcile two separately learned representations. 11 00:04:58,350 --> 00:05:19,150 [Hal Turing] Okay — so width via Mixture-of-Experts, a depth axis where plain residuals hit a ceiling, a million-token context, expert parallelism holding all of that up physically, and vision baked in natively instead of bolted on afterward. What does Moonshot actually claim they got out of stacking all of that together? 12 00:05:19,150 --> 00:05:40,150 [Dr. Ada Shannon] Roughly a two-and-a-half times improvement in overall scaling efficiency over their own prior model, Kimi K2. And it's attributed jointly — Kimi Delta Attention, Attention Residuals, Stable LatentMoE, and refined training and data recipes, all credited together rather than to any single piece. Untangling which one's actually doing the heavy lifting is exactly where we're headed next. 13 00:05:40,150 --> 00:06:02,525 [Hal Turing] Okay, let's open up the block diagram, because 'hybrid attention' is doing a lot of hand-waving in that abstract. What's actually stacked inside one of these ninety-three layers — some fixed ratio between the cheap long-context mechanism and the expensive global one? And what happened to positional encoding — I don't see RoPE mentioned anywhere in my notes. 14 00:06:02,525 --> 00:06:58,351 [Dr. Ada Shannon] Three to one, not alternating. Every block runs three Kimi Delta Attention layers, then one Gated MLA layer — a compressed-KV global attention design from the DeepSeek-V2 paper, DeepSeek-AI, 2024 — and that pattern repeats down the backbone, with an extra MLA layer at the very end so the last thing before output is always full global attention. KDA itself is a delta-rule recurrence with a channel-wise forget gate: instead of a key-value cache that keeps growing with sequence length, it keeps one fixed-size recurrent state and selectively decays old information per channel. The tricky part was their decay term, which could blow up in low precision. K3's fix is a lower-bounded sigmoid decay that floors every retention factor above a fixed value, letting both diagonal and off-diagonal tiles run as dense Tensor Core matrix multiplies instead of falling back to a slow position-by-position path. 15 00:06:58,351 --> 00:07:20,276 [Hal Turing] Oh, wait, hold on — sorry to cut in, but 'fixed-size state instead of a growing cache' is the part I want to sit on. That's memory that doesn't scale with conversation length at all. But depth is still bugging me — how does that same selective-retrieval idea play out for Attention Residuals? Full versus Block, since I know that's the actual production choice. 16 00:07:20,276 --> 00:07:55,801 [Dr. Ada Shannon] Right — you're compressing everything into that one fixed slot instead of keeping a full history around, that's the cost side. On depth: Full AttnRes gives every layer a learned pseudo-query that attends individually over every preceding layer's output — expensive to keep resident in memory, but it lets a layer selectively retrieve from any single earlier layer. Block AttnRes is the production compromise: group layers into blocks, sum each into one representation, and only attend across those block-level summaries. K3 ships Block, split into eight blocks of twelve layers each. 17 00:07:55,801 --> 00:08:20,551 [Hal Turing] Eight blocks specifically feels oddly precise — not four, not sixteen. Did they actually run that sweep at K3's own ninety-three-layer, 2.8-trillion-parameter scale and watch it flatten out right there? Or does that number come from somewhere else in their own research pipeline — is it tested at this exact scale, or extrapolated from something smaller and just carried over? Genuinely curious which one it is. 18 00:08:20,551 --> 00:09:00,777 [Dr. Ada Shannon] Somewhere else — worth stating plainly: the paper's justification is 'N approximately eight recovers most of the benefit across model scales,' cited to a separate companion paper, not derived from K3-scale experiments in this report. Same goes for AttnRes's ablations generally — they live in that other paper. Not passing judgment right now, just flagging it as a fact. What Block AttnRes does buy them operationally is concrete, though: overhead drops from scaling with the full layer count down to scaling with just the block count, and it lets them merge block-level results with the sequential intra-block computation through online softmax instead of keeping every layer's output alive. 19 00:09:00,777 --> 00:09:23,902 [Hal Turing] Filed for later, then. Let's go wide instead of deep — eight hundred ninety-six routed experts, only sixteen active per token. That's a sparsity number I want you to actually put in plain terms. And separately, I noticed the vision encoder got trained a completely different way this time, not the pretrain-and-bolt-on approach we talked about earlier. What changed, and why? 20 00:09:23,902 --> 00:10:20,527 [Dr. Ada Shannon] That's a sparsity ratio of fifty-six, and at that extreme two things break. First, the routed path chains nearly four matrix multiplications together, and at this scale that ill-conditioned chain produces exploding activations. The fix: RMSNorm before the up-projection, plus a new activation called SiTU-GLU that soft-caps both branches of the gate instead of running unbounded like the SwiGLU most models use. Second, load balancing breaks down with nearly nine hundred experts — the usual bias-update rule oscillates or adapts too slowly. They swap in Quantile Balancing, setting each expert's bias from the score quantile matching its target load, estimated via histograms so it's cheap at global-batch scale. On vision, MoonViT-V2 is trained entirely from scratch with next-token prediction instead of initializing from a contrastive model like SigLIP — purely for training stability. The SigLIP-initialized version showed spiky gradient norms; the from-scratch version stayed smooth, and still matches the SigLIP baseline on vision evals. 21 00:10:20,527 --> 00:10:48,402 [Hal Turing] So what's the actual payoff, once you tally all of that up? I want the real numbers now — the scaling curve against K2, the architecture delta table, and where K3 actually lands against real competitors instead of just the abstract's headline claim. Because stacking three separate architectural bets into one 2.8-trillion-parameter model and having it train at all is, credit where it's due, a genuinely impressive engineering feat regardless of how the numbers shake out. 22 00:10:48,402 --> 00:11:56,077 [Dr. Ada Shannon] Figure 7's scaling curves put that validation-loss gain at 2.5x over K2 — and worth flagging, that number bundles KDA, AttnRes, Stable LatentMoE, and the data recipe together, with no per-component ablation isolating what AttnRes alone contributes. Table 1 spells out where the parameters went: layers sixty-one to ninety-three, total params 1.04 trillion to 2.78 trillion, routed experts 384 to 896, active experts eight to sixteen, context 128K to a million tokens. On the eval suite, K3 trails Claude Fable 5 and GPT-5.6 Sol but leads every other open and proprietary model tested, across coding, agentic, knowledge, reasoning, and vision. MoonEP holds the training side up — perfect load balance with a proven bound on redundant experts per rank, instead of the capped, sometimes-stalling schemes like DeepEP, ECHO, or UltraEP. KDA Context Parallelism is what makes million-token training tractable, splitting the recurrent state's propagation across devices. 23 00:11:56,077 --> 00:12:19,752 [Hal Turing] But here's what's nagging at me: KDA, AttnRes, Stable LatentMoE, and the whole refined data and training recipe all changed together between K2 and K3. Figure 7 gives us one clean 2.5x line. If I'm being skeptical, how much of that number can we actually credit to Attention Residuals specifically, versus the other three moving at once? 24 00:12:19,752 --> 00:13:03,777 [Dr. Ada Shannon] Honestly, none of it cleanly. This report gives you no lever to isolate AttnRes — the 2.5x is one number covering four simultaneous bets. It gets thinner from there: AttnRes itself, the Full-versus-Block comparison, and that 'N approximately 8' claim are all sourced to citation 57, Attention Residuals, Kimi Team, Moonshot AI, 2026 — a preprint from the same team, same year, that this report never re-derives or independently checks. Compare that to what AttnRes is explicitly positioned against: Deep Residual Learning for Image Recognition, Kaiming He and colleagues, Microsoft Research, 2015. That mechanism earned a decade of outside replication before anyone trusted it at scale. AttnRes is going into a 2.8-trillion-parameter production model on the strength of its own inventors' word alone. 25 00:13:03,777 --> 00:13:25,777 [Hal Turing] Oh wait, hold on — sorry, but that's the thread I want to pull. The paper justifies Full AttnRes's cost as fine 'since network depth is modest, L under 100' — and K3 is already at 93 layers. So by the paper's own logic, isn't it basically conceding AttnRes hits a wall right where frontier depth is heading next? 26 00:13:25,777 --> 00:14:02,277 [Dr. Ada Shannon] Your ceiling read is fair: if L under 100 is the affordability condition for the full form, that's a real constraint being waved past, not confronted. There's a second confound too — K3 didn't just add AttnRes, it grew from 61 to 93 layers and 1 to 2.8 trillion parameters simultaneously. AttnRes adds its own learned pseudo-queries and attention weights per layer. Nothing here — no attention-weight analysis, no probing for what earlier layers actually get retrieved — separates 'the model got smarter about routing across depth' from 'the model just got more capacity, some of which happened to land in AttnRes.' 27 00:14:02,277 --> 00:14:54,777 [Hal Turing] That distinction matters practically, because the one AttnRes-specific number I found isn't a quality metric at all — Figure 14, where Kimi K3 acts as a coding agent and optimizes its own AttnRes CUDA kernel, 283.6 milliseconds down to 114.4. Neat demo of K3-the-agent writing fast code. Tells you nothing about whether AttnRes-the-mechanism improves quality, and there's something almost too on-the-nose about handing an unvalidated depth-routing trick to the model built on it and asking it to show off. Meanwhile Block AttnRes needed real new infrastructure — cache-based cross-stage block transfer, checkpointing wrapped around the computation, remote activation offload just to stay memory-tractable. Genuine engineering cost, and nowhere is it weighed against a measured quality contribution. 28 00:14:54,777 --> 00:15:42,127 [Dr. Ada Shannon] Worth separating the credit one level further, too — 'improves information flow' gets used for KDA and AttnRes almost interchangeably here. KDA traces to Kimi Linear, Kimi Team, Moonshot AI, 2025, arXiv 2510.26692 — that's the sequence axis. Gated MLA descends from DeepSeek-V2, DeepSeek-AI, 2024 — global content mixing. AttnRes is supposed to be the depth axis stacked on both. Three 'information flow' stories, one abstract sentence. What would settle this is independent replication, and there's a useful contrast right in this report: KDA's speculative-decoding replay trick was independently and concurrently invented elsewhere, by Dao AI Lab's ReplaySSM. That's real convergent validation. AttnRes has none — nobody outside Moonshot, the same team behind last cycle's Kimi K2.5 visual-agent work, has apparently tried reproducing it yet. 29 00:15:42,127 --> 00:16:17,553 [Hal Turing] So where that leaves us: KDA's fixed-state efficiency story is well-grounded, Stable LatentMoE's load-balancing fix is concrete and testable, but AttnRes is the one piece riding almost entirely on the authors' own word, at a depth their own justification says might already be strained. That's not a reason to dismiss this model — the benchmark numbers are real, and the weights are open for anyone to go test these questions themselves. It's a reason to watch whether anyone outside Moonshot actually does. Thanks for listening, and we'll catch you next time.