1 00:00:01,000 --> 00:00:46,325 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today's paper is "On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability" from the Qwen Team at Alibaba. First author Zihan Qiu plus 35 co-authors, on arXiv 31 August 2026. It's a design-rationale report on Qwen3.8-Flash-Next, an ablation log more than one new method, and it covers the base model, not post-training. It sets up its own puzzle: a bigger n-gram vocabulary lowers loss every time, while accuracy stops moving. Ada, discovery or measurement problem? 2 00:00:46,325 --> 00:01:22,925 [Dr. Ada Shannon] Pick architectures by loss alone and you can ship the wrong model, so that question is the spine of the paper. Every change is judged on three axes: loss plus benchmarks, cost in training, prefill and decode, and effect on optimal hyperparameters and stability. The headline: on fourteen pre-training benchmarks, the base model leads the 397B-A17B predecessor on eight and trails on six by at most 2.6 points, with a third of the activated parameters and a third of the tokens. One-third times one-third is roughly a ninth of the training FLOPs. 3 00:01:22,925 --> 00:01:48,950 [Hal Turing] A ninth is a big number. The intro calls that predecessor the Qwen3.5 flagship, and Table 11 labels the same 397B baseline Qwen3.7-Plus-Base. Somebody in this comparison is in witness protection. Also, the model is 125B total, 6B activated, plus 51B of n-gram embeddings off the accelerator. Which of those is the size? 4 00:01:48,950 --> 00:02:32,100 [Dr. Ada Shannon] None alone. 125B is what the accelerators hold, 6B is what the router activates per token, so it sets the FLOPs, and the 51B sit in host memory outside both. "Total" and "activated" need a third number. Now attention. Full attention is quadratic in compute, with a KV cache that grows linearly. Sliding windows aren't a free fix, since information outside the window only travels through depth. The paper's answer: three Gated DeltaNet layers, from Yang and colleagues at MIT in 2024, then one full-attention layer, repeated. The recurrent layers squeeze the prefix into a fixed-size state at linear cost. The periodic attention layer keeps the exact token retrieval a finite state can't guarantee. 5 00:02:32,100 --> 00:02:43,125 [Hal Turing] Genuine question: that one attention layer in four is still quadratic. At a million tokens, doesn't it eat the savings? Is that the gap sparse attention fills? 6 00:02:43,125 --> 00:03:03,050 [Dr. Ada Shannon] Yes. DeepSeek's DSA, Liu and colleagues at DeepSeek-AI in 2025, uses a cheap indexer to score which context matters, then runs real attention on only the top-k. But the indexer is itself O(n squared), so at long context it becomes the bill. Qwen's QSA attacks exactly that, and the way it does it is— 7 00:03:03,050 --> 00:03:13,525 [Hal Turing] Sorry to cut in, but hold that for later, because the residual stream is nagging me. The paper widens it to four branches. Is that highway networks again? 8 00:03:13,525 --> 00:03:49,750 [Dr. Ada Shannon] Same family. Highway networks, Srivastava and Schmidhuber at IDSIA in 2015, gated the skip path. DenseNet, Huang at Cornell in 2017, wired each layer to earlier ones. AltUp from Baykal at Google in 2023, Hyper-Connections from Zhu at ByteDance in 2024, and mHC replace one residual vector with parallel branches. Three operators: H_mix reads the branches into a layer, H_combine writes the layer's output back, H_res mixes branches with each other. Pre-norm attenuates signal with depth, and branches let layers read and write at different strengths. 9 00:03:49,750 --> 00:03:59,950 [Hal Turing] Now the 51B. An expert is capacity because it computes something. How is a lookup table capacity, and why does it live off the accelerator? 10 00:03:59,950 --> 00:04:54,350 [Dr. Ada Shannon] It's memory instead of computation. Short n-grams ending at each token are hashed into a very large table, and the retrieved vector enters the residual stream through a gate. The address depends only on token IDs, never hidden states, so the table can sit in host memory and be prefetched while earlier layers compute. An expert is learned computation picked by a router. This is learned memory with deterministic addressing. The lineage runs through Gemma 3n's per-layer embeddings at Google DeepMind, ByteDance Seed's Over-Tokenized Transformer in 2025, DeepSeek's Engram by Cheng and colleagues in 2026, and Meituan LongCat's "Scaling Embeddings Outperforms Scaling Experts" by Liu and colleagues in 2026. Those last two pull apart. Engram reports a U-shaped optimum when a fixed sparse budget is split between experts and n-grams, while LongCat argues embeddings are the better home for parameters. Park that fork. 11 00:04:54,350 --> 00:04:57,900 [Hal Turing] Parked. Last ingredient: the optimizer. 12 00:04:57,900 --> 00:05:55,225 [Dr. Ada Shannon] Muon, from Keller Jordan and collaborators in 2024, orthogonalizes each 2D weight matrix's momentum with Newton-Schulz iterations, so the update is shape-aware in a way AdamW isn't. AdamW stays on embeddings, head and router. The paper claims architecture and optimizer together shift the optimal learning rate and batch size upward. So: five components, GDN hybrid, QSA, Gated Residual, the n-gram layer and Muon, and one recurring question: do loss, accuracy, cost and stability agree? The paper flags three places where loss and benchmarks split. First, the evidence standard. Most ablations are single runs at moderate scale, from 25B-A3B up to a 48-layer 156B-A7B, with no reported seeds. Two identical runs give slightly different accuracy, and that gap is the noise floor. Section 3.2 says single-evaluation margins among the top settings "likely fall within standard evaluation noise." 13 00:05:55,225 --> 00:06:14,825 [Hal Turing] Credit where it's due: writing that sentence into your own report is honest, and it hands us the ruler. We'll hold the paper to it. Receipts, then. Hybrid first. A fixed-size state sounds like a whiteboard that fills up. How does Gated DeltaNet avoid stacking every repeated key on top of the last one? 14 00:06:14,825 --> 00:07:12,100 [Dr. Ada Shannon] It edits instead of appending. Picture a whiteboard of key-to-value notes. A decay, alpha, erases some of the whole board each step. Then the delta term reads what's already written under the current key, subtracts it, and writes only the error. Same key twice, one corrected note. That's the gated delta rule. Qwen adds a sigmoid output gate, zero-centered RMSNorm and short causal convolutions. In Table 1, 25B-A3B MoEs on 400B tokens, the GDN hybrid beats full attention on eight of nine benchmarks and the sliding-window hybrid on seven. Averages are 53.81, 49.87 and 51.15. MATH is 53.98 against 49.40, MultiPL-E 47.48 against 39.73. The paper says outright it doesn't isolate the cause of each gain. Also, RoPE and NoPE tie in pretraining, but NoPE gives far more endless generation after post-training. FlashQLA kernels run two to three times faster forward. 15 00:07:12,100 --> 00:07:18,650 [Hal Turing] Now the skim. Pooling keys to make the indexer cheap, doesn't that smear positions together? 16 00:07:18,650 --> 00:08:21,775 [Dr. Ada Shannon] Not if you pool before rotation. The indexer is MQA with four query heads and one shared key head. Keys are average-pooled in blocks of four before RoPE, so nothing averages across rotary phases. ReLU block-causal scores feed a top-k of 2048 tokens, 512 blocks, plus the incomplete tail block. Indexer cost falls from n squared to n squared over r. Training has two stages: about 2B tokens distilling the indexer alone against max-pooled teacher attention, then roughly 200B tokens of joint training at 256K. Dense-initialized sparse attention drops sharply before that joint phase recovers it. Loss gap to full attention is about 1e-4. Table 2 average goes 75.9 to 76.8, MRCR at 512K goes 30.66 to 40.53. QSA matches full-attention RULER at relative indexer latency 0.25, while IndexShare stays below at 0.5. Kernel-level speedups at 1M are 7.6x prefill and 4.9x decode. 17 00:08:21,775 --> 00:08:27,775 [Hal Turing] Now the four mailboxes. Before any clever operators, how much is just width? 18 00:08:27,775 --> 00:08:51,850 [Dr. Ada Shannon] Static scalars alone cut loss about 0.01. Then Table 5, 560B tokens: pre-norm 1.617 loss and 50.91 average; mHC static 1.596 and 52.49; mHC dynamic 1.594 and 54.47; Gated Residual 1.590 and 54.66. Going from static to dynamic moves loss by— 19 00:08:51,850 --> 00:09:02,774 [Hal Turing] Hold on, 0.002? For almost two points of accuracy? Every ablation I trust says a loss gap that small is a rounding error downstream. 20 00:09:02,774 --> 00:09:45,125 [Dr. Ada Shannon] That's the paper's point: the ratio reverses. Baseline to static was 0.021 loss for 1.58 points. Static to dynamic is 0.002 for 1.98, MATH 55.08 to 59.54. Lessons: sigmoid beats tanh, per-channel read matters, write granularity doesn't, and the branch-mixing operator adds little, so it's dropped. GR is per-branch RMSNorm plus an elementwise sigmoid gate on the read, one scalar per branch on the write, replacing pre-norm. Against dynamic mHC the paper says 'comparable'; GR's edge is less memory traffic. Sparse top-2 branch reads looked free in pretraining and degraded after post-training, so they were dropped. 21 00:09:45,125 --> 00:09:52,200 [Hal Turing] Genuine question before the phone book tables: why put the lookup at layer 2 and not layer 0? 22 00:09:52,200 --> 00:11:08,325 [Dr. Ada Shannon] So the host can prefetch while layer 1 computes. One layer, multi-head hashing, contextual gate, 300 tokens per active parameter. Table 7: every single-layer placement gets loss 1.541 to 1.544 against 1.585 without, and multi-layer adds nothing consistent. Table 9 adds parameters, vocabulary 20x to 200x of the 250K base: loss 1.585, 1.553, 1.541, 1.534, 1.526. Accuracy jumps from none to 20x, GSM8K 59.21 to 65.09, MATH 32.52 to 37.38. Then it stays within about 1.5 points across 50x to 200x, while C-Eval and CMMLU keep rising. The paper's words: 'saturates or fluctuates'. Table 8 fixes the total budget by removing experts: loss 1.202, 1.200, 1.197, 1.201 at none, 5x, 10x, 30x, and MMLU-Pro 44.38, 44.49, 44.66, 42.61. The paper concludes n-grams and experts play distinct roles. The excerpt never says which row the shipped 51B matches. 23 00:11:08,325 --> 00:11:11,050 [Hal Turing] And Muon. What doesn't it touch? 24 00:11:11,050 --> 00:12:14,825 [Dr. Ada Shannon] The gates stay on AdamW along with the router, embeddings and output head, because Muon destabilised early router training. Everything else that's a 2D linear map gets Muon, with eight Newton-Schulz steps chosen for fewer grad-norm spikes, and fused matrices are split first. Canzona repartitions whole parameters by estimated orthogonalization cost across data-parallel ranks. The refit scaling law predicts a bigger batch: 25.2M gives loss 1.5702 against 1.5774 for the old 12.6M, and batch warmup costs 18.8% more steps. At 4x optimal learning rate, AdamW hits 183 spikes per 10k steps while Muon never crosses the clip threshold. The single-variable gate ablation cuts spikes from 32.0 to 3.2. Table 11: Flash-Next wins all 14 against the 27B base. Against the 397B it gets MMLU-Pro 73.23 versus 70.90 but MATH 72.78 versus 74.38, and SWEBench-Pretrain is a proxy the paper built itself. 25 00:12:14,825 --> 00:12:42,150 [Hal Turing] Ada, time to cash the cold-open check. Table 9, rows 50x to 200x: loss keeps falling, 1.541, 1.534, 1.526, but MMLU reads 64.71, 64.70, 64.85 and GSM8K reads 64.00, 63.08, 62.96. What would you need to see to tell saturation from fluctuation? 26 00:12:42,150 --> 00:13:38,275 [Dr. Ada Shannon] Seeds. The paper never says how much accuracy moves between two identical runs, so the honest verdict is 'not established', which is not 'refuted'. Table 7 gives a proxy. The same n-gram parameters at different layers land at near-identical loss, yet MMLU spans 63.20 to 65.07, GSM8K 61.26 to 65.73, and the nine-benchmark average 46.62 to 47.94. That's spread at equal loss, an upper bound rather than a seed study, and it's as large as the 50x-to-200x differences. Section 3.2 invokes evaluation noise to refuse to rank Table 10's top settings, then Table 9 gets no such treatment. Loss does have a scale: Figure 9 calls 7e-4 near noise, and Table 9's steps are 0.007 to 0.008. Different experiment, different scale, but loss falling is credible. Accuracy flatness is unmeasured. 27 00:13:38,275 --> 00:14:20,475 [Hal Turing] I averaged the nine columns myself, so this is our arithmetic, not the paper's: about 50.15 with no table, 53.33 at 20x, then 53.57, 53.52, 53.30. Nearly all the gain is the first step. But a flat line and a noisy flat line look the same. Here's my genuine question. C-Eval goes 72.12, 73.75, 74.94 and CMMLU creeps up too, while English reasoning sits still. Real Chinese knowledge, or a big hashed table memorizing multi-character entities and templated exam formats? 28 00:14:20,475 --> 00:15:03,675 [Dr. Ada Shannon] Nobody can say, because there's no contamination or n-gram-overlap analysis. A hashed table is the component best placed to turn leakage into accuracy. Carlini and colleagues at Google Brain showed in 2023, in Quantifying Memorization Across Neural Language Models, that memorization grows with capacity and duplication. A phone book that contains the answer key is still a phone book. There are tamer readings. English reasoning saturates while Chinese recall keeps climbing, which per-language breakdowns could test. Or 300 tokens per active parameter is fixed, so bigger tables get fewer updates per slot, and saturation belongs to this recipe. Also, Uncheatable perplexity appears for Table 8 but not Table 9, where loss drops most. 29 00:15:03,675 --> 00:15:22,525 [Hal Turing] Hold on, Table 8 is the one I can't square with the headline. Experts were removed to pay for n-gram slots, and MMMLU falls from 58.32 to 56.22 by 30x. Yet LongCat says embeddings are the better place for parameters. So who's right? 30 00:15:22,525 --> 00:15:58,825 [Dr. Ada Shannon] Neither is established. The defensible claim is narrow: n-gram parameters are additive capacity if you keep the experts and can host tables off-accelerator. They aren't a swap for experts. The two tables also use different backbones, baseline loss 1.202 versus 1.585, and Table 9's 50x row is Table 7's layer-2 row, the same run reused. The excerpt doesn't say whether FLOPs were held equal when experts were removed. Qwen's loss curve echoes Engram's U-shape without any downstream benefit, and there's no head-to-head. Recipe and baseline strength differ, so the disagreement is unresolved, not settled. 31 00:15:58,825 --> 00:16:14,775 [Hal Turing] Does the same worry hit QSA? It beats the dense model it replaces at 512K with a 1e-4 loss gap, trained at 256K. And the stress tests favor Muon so heavily. How much do I believe those? 32 00:16:14,775 --> 00:16:57,200 [Dr. Ada Shannon] Sparse beating dense by ten points is regularisation, noise in a high-variance eval, or unequal training history, and the excerpt doesn't say CPT was matched. The stress test is one run per arm at the same absolute learning rate for both optimizers, so 'x optimal' is ambiguous. The cleanest evidence is the single-variable gate pair, 32.0 down to 3.2 spikes. The 7.6x and 4.9x are kernel-level at batch 4, and the ninth is arithmetic rather than GPU-hours. What holds up is the three-axis discipline. Static-to-dynamic residuals, sparse top-2 reads breaking after post-training, NoPE: loss alone misled each time. The n-gram layer never got that late-stage check. 33 00:16:57,200 --> 00:17:02,600 [Hal Turing] So what does a practitioner do differently, and what should the next report contain? 34 00:17:02,600 --> 00:17:42,425 [Dr. Ada Shannon] With host memory and a prefetch path, a single n-gram layer is a cheap gain, but past 20x to 50x isn't shown to help English reasoning. Don't pay for tables by cutting experts. If you widen the residual stream, spend on the read, not branch mixing. Muon wants split fused matrices. And measure run-to-run spread on your own harness before believing a smaller difference. For the authors: seed replicates of Table 9, matched updates per slot, Table 8 rerun on the Table 9 backbone, real prefetch numbers, since decode addresses depend on the token just sampled, and post-training runs. Their stated bottleneck is evaluation throughput, and a cheaper probe only helps if its variance is known. 35 00:17:42,425 --> 00:18:05,050 [Hal Turing] So, the cold open. Loss falling is well supported. Accuracy saturation is plausible but unproven, and the strongest claim the data supports is that n-gram tables are extra capacity, not a swap for experts. The habit to keep: ask what the run-to-run variance is before believing any change smaller than it. Thanks for listening, and goodbye.