← All episodes Qwen3.8-Next Design: Hybrid Attention, Residuals, and N-gram Embeddings

Qwen3.8-Next Design: Hybrid Attention, Residuals, and N-gram Embeddings

Sep 29, 2026
This episode examines the Qwen team's design-rationale report for Qwen3.8-Flash-Next, a 125B-total, 6B-activated base model with 51B of n-gram embeddings held in host memory. It walks through the five components: a hybrid of Gated DeltaNet layers with periodic full attention, Qwen's sparse attention variant (QSA) as a response to the quadratic cost of DeepSeek-style indexers, a four-branch gated residual stream, a hashed n-gram lookup layer, and the Muon optimizer. The recurring question is whether loss, benchmark accuracy, cost, and stability agree. The paper reports that larger n-gram vocabularies always lower loss while accuracy stays flat. The hosts also note that most ablations are single runs with no reported seeds, and that the paper itself says top-setting margins likely fall within evaluation noise. The headline claim, matching or beating a 397B-A17B predecessor on eight of fourteen benchmarks at roughly a ninth of the training compute, is set against the ablation evidence behind it. The episode is useful for anyone who wants to see how the design choices trace to prior work (highway networks, Hyper-Connections, Engram, LongCat) and how much weight those choices can bear.
Sources:
1. On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability — Zihan Qiu, Zekun Wang, Xiao Li, Yanpeng Li, Yang Xu, Yixuan Wang, Huaqing Zhang, Rui Men, Bochao Mao, Chengruidong Zhang, Fan Zhou, Hao Luo, Haofeng Huang, Haoran Lian, Haoyan Huang, Hongqing Chen, Jianwei Zhang, Jing Xu, Junjie Wang, Langshi Chen, Liangyu Wang, Linlang Jiang, Man Yuan, Minmin Sun, Peng Jin, Siqi Zhang, Siyu Wang, Xingzhang Ren, Yakai Wang, Yi Zhang, Yiming Dong, Yizhong Cao, Yubo Ma, Yunfei Mao, Bo Zheng, Dayiheng Liu, 2026
http://arxiv.org/abs/2608.30320v1
2. Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models (Engram) — Xin Cheng et al. (DeepSeek-AI, with Peking University), 2026
https://scholar.google.com/scholar?q=Conditional+Memory+via+Scalable+Lookup%3A+A+New+Axis+of+Sparsity+for+Large+Language+Models+%28Engram%29
3. Scaling Embedding Layers in Language Models (SCONE) — Da Yu, Edith Cohen, Badih Ghazi, Yangsibo Huang, Pritish Kamath, Ravi Kumar, Daogao Liu, Chiyuan Zhang (Google), 2025
https://scholar.google.com/scholar?q=Scaling+Embedding+Layers+in+Language+Models+%28SCONE%29
4. Over-Tokenized Transformer: Vocabulary is Generally Worth Scaling — Hongzhi Huang, Defa Zhu, Banggu Wu, Yutao Zeng, Ya Wang, Qiyang Min, Xun Zhou (ByteDance Seed), 2025
https://scholar.google.com/scholar?q=Over-Tokenized+Transformer%3A+Vocabulary+is+Generally+Worth+Scaling
5. Memory Layers at Scale — Vincent-Pierre Berges, Barlas Oguz, Daniel Haziza, Wen-tau Yih, Luke Zettlemoyer, Gargi Ghosh (Meta FAIR), 2024
https://scholar.google.com/scholar?q=Memory+Layers+at+Scale
6. Scaling Embeddings Outperforms Scaling Experts in Language Models — Hong Liu, Jiaqi Zhang, Chao Wang, ... Xunliang Cai (Meituan LongCat), 2026
https://scholar.google.com/scholar?q=Scaling+Embeddings+Outperforms+Scaling+Experts+in+Language+Models
7. N-Grammer: Augmenting Transformers with Latent n-grams — Aurko Roy, Rohan Anil, Guangda Lai, et al., 2022
https://scholar.google.com/scholar?q=N-Grammer%3A+Augmenting+Transformers+with+Latent+n-grams
8. Scaling Embedding Layers in Language Models — Da Yu, Edith Cohen, Badih Ghazi, Yangsibo Huang, Pritish Kamath, Ravi Kumar, Daogao Liu, Chiyuan Zhang, 2025
https://scholar.google.com/scholar?q=Scaling+Embedding+Layers+in+Language+Models
9. L3: Large Lookup Layers — Albert Tseng, Christopher De Sa, 2026
https://scholar.google.com/scholar?q=L3%3A+Large+Lookup+Layers
10. STEM: Scaling Transformers with Embedding Modules — Ranajoy Sadhukhan, Sheng Cao, Harry Dong, Changsheng Zhao, Attiano Purpura-Pontoniere, Yuandong Tian, Zechun Liu, Beidi Chen, 2026
https://scholar.google.com/scholar?q=STEM%3A+Scaling+Transformers+with+Embedding+Modules
11. Large Memory Layers with Product Keys — Guillaume Lample, Alexandre Sablayrolles, Marc'Aurelio Ranzato, Ludovic Denoyer, Hervé Jégou, 2019
https://scholar.google.com/scholar?q=Large+Memory+Layers+with+Product+Keys
12. Mixture of A Million Experts (PEER) — Xu Owen He, 2024
https://scholar.google.com/scholar?q=Mixture+of+A+Million+Experts+%28PEER%29
13. Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws — Zeyuan Allen-Zhu, Yuanzhi Li, 2024
https://scholar.google.com/scholar?q=Physics+of+Language+Models%3A+Part+3.3%2C+Knowledge+Capacity+Scaling+Laws
14. Quantifying Memorization Across Neural Language Models — Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramèr, Chiyuan Zhang, 2023
https://scholar.google.com/scholar?q=Quantifying+Memorization+Across+Neural+Language+Models
15. Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention — Jingyang Yuan et al. (DeepSeek), 2025
https://scholar.google.com/scholar?q=Native+Sparse+Attention%3A+Hardware-Aligned+and+Natively+Trainable+Sparse+Attention
16. Kimi Linear: An Expressive, Efficient Attention Architecture — Kimi Team, 2025
https://scholar.google.com/scholar?q=Kimi+Linear%3A+An+Expressive%2C+Efficient+Attention+Architecture
17. Small-scale Proxies for Large-scale Transformer Training Instabilities — Mitchell Wortsman et al., 2023
https://scholar.google.com/scholar?q=Small-scale+Proxies+for+Large-scale+Transformer+Training+Instabilities
Interactive Visualization: Qwen3.8-Next Design: Hybrid Attention, Residuals, and N-gram Embeddings