This episode examines LinearKV, a new approach to position-independent caching for hybrid Mamba-attention language models, and a striking result: the mathematically "correct" way to merge cached context — exact algebraic composition of recurrent states — badly degrades one tested model's output quality, while a simpler shortcut using only the most recent cached chunk performs reliably. The discussion contrasts standard prefix caching (used by vLLM's PagedAttention and SGLang's RadixAttention) with position-independent caching, which lets cached chunks be reused regardless of order, and explains why that trick breaks down for linear-recurrent layers like Mamba-2 and Gated DeltaNet, which compress history into a single fixed-size state rather than a token-indexed KV ledger. It traces the core problem to how each cached chunk's recurrent state was built in isolation, making "exact" composition exact only relative to a flawed reference rather than the true full-context computation. Listeners interested in LLM serving infrastructure, caching systems, or the tradeoffs of hybrid architectures will find this a concrete case study in how systems intuitions from full attention can actively mislead when applied to newer recurrent designs.

Sources:
1. LinearKV: One Cached State Suffices for Position-Independent Caching in Hybrid LLMs — Yirui Liu, Ruoling Qi, Longwen Wang, Xuaner Wu, Jian Chen, Yuxin Jin, Jiawei Shao, Xuelong Li, 2026
http://arxiv.org/abs/2608.11231
2. Efficient Memory Management for Large Language Model Serving with PagedAttention — Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, Ion Stoica, 2023
https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention
3. SGLang: Efficient Execution of Structured Language Model Programs — Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun, Jeff Huang, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonzalez, Clark Barrett, Ying Sheng, 2024
https://scholar.google.com/scholar?q=SGLang%3A+Efficient+Execution+of+Structured+Language+Model+Programs
4. CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowledge Fusion — Jiayi Yao, Hanchen Li, Yuhan Liu, Siddhant Ray, Yihua Cheng, Qizheng Zhang, Kuntai Du, Shan Lu, Junchen Jiang, 2025
https://scholar.google.com/scholar?q=CacheBlend%3A+Fast+Large+Language+Model+Serving+for+RAG+with+Cached+Knowledge+Fusion
5. EPIC: Efficient Position-Independent Context Caching for Serving Large Language Models — Junhao Hu et al., 2025
https://scholar.google.com/scholar?q=EPIC%3A+Efficient+Position-Independent+Context+Caching+for+Serving+Large+Language+Models
6. HYPIC: Accelerating hybrid-attention LLM serving with position-independent caching — Yifei Liu, Juntong Wu, Yang Liu, Junhao Hu, Minghao Li, Xiaoxu Chen, Weihang Chen, 2026
https://scholar.google.com/scholar?q=HYPIC%3A+Accelerating+hybrid-attention+LLM+serving+with+position-independent+caching
7. Marconi: Prefix caching for the era of hybrid LLMs — Rui Pan, Zhuang Wang, Zhen Jia, Can Karakus, Luca Zancato, Tri Dao, Yida Wang, Ravi Netravali, 2025
https://scholar.google.com/scholar?q=Marconi%3A+Prefix+caching+for+the+era+of+hybrid+LLMs
8. Gated Delta Networks: Improving Mamba2 with the Delta Rule — Songlin Yang, Jan Kautz, Ali Hatamizadeh, 2025 (ICLR)
https://scholar.google.com/scholar?q=Gated+Delta+Networks%3A+Improving+Mamba2+with+the+Delta+Rule
9. EPIC: Efficient position-independent caching for serving large language models — Junhao Hu, Wenrui Huang, Weidong Wang, et al., 2025 (ICML)
https://scholar.google.com/scholar?q=EPIC%3A+Efficient+position-independent+caching+for+serving+large+language+models
Interactive Visualization: LinearKV: When Exact State Merging Breaks Hybrid Models

This episode examines FLINT, a proposed hardware/software architecture from Huawei's Zurich research lab (with ETH Zürich and HUST) for closing the gap between multi-terabyte LLM weight sizes and the far smaller on-package memory of today's GPUs. It introduces high bandwidth flash (HBF), an emerging memory tier that stacks 3D NAND dies with through-silicon vias to sit directly beside HBM in the accelerator package, storing read-only model weights while HBM handles fast-changing KV cache and activations. The discussion walks through core NAND flash mechanics — dies, planes, blocks, and pages, along with the punishing asymmetry between microsecond reads and millisecond erases — to explain why naive flash designs stall under refresh operations and static prefetching. It then details how FLINT's burst-buffer controller replaces compiler-driven prefetch hints with real-time demand-based read coalescing, using HBF's built-in page and cache buffers instead of dedicated SRAM. Listeners interested in memory system design, inference hardware economics, and the practical engineering trade-offs of scaling capacity without wasting GPU compute will find this a detailed look at a genuinely emerging technology rather than a shipping product.

Sources:
1. FLINT: Efficiently Leveraging High Bandwidth Flash for Capacity-Scalable LLM Inference Acceleration — Geraldo F. Oliveira, Arash Tavakkol, Xiangyu Zhu, Ahmet Caner Yüzügüler, Vamanan Arulchelvan, Lukas Cavigelli, Renzo Andri, Mohammad Sadrosadati, Jia Xinglei, Onur Mutlu, Zhou Ke, Shai Bergman, Ji Zhang, 2026
http://arxiv.org/abs/2608.25062
2. ZeRO-Infinity: Breaking the GPU Memory Wall for Extreme Scale Deep Learning — Samyam Rajbhandari, Olatunji Ruwase, Jeff Rasley, Shaden Smith, Yuxiong He (Microsoft), 2021
https://scholar.google.com/scholar?q=ZeRO-Infinity%3A+Breaking+the+GPU+Memory+Wall+for+Extreme+Scale+Deep+Learning
3. FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU — Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, et al., 2023
https://scholar.google.com/scholar?q=FlexGen%3A+High-Throughput+Generative+Inference+of+Large+Language+Models+with+a+Single+GPU
4. Smart-Infinity: Fast Large Language Model Training using Near-Storage Processing on a Real System — Hongsun Jang, Jaeyong Song, Jaewon Jung, Jaeyoung Park, Youngsok Kim, Jinho Lee, 2024, HPCA
https://scholar.google.com/scholar?q=Smart-Infinity%3A+Fast+Large+Language+Model+Training+using+Near-Storage+Processing+on+a+Real+System
5. InstInfer: In-Storage Attention Offloading for Cost-Effective Long-Context LLM Inference — Xiurui Pan, Endian Li, Qiao Li, Shengwen Liang, Yizhou Shan, Ke Zhou, Yingwei Luo, Xiaolin Wang, Jie Zhang, 2024
https://scholar.google.com/scholar?q=InstInfer%3A+In-Storage+Attention+Offloading+for+Cost-Effective+Long-Context+LLM+Inference
6. DFTL: A Flash Translation Layer Employing Demand-based Selective Caching of Page-level Address Mappings — Aayush Gupta, Youngjae Kim, Bhuvan Urgaonkar, 2009, ASPLOS
https://scholar.google.com/scholar?q=DFTL%3A+A+Flash+Translation+Layer+Employing+Demand-based+Selective+Caching+of+Page-level+Address+Mappings
7. Design Tradeoffs for SSD Performance — Nitin Agrawal, Vijayan Prabhakaran, Ted Wobber, John D. Davis, Mark Manasse, Rina Panigrahy, 2008, USENIX ATC
https://scholar.google.com/scholar?q=Design+Tradeoffs+for+SSD+Performance
8. ZNS: Avoiding the Block Interface Tax for Flash-based SSDs — Matias Bjørling, Abutalib Aghayev, Hans Holmberg, Aravind Ramesh, Damien Le Moal, Gregory R. Ganger, George Amvrosiadis, 2021, USENIX ATC
https://scholar.google.com/scholar?q=ZNS%3A+Avoiding+the+Block+Interface+Tax+for+Flash-based+SSDs
9. RAIDR: Retention-Aware Intelligent DRAM Refresh — Jamie Liu, Ben Jaiyen, Richard Veras, Onur Mutlu, 2012, ISCA
https://scholar.google.com/scholar?q=RAIDR%3A+Retention-Aware+Intelligent+DRAM+Refresh
10. Threshold Voltage Distribution in MLC NAND Flash Memory: Characterization, Analysis, and Modeling — Yu Cai, Erich F. Haratsch, Onur Mutlu, Ken Mai, 2013, DATE
https://scholar.google.com/scholar?q=Threshold+Voltage+Distribution+in+MLC+NAND+Flash+Memory%3A+Characterization%2C+Analysis%2C+and+Modeling
11. Error Analysis and Retention-Aware Error Management for NAND Flash Memory — Yu Cai, Erich F. Haratsch, Onur Mutlu, Ken Mai, 2013, Intel Technology Journal
https://scholar.google.com/scholar?q=Error+Analysis+and+Retention-Aware+Error+Management+for+NAND+Flash+Memory
12. H3: Hybrid Architecture Using High Bandwidth Memory and High Bandwidth Flash for Cost-Efficient LLM Inference — M. Ha, E. Kim, H. Kim, 2026
https://scholar.google.com/scholar?q=H3%3A+Hybrid+Architecture+Using+High+Bandwidth+Memory+and+High+Bandwidth+Flash+for+Cost-Efficient+LLM+Inference
13. LLM in a Flash: Efficient Large Language Model Inference with Limited Memory — K. Alizadeh, S. I. Mirzadeh, et al., 2024
https://scholar.google.com/scholar?q=LLM+in+a+Flash%3A+Efficient+Large+Language+Model+Inference+with+Limited+Memory
14. PRESERVE: Prefetching Model Weights and KV-Cache in Distributed LLM Serving — A. C. Yüzügüler, J. Zhuang, L. Cavigelli, 2025
https://scholar.google.com/scholar?q=PRESERVE%3A+Prefetching+Model+Weights+and+KV-Cache+in+Distributed+LLM+Serving
15. MemExplorer: Navigating the Heterogeneous Memory Design Space for Agentic Inference NPUs — H. Wu, Z. Cao, et al., 2026
https://scholar.google.com/scholar?q=MemExplorer%3A+Navigating+the+Heterogeneous+Memory+Design+Space+for+Agentic+Inference+NPUs
16. Lincoln: Real-Time 50~100B LLM Inference on Consumer Devices with LPDDR-Interfaced, Compute-Enabled Flash Memory — W. Sun, M. Gao, et al., 2025
https://scholar.google.com/scholar?q=Lincoln%3A+Real-Time+50~100B+LLM+Inference+on+Consumer+Devices+with+LPDDR-Interfaced%2C+Compute-Enabled+Flash+Memory
Interactive Visualization: FLINT: Sharing High Bandwidth Flash for Scalable LLM Inference

This episode examines "Decoupling KL and Trajectories," which challenges the assumption that off-policy training must pair with forward KL divergence and on-policy training with reverse KL — a coupling used by DeepSeek-R1, Qwen3, MiMo, and GLM-5 without ever being tested. The hosts unpack the two independent design axes at play: prefix source (teacher-generated versus student-generated rollouts) and KL direction (forward's distribution-covering behavior versus reverse's mode-seeking concentration), showing that standard SFT is actually just forward-KL distillation on teacher text in disguise. The paper's real contribution is exploring two previously unstudied combinations — teacher-prefix reverse KL and student-prefix forward KL — turning an assumed two-option choice into a full four-quadrant design space. Grounding the discussion is a concrete experimental setup using Qwen3-0.6B-Base as a student distilling from Qwen3-4B and Qwen3-8B teachers. Listeners interested in LLM distillation, on-policy training, or exposure bias will find this a sharp dismantling of an industry-wide convention nobody had thought to question.

Sources:
1. Decoupling KL and Trajectories: A Unified Perspective for SFT, DAgger, Offline RL, and OPD in LLM Distillation — Anhao Zhao, Haoran Xin, Yingqi Fan, Junlong Tong, Wenjie Li, Xiaoyu Shen, 2026
http://arxiv.org/abs/2605.16826
2. A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning — Stéphane Ross, Geoffrey Gordon, J. Andrew Bagnell, 2011
https://scholar.google.com/scholar?q=A+Reduction+of+Imitation+Learning+and+Structured+Prediction+to+No-Regret+Online+Learning
3. Scheduled Sampling for Sequence Prediction with Recurrent Neural Networks — Samy Bengio, Oriol Vinyals, Navdeep Jaitly, Noam Shazeer, 2015
https://scholar.google.com/scholar?q=Scheduled+Sampling+for+Sequence+Prediction+with+Recurrent+Neural+Networks
4. Sequence Level Training with Recurrent Neural Networks — Marc'Aurelio Ranzato, Sumit Chopra, Michael Auli, Wojciech Zaremba, 2016
https://scholar.google.com/scholar?q=Sequence+Level+Training+with+Recurrent+Neural+Networks
5. On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes — Rishabh Agarwal, Nino Vieillard, Yongchao Zhou, Piotr Stanczyk, Sabela Ramos, Matthieu Geist, Olivier Bachem (Google DeepMind), 2024
https://scholar.google.com/scholar?q=On-Policy+Distillation+of+Language+Models%3A+Learning+from+Self-Generated+Mistakes
6. MiniLLM: Knowledge Distillation of Large Language Models — Yuxian Gu, Li Dong, Furu Wei, Minlie Huang, 2024
https://scholar.google.com/scholar?q=MiniLLM%3A+Knowledge+Distillation+of+Large+Language+Models
7. Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models — Feng Luo, Yu-Neng Chuang, Guanchu Wang, Zicheng Xu, Xiaotian Han, Tianyi Zhang, Vladimir Braverman, 2026
https://scholar.google.com/scholar?q=Demystifying+OPD%3A+Length+Inflation+and+Stabilization+Strategies+for+Large+Language+Models
8. A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning (DAgger) — Stéphane Ross, Geoffrey J. Gordon, J. Andrew Bagnell, 2011
https://scholar.google.com/scholar?q=A+Reduction+of+Imitation+Learning+and+Structured+Prediction+to+No-Regret+Online+Learning+%28DAgger%29
9. Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting — Xinyuan Chen, Noam Razin, Karthik Narasimhan, Danqi Chen, 2025
https://scholar.google.com/scholar?q=Retaining+by+Doing%3A+The+Role+of+On-Policy+Data+in+Mitigating+Forgetting
Interactive Visualization: Decoupling KL Direction from Rollout Source in LLM Distillation

This episode examines "Not All Disagreement Is Learnable: Token Teachability in On-Policy Distillation," which challenges a core assumption in on-policy knowledge distillation: that raw KL divergence between teacher and student token predictions is a reliable signal for which tokens deserve training focus. The discussion traces the lineage from Hinton's original distillation work through Google DeepMind's on-policy approach, then explains the paper's key insight — large disagreement can mean either a small, actionable correction the student can use, or a "incompatible" mismatch pointing toward options the student assigns near-zero probability, and raw KL can't distinguish the two. Building on this distinction, the authors introduce "token teachability" as a better selection criterion and a method called TA-OPD that trains only on the most teachable tokens, reportedly matching or beating full-dataset training while using just 5% of the tokens. Listeners interested in efficient model training, distillation techniques, or the gap between statistical salience and actual learnability will find the reframing of a decade-old assumption particularly compelling.

Sources:
1. Not All Disagreement Is Learnable: Token Teachability in On-Policy Distillation — Yuanyi Wang, Su Lu, Yanggan Gu, Pengkai Wang, Yifan Yang, Zhaoyi Yan, Congkai Xie, Jianmin Wu, Hongxia Yang, 2026
http://arxiv.org/abs/2605.26844
2. Distilling the Knowledge in a Neural Network — Geoffrey Hinton, Oriol Vinyals, Jeff Dean, 2015
https://scholar.google.com/scholar?q=Distilling+the+Knowledge+in+a+Neural+Network
3. Sequence-Level Knowledge Distillation — Yoon Kim, Alexander M. Rush, 2016
https://scholar.google.com/scholar?q=Sequence-Level+Knowledge+Distillation
4. On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes — Rishabh Agarwal, Nino Vieillard, et al. (Google DeepMind), 2024
https://scholar.google.com/scholar?q=On-Policy+Distillation+of+Language+Models%3A+Learning+from+Self-Generated+Mistakes
5. MiniLLM: Knowledge Distillation of Large Language Models — Yuxian Gu, Li Dong, Furu Wei, Minlie Huang, 2024
https://scholar.google.com/scholar?q=MiniLLM%3A+Knowledge+Distillation+of+Large+Language+Models
6. A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning (DAgger) — Stéphane Ross, Geoffrey J. Gordon, J. Andrew Bagnell, 2011
https://scholar.google.com/scholar?q=A+Reduction+of+Imitation+Learning+and+Structured+Prediction+to+No-Regret+Online+Learning+%28DAgger%29
7. Not All Tokens Are What You Need for Pretraining (Rho-1) — Zhenghao Lin, Zhibin Gou, Yeyun Gong, Xiao Liu, Yelong Shen, Ruochen Xu, Chen Lin, Yujiu Yang, Jian Jiao, Nan Duan, Weizhu Chen, 2024
https://scholar.google.com/scholar?q=Not+All+Tokens+Are+What+You+Need+for+Pretraining+%28Rho-1%29
8. Contrastive Decoding: Open-ended Text Generation as Optimization — Xiang Lisa Li, Ari Holtzman, Daniel Fried, Percy Liang, Jason Eisner, Tatsunori Hashimoto, Luke Zettlemoyer, Mike Lewis, 2023
https://scholar.google.com/scholar?q=Contrastive+Decoding%3A+Open-ended+Text+Generation+as+Optimization
9. DistiLLM: Towards Streamlined Distillation for Large Language Models — Jongwoo Ko, Sungnyun Kim, Tianyi Chen, Se-Young Yun, 2024
https://scholar.google.com/scholar?q=DistiLLM%3A+Towards+Streamlined+Distillation+for+Large+Language+Models
10. Fast Inference from Transformers via Speculative Decoding — Yaniv Leviathan, Matan Kalman, Yossi Matias, 2023
https://scholar.google.com/scholar?q=Fast+Inference+from+Transformers+via+Speculative+Decoding
11. TIP: Token Importance in On-Policy Distillation — Yuanda Xu, Hejian Sang, Zhengze Zhou, Ran He, Zhipeng Wang, Alborz Geramifard, 2026
https://scholar.google.com/scholar?q=TIP%3A+Token+Importance+in+On-Policy+Distillation
12. Entropy-Aware On-Policy Distillation of Language Models — Woogyeol Jin, Taywon Min, Yongjin Yang, Swanand Ravindra Kadhe, Yi Zhou, Dennis Wei, Nathalie Baracaldo, Kimin Lee, 2026
https://scholar.google.com/scholar?q=Entropy-Aware+On-Policy+Distillation+of+Language+Models
13. Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe — Yaxuan Li, Yuxin Zuo, Bingxiang He, Jinqian Zhang, Chaojun Xiao, Cheng Qian, et al., 2026
https://scholar.google.com/scholar?q=Rethinking+On-Policy+Distillation+of+Large+Language+Models%3A+Phenomenology%2C+Mechanism%2C+and+Recipe
14. Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning — Shenzhi Wang, Le Yu, Chang Gao, Chujie Zheng, Shixuan Liu, Rui Lu, et al., 2026
https://scholar.google.com/scholar?q=Beyond+the+80%2F20+Rule%3A+High-Entropy+Minority+Tokens+Drive+Effective+Reinforcement+Learning+for+LLM+Reasoning
15. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning — Daya Guo, Dejian Yang, Haowei Zhang, et al., 2025
https://scholar.google.com/scholar?q=DeepSeek-R1%3A+Incentivizing+Reasoning+Capability+in+LLMs+via+Reinforcement+Learning
Interactive Visualization: Token Teachability: Rethinking Disagreement in On-Policy Distillation

This episode examines Weak-to-Strong On-Policy Distillation, a paper from University of Maryland, Microsoft Research, and MBZUAI showing that an 8-billion-parameter student model can outperform every teacher used to train it on math benchmarks. The discussion traces the method's lineage from Hinton's original 2015 knowledge distillation through DAgger's 2011 on-policy correction idea to 2023 weak-to-strong generalization work from OpenAI's Superalignment team, explaining how the paper inverts DAgger's core assumption by using a supervisor weaker than the student rather than a trusted expert. It also contrasts this approach with reinforcement learning from verifiable rewards, which gives only a sparse end-of-rollout signal, versus on-policy distillation's dense per-token feedback on the student's own generated trajectories. The hosts situate the work against industry precedents like Qwen3's large-to-small distillation and multi-teacher on-policy distillation, both of which still require a teacher at least as capable as the student — a constraint this paper's method aims to eliminate. Listeners interested in how weaker models might keep training stronger ones as the field runs out of better supervisors will find the mechanism and its research lineage laid out in detail.

Sources:
1. Weak-to-Strong On-Policy Distillation — Fangxu Yu, Weijia Xu, Michael Xu, Tianyi Zhou, Zinan Lin, 2026
http://arxiv.org/abs/2607.26246
2. Distilling the Knowledge in a Neural Network — Geoffrey Hinton, Oriol Vinyals, Jeff Dean, 2015
https://scholar.google.com/scholar?q=Distilling+the+Knowledge+in+a+Neural+Network
3. A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning (DAgger) — Stéphane Ross, Geoffrey Gordon, J. Andrew Bagnell, 2011
https://scholar.google.com/scholar?q=A+Reduction+of+Imitation+Learning+and+Structured+Prediction+to+No-Regret+Online+Learning+%28DAgger%29
4. On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes (Generalized Knowledge Distillation, GKD) — Rishabh Agarwal, Nino Vieillard, et al. (Google DeepMind), 2024
https://scholar.google.com/scholar?q=On-Policy+Distillation+of+Language+Models%3A+Learning+from+Self-Generated+Mistakes+%28Generalized+Knowledge+Distillation%2C+GKD%29
5. On-Policy Distillation (blog post / technical report) — Thinking Machines Lab (cited in the paper as 'Lu & Lab'), 2025
https://scholar.google.com/scholar?q=On-Policy+Distillation+%28blog+post+%2F+technical+report%29
6. Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision — Collin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker, Leo Gao, Leopold Aschenbrenner, Yining Chen, Adrien Ecoffet, Manas Joglekar, Jan Leike, Ilya Sutskever, Jeff Wu, 2023
https://scholar.google.com/scholar?q=Weak-to-Strong+Generalization%3A+Eliciting+Strong+Capabilities+With+Weak+Supervision
7. Weak-to-Strong Generalization beyond Accuracy: a Roadmap in Codegen, Safety, and Beyond (or closely related 2024 weak-to-strong follow-up work, cited in the paper as 'Ning et al., 2024') — Ning et al., 2024
https://scholar.google.com/scholar?q=Weak-to-Strong+Generalization+beyond+Accuracy%3A+a+Roadmap+in+Codegen%2C+Safety%2C+and+Beyond+%28or+closely+related+2024+weak-to-strong+follow-up+work%2C+cited+in+the+paper+as+%27Ning+et+al.%2C+2024%27%29
8. Illustrating Reinforcement Learning from Human Feedback / scalable oversight lineage (e.g. Christiano et al., 'Deep Reinforcement Learning from Human Preferences', and related scalable-oversight work) — Paul Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, Dario Amodei (RLHF); related scalable oversight literature, 2017 (RLHF) / ongoing scalable oversight literature
https://scholar.google.com/scholar?q=Illustrating+Reinforcement+Learning+from+Human+Feedback+%2F+scalable+oversight+lineage+%28e.g.+Christiano+et+al.%2C+%27Deep+Reinforcement+Learning+from+Human+Preferences%27%2C+and+related+scalable-oversight+work%29
9. Contrastive Decoding: Open-ended Text Generation as Optimization — Li, Holtzman, Fried, Liang, Eisner, Hashimoto, Zettlemoyer, Lewis, 2023
https://scholar.google.com/scholar?q=Contrastive+Decoding%3A+Open-ended+Text+Generation+as+Optimization
10. Distillation Scaling Laws — Busbridge, Ramapuram, Ablin, et al. (Apple), 2024
https://scholar.google.com/scholar?q=Distillation+Scaling+Laws
11. Speculative Decoding papers (e.g., Leviathan et al., 'Fast Inference from Transformers via Speculative Decoding', 2023) — Leviathan, Kalman, Matias, 2023
https://scholar.google.com/scholar?q=Speculative+Decoding+papers+%28e.g.%2C+Leviathan+et+al.%2C+%27Fast+Inference+from+Transformers+via+Speculative+Decoding%27%2C+2023%29
Interactive Visualization: Weak-to-Strong On-Policy Distillation Beats the Teacher

This episode examines why on-policy distillation (OPD) of large language models can fail catastrophically even when the teacher model is objectively stronger by every benchmark — a 7-billion-parameter teacher completely failed to improve a 1.5-billion-parameter student, while a smaller, weaker teacher succeeded. The discussion traces OPD's mechanics: unlike classic distillation, which trains students on teacher-generated text and suffers from exposure bias, OPD has the student generate its own rollouts and scores them against the teacher's token-level probability distribution via reverse KL divergence, yielding a dense per-token reward without a verifier. The central finding is that teacher-student "overlap ratio" — how much the teacher's likely next tokens actually match the student's own candidate set — determines whether that dense signal teaches anything, meaning benchmark strength and teachability are fundamentally different axes. The paper's three-part structure (phenomenology, mechanism, recipe) is illustrated through a controlled comparison of two similarly-scored Qwen3-4B teacher variants, isolating thinking-pattern compatibility as the real driver of distillation success. Listeners working on post-training or distillation pipelines will find this a direct challenge to the common assumption that upgrading the teacher model is a safe, unconditional improvement.

Sources:
1. Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe — Yaxuan Li, Yuxin Zuo, Bingxiang He, Jinqian Zhang, Chaojun Xiao, Cheng Qian, Tianyu Yu, Huan-ang Gao, Wenkai Yang, Zhiyuan Liu, Ning Ding, 2026
http://arxiv.org/abs/2604.13016
2. A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning — Stéphane Ross, Geoffrey Gordon, J. Andrew Bagnell, 2011
https://scholar.google.com/scholar?q=A+Reduction+of+Imitation+Learning+and+Structured+Prediction+to+No-Regret+Online+Learning
3. MiniLLM: Knowledge Distillation of Large Language Models — Yuxian Gu, Li Dong, Furu Wei, Minlie Huang, 2023
https://scholar.google.com/scholar?q=MiniLLM%3A+Knowledge+Distillation+of+Large+Language+Models
4. On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes (GKD) — Rishabh Agarwal, Nino Vieillard, and colleagues (Google DeepMind), 2024
https://scholar.google.com/scholar?q=On-Policy+Distillation+of+Language+Models%3A+Learning+from+Self-Generated+Mistakes+%28GKD%29
5. Qwen3 Technical Report — An Yang and the Qwen Team (Alibaba), 2025
https://scholar.google.com/scholar?q=Qwen3+Technical+Report
6. On-policy distillation of language models: Learning from self-generated mistakes (MiniLLM) — Yuxian Gu, Li Dong, Furu Wei, Minlie Huang, 2023
https://scholar.google.com/scholar?q=On-policy+distillation+of+language+models%3A+Learning+from+self-generated+mistakes+%28MiniLLM%29
7. Distillation scaling laws — Dan Busbridge, Amitis Shidani, Floris Weers, Jason Ramapuram, Etai Littwin, Russ Webb, 2025
https://scholar.google.com/scholar?q=Distillation+scaling+laws
8. On the efficacy of knowledge distillation — Jang Hyun Cho, Bharath Hariharan, 2019
https://scholar.google.com/scholar?q=On+the+efficacy+of+knowledge+distillation
9. Small models struggle to learn from strong reasoners — Yuetai Li, Xiang Yue, Zhangchen Xu, Fengqing Jiang, Luyao Niu, Bill Yuchen Lin, Bhaskar Ramasubramanian, Radha Poovendran, 2025
https://scholar.google.com/scholar?q=Small+models+struggle+to+learn+from+strong+reasoners
10. On-policy distillation (Thinking Machines Lab blog) — Kevin Lu and Thinking Machines Lab, 2025
https://scholar.google.com/scholar?q=On-policy+distillation+%28Thinking+Machines+Lab+blog%29
11. Learning beyond teacher: Generalized on-policy distillation with reward extrapolation — Wenkai Yang, Weijie Liu, Ruobing Xie, Kai Yang, Saiyong Yang, Yankai Lin, 2026
https://scholar.google.com/scholar?q=Learning+beyond+teacher%3A+Generalized+on-policy+distillation+with+reward+extrapolation
Interactive Visualization: On-Policy Distillation: Why a Stronger Teacher Can Backfire

This episode examines "Hidden in Memory: Sleeper Memory Poisoning in LLM Agents," a study showing how attackers can plant fabricated facts into an AI assistant's persistent memory that lie dormant until triggered in an unrelated future conversation. Unlike traditional prompt injection, which manipulates a model's behavior only within a single session, this attack targets the memory-write step itself, allowing a single black-box, universal payload template—refined through an actor-critic search between attacker and critic LLMs—to succeed across arbitrary goals with startlingly high rates (up to 99.8% on GPT-5.5). The discussion breaks down the three-stage pipeline attackers must clear (injection, retrieval, and usage), and highlights a clever technique for maximizing the odds a poisoned memory resurfaces later: rewriting it to boost embedding similarity with plausible future queries while a semantic-consistency judge guards against the rewrite drifting from the original intent. Testing spans 700 document-goal pairs across 15 source types and multiple commercial memory architectures, revealing that whether the model or a separate manager process controls memory writes dramatically changes how exploitable a system is. It's a sobering look at how "memory" — now a default feature across ChatGPT, Claude, Gemini, and agent frameworks like Mem0 — introduces a persistent, hard-to-detect attack surface that outlives the malicious content that created it.

Sources:
1. Hidden in Memory: Sleeper Memory Poisoning in LLM Agents — Sidharth Pulipaka, Stanislau Hlebik, Leonidas Raghav, Sahar Abdelnabi, Vyas Raina, Ivaxi Sheth, Mario Fritz, 2026
http://arxiv.org/abs/2605.15338
2. Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training — Evan Hubinger et al. (Anthropic), 2024
https://scholar.google.com/scholar?q=Sleeper+Agents%3A+Training+Deceptive+LLMs+that+Persist+Through+Safety+Training
3. Not What You've Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection — Kai Greshake, Sahar Abdelnabi, et al., 2023
https://scholar.google.com/scholar?q=Not+What+You%27ve+Signed+Up+For%3A+Compromising+Real-World+LLM-Integrated+Applications+with+Indirect+Prompt+Injection
4. Prompt Injection Attack Against LLM-Integrated Applications — Yi Liu et al., 2023
https://scholar.google.com/scholar?q=Prompt+Injection+Attack+Against+LLM-Integrated+Applications
5. Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory — Prateek Chhikara et al., 2025
https://scholar.google.com/scholar?q=Mem0%3A+Building+Production-Ready+AI+Agents+with+Scalable+Long-Term+Memory
6. Universal and Transferable Adversarial Attacks on Aligned Language Models (GCG) — Zou et al., 2023
https://scholar.google.com/scholar?q=Universal+and+Transferable+Adversarial+Attacks+on+Aligned+Language+Models+%28GCG%29
7. AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases — Chen, Xiang, Xiao, Song, Li, 2024 (NeurIPS)
https://scholar.google.com/scholar?q=AgentPoison%3A+Red-teaming+LLM+Agents+via+Poisoning+Memory+or+Knowledge+Bases
8. GEPA: Efficient Textual Optimization via LLM-based Reflection and Pareto-Efficient Evolutionary Search — Agrawal, Khattab, Potts, 2025
https://scholar.google.com/scholar?q=GEPA%3A+Efficient+Textual+Optimization+via+LLM-based+Reflection+and+Pareto-Efficient+Evolutionary+Search
9. Injection through web agents that fetch pages with hidden HTML instructions — Raghav and Choong, 2026
https://scholar.google.com/scholar?q=Injection+through+web+agents+that+fetch+pages+with+hidden+HTML+instructions
10. The Trigger in the Haystack: Extracting and Reconstructing LLM Backdoor Triggers — Bullwinkel, Severi, Hines, Minnich, Kumar, Zunger, 2026
https://scholar.google.com/scholar?q=The+Trigger+in+the+Haystack%3A+Extracting+and+Reconstructing+LLM+Backdoor+Triggers
Interactive Visualization: Sleeper Memory Poisoning: When Assistants Remember Lies

This episode examines a paper by Emmanuel Dupoux, Yann LeCun, and Jitendra Malik arguing that deployed AI models learn nothing after training, unlike a toddler who continuously experiments through action, observation, imitation, and inquiry. The discussion breaks down the paper's core distinction between System A (passive, observation-based statistical learning like self-supervised training) and System B (action-based reinforcement learning through feedback), and explains why neither alone can produce autonomous intelligence. It then covers the paper's proposed fix, System M, an orchestrator modeled on software-defined networking that monitors low-bandwidth "meta-state" signals like prediction error and confidence to dynamically route between learning systems, automating what human MLOps engineers currently do by hand. The conversation also connects this framework to LeCun's 2022 autonomous machine intelligence proposal and the ongoing debate sparked by Silver and Sutton's "Era of Experience" critique about AI hitting a data wall. Listeners interested in the architecture of autonomous learning and what's actually missing between today's static models and genuinely adaptive intelligence will find the systems-level framing illuminating.

Sources:
1. Why AI systems don't learn and what to do about it: Lessons on autonomous learning from cognitive science — Emmanuel Dupoux, Yann LeCun, Jitendra Malik, 2026
http://arxiv.org/abs/2603.15381
2. A Path Towards Autonomous Machine Intelligence — Yann LeCun, 2022
https://scholar.google.com/scholar?q=A+Path+Towards+Autonomous+Machine+Intelligence
3. Welcome to the Era of Experience — David Silver, Richard Sutton, 2025
https://scholar.google.com/scholar?q=Welcome+to+the+Era+of+Experience
4. Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model (MuZero) — Julian Schrittwieser et al., 2020
https://scholar.google.com/scholar?q=Mastering+Atari%2C+Go%2C+Chess+and+Shogi+by+Planning+with+a+Learned+Model+%28MuZero%29
5. V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning — Mido Assran et al. (incl. LeCun), 2025
https://scholar.google.com/scholar?q=V-JEPA+2%3A+Self-Supervised+Video+Models+Enable+Understanding%2C+Prediction+and+Planning
6. Coordination Among Neural Modules Through a Shared Global Workspace — Anirudh Goyal, Aniket Didolkar, et al., 2022
https://scholar.google.com/scholar?q=Coordination+Among+Neural+Modules+Through+a+Shared+Global+Workspace
7. Embodied AI Agents: Modeling the World — Pascale Fung, Emmanuel Dupoux, Jitendra Malik, et al., 2025
https://scholar.google.com/scholar?q=Embodied+AI+Agents%3A+Modeling+the+World
Interactive Visualization: Why AI Systems Don't Learn After Deployment

This episode examines "Runtime-Orchestrated Second-Order Optimization for Scalable LLM Training," a May 2026 Oxford paper introducing Asteria, a runtime system rather than a new optimizer. The discussion covers why curvature-aware methods like Shampoo and SOAP have never displaced AdamW despite converging in fewer steps, tracing the lineage from K-FAC through Distributed Shampoo to SOAP and explaining the Kronecker-factorization tricks that make tracking curvature tractable at all. The hosts unpack the paper's "three physical walls" framework — a vertical capacity wall from single-GPU memory limits, an overlap disruption wall where cubic-cost matrix operations stall compute-communication overlap, and a global consensus wall from synchronous full-state updates across mismatched network speeds — and debate whether reengineering the plumbing around an unchanged optimizer counts as a genuine research contribution. Listeners interested in distributed training infrastructure, optimizer design trade-offs, or the gap between algorithmic elegance and practical deployability will find the back-and-forth over real benchmark numbers (96 seconds versus 1.5 seconds per step) especially grounded.

Sources:
1. Runtime-Orchestrated Second-Order Optimization for Scalable LLM Training — Yishun Lu, Junhao Zhang, Zeyu Yang, Wes Armour, 2026
http://arxiv.org/abs/2605.16184
2. Optimizing Neural Networks with Kronecker-factored Approximate Curvature — James Martens, Roger Grosse, 2015
https://scholar.google.com/scholar?q=Optimizing+Neural+Networks+with+Kronecker-factored+Approximate+Curvature
3. Shampoo: Preconditioned Stochastic Tensor Optimization — Vineet Gupta, Tomer Koren, Yoram Singer, 2018
https://scholar.google.com/scholar?q=Shampoo%3A+Preconditioned+Stochastic+Tensor+Optimization
4. Scalable Second Order Optimization for Deep Learning — Rohan Anil, Vineet Gupta, Tomer Koren, Kevin Regan, Yoram Singer, 2020
https://scholar.google.com/scholar?q=Scalable+Second+Order+Optimization+for+Deep+Learning
5. SOAP: Improving and Stabilizing Shampoo using Adam — Nikhil Vyas, Depen Morwani, Rosie Zhao, Itai Shapira, David Brandfonbrener, Sham Kakade, et al., 2024
https://scholar.google.com/scholar?q=SOAP%3A+Improving+and+Stabilizing+Shampoo+using+Adam
6. A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale — H.-J. M. Shi, T.-H. Lee, S. Iwasaki, J. Gallego-Posada, Z. Li, K. Rangadurai, D. Mudigere, M. Rabbat, 2023
https://scholar.google.com/scholar?q=A+Distributed+Data-Parallel+PyTorch+Implementation+of+the+Distributed+Shampoo+Optimizer+for+Training+Neural+Networks+At-Scale
7. Deep Optimizer States: Towards Scalable Training of Transformer Models Using Interleaved Offloading — A. Maurya, J. Ye, M. M. Rafique, F. Cappello, B. Nicolae, 2024
https://scholar.google.com/scholar?q=Deep+Optimizer+States%3A+Towards+Scalable+Training+of+Transformer+Models+Using+Interleaved+Offloading
8. ZeRO-Infinity: Breaking the GPU Memory Wall for Extreme Scale Deep Learning — S. Rajbhandari, O. Ruwase, J. Rasley, S. Smith, Y. He, 2021
https://scholar.google.com/scholar?q=ZeRO-Infinity%3A+Breaking+the+GPU+Memory+Wall+for+Extreme+Scale+Deep+Learning
9. Understanding and Improving Shampoo and SOAP via Kullback-Leibler Minimization — W. Lin, S. C. Lowe, F. Dangel, R. Eschenhagen, Z. Xu, R. B. Grosse, 2026
https://scholar.google.com/scholar?q=Understanding+and+Improving+Shampoo+and+SOAP+via+Kullback-Leibler+Minimization
10. Beyond the Mean: Fisher-Orthogonal Projection for Natural Gradient Descent in Large Batch Training — Y. Lu, W. Armour, 2026
https://scholar.google.com/scholar?q=Beyond+the+Mean%3A+Fisher-Orthogonal+Projection+for+Natural+Gradient+Descent+in+Large+Batch+Training
Interactive Visualization: Second-Order Optimization Meets Runtime Scheduling at Scale

This episode digs into a paper testing whether hand-written PTX assembly beats NVIDIA's WMMA API for Tensor Core GEMM kernels on an L4 GPU, finding that the answer flips depending on numeric precision rather than holding as a universal rule. The discussion covers the hardware distinction between Tensor Cores and regular CUDA cores, and contrasts the convenience of WMMA against the finer control PTX offers through instructions like cp.async, ldmatrix, and mma.sync. A key thread traces why this matters in practice: quantized LLM serving at INT8 or INT4 shifts kernels from compute-bound to memory-bound, making the precision-dependent payoff of hand-tuned PTX directly relevant to running open-weight models cheaply. The episode also addresses the methodological choice to test on a single GPU, arguing that isolating precision and instruction-set effects requires holding hardware constant rather than spreading across devices. Listeners get a concrete framework for deciding when the extra engineering effort of writing raw PTX is worth it versus when it's wasted work.

Sources:
1. Hand-Written PTX Tensor-Core GEMM Kernels: A Multi-Precision Study on NVIDIA L4 — Matt J. Borowski, Blazej Osinski, 2026
http://arxiv.org/abs/2608.10103
2. NVIDIA Tensor Core Programmability, Performance & Precision — Stefano Markidis, Steven W. D. Chien, Erwin Laure, Ivy B. Peng, Jeffrey S. Vetter, 2018
https://scholar.google.com/scholar?q=NVIDIA+Tensor+Core+Programmability%2C+Performance+%26+Precision
3. Dissecting the NVIDIA Volta GPU Architecture via Microbenchmarking — Zhe Jia, Marco Maggioni, Jeffrey Smith, Daniele Paolo Scarpazza, 2018
https://scholar.google.com/scholar?q=Dissecting+the+NVIDIA+Volta+GPU+Architecture+via+Microbenchmarking
4. CUTLASS: CUDA Templates for Linear Algebra Subroutines — Andrew Kerr, Duane Merrill, Julien Demouth, John Tran (NVIDIA), with ongoing project contributors, 2018 (initial release, actively maintained since)
https://scholar.google.com/scholar?q=CUTLASS%3A+CUDA+Templates+for+Linear+Algebra+Subroutines
5. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness — Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, Christopher Ré, 2022
https://scholar.google.com/scholar?q=FlashAttention%3A+Fast+and+Memory-Efficient+Exact+Attention+with+IO-Awareness
6. Triton: An Intermediate Language and Compiler for Tiled Neural Network Computations — Philippe Tillet, H. T. Kung, David Cox, 2019
https://scholar.google.com/scholar?q=Triton%3A+An+Intermediate+Language+and+Compiler+for+Tiled+Neural+Network+Computations
7. Ansor: Generating High-Performance Tensor Programs for Deep Learning — Lianmin Zheng, Chengfan Jia, Minmin Sun, Zhao Wu, Cody Hao Yu, Ameer Haj-Ali, Yida Wang, Jun Yang, Danyang Zhuo, Koushik Sen, Joseph E. Gonzalez, Ion Stoica, 2020 (OSDI)
https://scholar.google.com/scholar?q=Ansor%3A+Generating+High-Performance+Tensor+Programs+for+Deep+Learning
8. Dissecting the Ampere GPU Architecture via Microbenchmarking — Wei Sun, Ang Li, Tong Geng, Sander Stuijk, Henk Corporaal, 2022
https://scholar.google.com/scholar?q=Dissecting+the+Ampere+GPU+Architecture+via+Microbenchmarking
9. TVM: An Automated End-to-End Optimizing Compiler for Deep Learning — Tianqi Chen, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Yan, Haichen Shen, Meghan Cowan, Leyuan Wang, Yuwei Hu, Luis Ceze, Carlos Guestrin, Arvind Krishnamurthy, 2018 (OSDI)
https://scholar.google.com/scholar?q=TVM%3A+An+Automated+End-to-End+Optimizing+Compiler+for+Deep+Learning
10. Understanding Latency Hiding on GPUs — V. Volkov, 2016
https://scholar.google.com/scholar?q=Understanding+Latency+Hiding+on+GPUs
11. AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration — J. Lin et al., 2024
https://scholar.google.com/scholar?q=AWQ%3A+Activation-aware+Weight+Quantization+for+LLM+Compression+and+Acceleration
12. FlashInfer: Kernel Library for LLM Serving — Z. Ye et al., 2024
https://scholar.google.com/scholar?q=FlashInfer%3A+Kernel+Library+for+LLM+Serving

This episode examines "Mechanist," a multi-agent system built by researchers at Zhejiang University, NUS, Southern University of Science and Technology, Heriot-Watt, UC San Diego, and Northeastern University to automate mechanistic interpretability research itself, rather than automating experiments in an external domain like chemistry or biology. The discussion covers how a central orchestrator coordinates four agents—hypothesis, experiment, verification, and iteration—drawing on a 13,000-paper interpretability knowledge graph and a 43-million-paper cross-disciplinary graph called SciAtlas to generate and test theories about how models actually compute. Key concepts explored include subliminal learning, where a trait transfers from teacher to student model through data that looks unrelated to it, and a three-frame belief decomposition (World Knowledge, Personal Belief, Attributed Belief) used to probe whether models genuinely separate fact from attributed belief. The episode previews four escalating case studies, starting with the discovery of a previously unflagged multimodal safety risk and building toward using mechanistic theories to directly intervene on model internals and even steer a biological system, raising the question of whether an AI system can meaningfully explain the black box that produced it.

Sources:
1. Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence — Mengru Wang, Junfeng Fang, Shuofei Qiao, Zhenqian Xu, Haoming Xu, Haoxiong Wang, Shumin Deng, Linyi Yang, Zhixiang Cui, Xin Xu, Yunzhi Yao, Buqiang Xu, Fei Shen, Haozhe Luo, Yunxiang Wei, Ningyu Zhang, Julian McAuley, Tat Seng Chua, Huajun Chen, 2026
http://arxiv.org/abs/2608.12036
2. Language models transmit behavioural traits through hidden signals in data — Alex Cloud, Minh Le, James Chua, Jan Betley, Anna Sztyber-Betley, Sören Mindermann, Jacob Hilton, Samuel Marks, Owain Evans, 2026 (Nature)
https://scholar.google.com/scholar?q=Language+models+transmit+behavioural+traits+through+hidden+signals+in+data
3. Subliminal learning is a lora artifact — Todd Nief, Harvey Yiyun Fu, Mark Muchane, Ari Holtzman, 2026
https://scholar.google.com/scholar?q=Subliminal+learning+is+a+lora+artifact
4. Language models cannot reliably distinguish belief from knowledge and fact — Mirac Suzgun, Tayfun Gur, Federico Bianchi, Daniel E. Ho, Thomas Icard, Dan Jurafsky, James Zou, 2025 (Nature Machine Intelligence)
https://scholar.google.com/scholar?q=Language+models+cannot+reliably+distinguish+belief+from+knowledge+and+fact
5. Genome modelling and design across all domains of life with evo 2 — Garyk Brixi, Matthew G. Durrant, Jerome Ku, et al., 2026 (Nature)
https://scholar.google.com/scholar?q=Genome+modelling+and+design+across+all+domains+of+life+with+evo+2
6. Towards end-to-end automation of ai research — Chris Lu, Cong Lu, Robert Tjarko Lange, Yutaro Yamada, Shengran Hu, Jakob Foerster, David Ha, Jeff Clune, 2026 (Nature)
https://scholar.google.com/scholar?q=Towards+end-to-end+automation+of+ai+research
7. Sleeper agents: Training deceptive llms that persist through safety training — Evan Hubinger et al., 2024
https://scholar.google.com/scholar?q=Sleeper+agents%3A+Training+deceptive+llms+that+persist+through+safety+training
8. Conditional misalignment: common interventions can hide emergent misalignment behind contextual triggers — Jan Dubinski, Jan Betley, Anna Sztyber-Betley, Daniel Tan, Owain Evans, 2026
https://scholar.google.com/scholar?q=Conditional+misalignment%3A+common+interventions+can+hide+emergent+misalignment+behind+contextual+triggers

This episode dissects the internal architecture of Claude Code by examining its extracted TypeScript source (v2.1.88) alongside two other agent systems, OpenClaw and Hermes Agent, drawing on the paper "Dive into Claude Code" by Jiacheng Liu et al. from VILA Lab at MBZUAI. It reveals that the model's reasoning core is essentially a single while-loop — literally called queryLoop() — with everything else (permissions, context management, tools, subagents) built as scaffolding around it. The discussion covers deny-first permission rules, the five-stage compaction pipeline for managing context windows, subagent delegation with isolated context windows, and how the Model Context Protocol connects to external tool servers. The hosts also extract five human values embedded directly in the code — human decision authority, safety/security/privacy, reliable execution, capability amplification, and contextual adaptability — framing the episode as less about how Claude Code works and more about what its designers chose to prioritize, made visible through actual implementation choices.

Sources:
1. Dive into Claude Code: The Design Space of Today's and Future AI Agent Systems — Jiacheng Liu, Xiaohan Zhao, Xinyi Shang, Zhiqiang Shen, 2026
http://arxiv.org/abs/2604.14228
2. Not What You've Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection — Kai Greshake, Sahar Abdelnabi, Shailesh Mishra, Christoph Endres, Thorsten Holz, Mario Fritz, 2023
https://scholar.google.com/scholar?q=Not+What+You%27ve+Signed+Up+For%3A+Compromising+Real-World+LLM-Integrated+Applications+with+Indirect+Prompt+Injection
3. GAIA: A Benchmark for General AI Assistants — Grégoire Mialon, Clémentine Fourrier, Craig Swift, Thomas Wolf, Yann LeCun, Thomas Scialom, 2023
https://scholar.google.com/scholar?q=GAIA%3A+A+Benchmark+for+General+AI+Assistants
4. Constitutional AI: Harmlessness from AI Feedback — Yuntao Bai, Saurav Kadavath, Sandipan Kundu, et al., 2022
https://scholar.google.com/scholar?q=Constitutional+AI%3A+Harmlessness+from+AI+Feedback
5. AgentBench: Evaluating LLMs as Agents — Xiao Liu, Hao Yu, Hanchen Zhang, et al., 2023
https://scholar.google.com/scholar?q=AgentBench%3A+Evaluating+LLMs+as+Agents

This episode examines "Is Grep All You Need? How Agent Harnesses Reshape Agentic Search," which tests whether simple regex-based retrieval can outperform vector search inside agentic pipelines like Chronos, Claude Code, and Codex CLI. The hosts dig into how retrieval mode interacts with harness architecture, delivery method (inline vs. programmatic), and backbone model choice, finding that inline grep beats inline vector search across every harness-model pairing tested — with gaps as wide as twenty points and swings as large as switching harnesses entirely. A striking case shows the same model scoring 93.1% on one harness but only 76.7% on another, suggesting orchestration and prompt construction matter as much as the retrieval algorithm itself. The discussion also surfaces a counterintuitive twist: forcing an agent to read retrieved results from a file instead of getting them dumped inline can nearly halve accuracy, even with identical underlying search. Listeners interested in RAG, agent design, or LLM evaluation methodology will find the paper's tangled-but-honest approach to measuring real deployed systems a useful corrective to cleaner but less realistic ablation studies.

Sources:
1. Is Grep All You Need? How Agent Harnesses Reshape Agentic Search — Sahil Sen, Akhil Kasturi, Elias Lumer, Anmol Gulati, Vamse Kumar Subbiah, 2026
http://arxiv.org/abs/2605.15184
2. ReAct: Synergizing Reasoning and Acting in Language Models — Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, Yuan Cao, 2022
https://scholar.google.com/scholar?q=ReAct%3A+Synergizing+Reasoning+and+Acting+in+Language+Models
3. WebGPT: Browser-assisted question-answering with human feedback — Reiichiro Nakano, Jacob Hilton, Suchir Balaji, et al. (OpenAI), 2021
https://scholar.google.com/scholar?q=WebGPT%3A+Browser-assisted+question-answering+with+human+feedback
4. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks — Patrick Lewis, Ethan Perez, Aleksandra Piktus, et al. (Facebook AI Research), 2020
https://scholar.google.com/scholar?q=Retrieval-Augmented+Generation+for+Knowledge-Intensive+NLP+Tasks
5. LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory — Di Wu, Hongwei Wang, Wenhao Yu, et al., 2024
https://scholar.google.com/scholar?q=LongMemEval%3A+Benchmarking+Chat+Assistants+on+Long-Term+Interactive+Memory
6. The Probabilistic Relevance Framework: BM25 and Beyond — Stephen Robertson, Hugo Zaragoza, 2009
https://scholar.google.com/scholar?q=The+Probabilistic+Relevance+Framework%3A+BM25+and+Beyond
7. Dense Passage Retrieval for Open-Domain Question Answering — Vladimir Karpukhin, Barlas Oğuz, Sewon Min, et al. (Facebook AI Research), 2020
https://scholar.google.com/scholar?q=Dense+Passage+Retrieval+for+Open-Domain+Question+Answering
8. Lost in the Middle: How Language Models Use Long Contexts — Nelson F. Liu, Kevin Lin, John Hewitt, et al., 2023
https://scholar.google.com/scholar?q=Lost+in+the+Middle%3A+How+Language+Models+Use+Long+Contexts
9. SPLADE: Sparse Lexical and Expansion Model for First Stage Ranking — Thibault Formal, Benjamin Piwowarski, Stéphane Clinchant, 2021
https://scholar.google.com/scholar?q=SPLADE%3A+Sparse+Lexical+and+Expansion+Model+for+First+Stage+Ranking
10. BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models — Nandan Thakur, Nils Reimers, Andreas Rücklé, Abhishek Srivastava, Iryna Gurevych, 2021
https://scholar.google.com/scholar?q=BEIR%3A+A+Heterogenous+Benchmark+for+Zero-shot+Evaluation+of+Information+Retrieval+Models
11. MemGPT: Towards LLMs as Operating Systems — Charles Packer, Vivian Fang, Shishir G. Patil, Kevin Lin, Sarah Wooders, Joseph E. Gonzalez, 2023
https://scholar.google.com/scholar?q=MemGPT%3A+Towards+LLMs+as+Operating+Systems
12. Chronos: Temporal-Aware Conversational Agents with Structured Event Retrieval for Long-Term Memory — Sahil Sen, Elias Lumer, Anmol Gulati, Vamse Kumar Subbiah, 2026
https://scholar.google.com/scholar?q=Chronos%3A+Temporal-Aware+Conversational+Agents+with+Structured+Event+Retrieval+for+Long-Term+Memory
Interactive Visualization: Grep vs Vector Search: How Agent Harnesses Shape Retrieval Accuracy

This episode dissects a common but misleading claim in LLM serving benchmarks: that swapping to a faster stack alone explains a headline speedup number. It walks through the mechanics separating prefill's compute-bound math from decode's memory-bandwidth-bound token generation, then explains how continuous batching and PagedAttention keep GPUs saturated, and how GPTQ quantization paired with the Marlin kernel avoids costly dequantization overhead. By introducing a matched FP16 intermediate stack, the paper cleanly splits a reported speedup into a runtime factor and a kernel-plus-quantization factor that multiply back to the observed total. The headline finding is striking: a controlled, apples-to-apples comparison yields a 2.58x speedup dominated by runtime (68%), while a more common operational comparison — pitting a batch-capped baseline against a fully loaded modern stack — inflates that to 10.62x, with runtime's share climbing to 86%. Listeners get a rare, rigorous look at how much of the "free lunch" in inference speedups is real architecture versus benchmark framing.

Sources:
1. Decomposing Speedups Across Runtime, Kernel, and Quantization
https://arxiv.org/pdf/2607.11368
2. FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU — Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, et al., 2023
https://scholar.google.com/scholar?q=FlexGen%3A+High-Throughput+Generative+Inference+of+Large+Language+Models+with+a+Single+GPU
3. DeepSpeed-Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale — Reza Yazdani Aminabadi, Samyam Rajbhandari, Ammar Ahmad Awan, et al., 2022
https://scholar.google.com/scholar?q=DeepSpeed-Inference%3A+Enabling+Efficient+Inference+of+Transformer+Models+at+Unprecedented+Scale
4. SqueezeLLM: Dense-and-Sparse Quantization — Sehoon Kim, Coleman Hooper, Amir Gholami, et al., 2023
https://scholar.google.com/scholar?q=SqueezeLLM%3A+Dense-and-Sparse+Quantization
5. Blink: Fast and Generic Collectives for Distributed ML — Guanhua Wang, Shivaram Venkataraman, Amar Phanishayee, et al., 2020
https://scholar.google.com/scholar?q=Blink%3A+Fast+and+Generic+Collectives+for+Distributed+ML
Interactive Visualization: Decomposing Speedups Across Runtime, Kernel, and Quantization

This episode examines a study analyzing the growth of AI-generated content across the open web, drawing on 33 monthly samples from the Internet Archive's Wayback Machine between August 2022 and May 2025. It highlights the paper's central finding that AI-generated or AI-assisted content on newly published websites rose from zero before ChatGPT's launch to roughly 35 percent by mid-2025, and explores how the authors transform "Dead Internet Theory" from internet folklore into six testable hypotheses — including semantic contraction, truth decay, positivity shift, epistemic islands, entropy dilution, and stylistic monoculture. The discussion covers the methodology behind sampling a representative slice of the internet, including logarithmic downsampling and stratification across time, MIME type, and domain to avoid bias toward heavily-crawled sites. It also connects the findings to the concept of model collapse, framing the 35 percent figure as empirical evidence for a previously theoretical concern about AI models training on their own synthetic output. Listeners interested in web ecosystem health, LLM training data quality, or the intersection of internet culture and rigorous data science will find the episode's blend of meme-to-metric translation particularly compelling.

Sources:
1. The Impact of AI-Generated Text on the Internet — Jonas Dolezal, Sawood Alam, Mark Graham, Maty Bohacek, 2026
http://arxiv.org/abs/2604.26965
2. DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability Curvature — Eric Mitchell, Yoonho Lee, Alexander Khazatsky, Christopher D. Manning, Chelsea Finn, 2023
https://scholar.google.com/scholar?q=DetectGPT%3A+Zero-Shot+Machine-Generated+Text+Detection+using+Probability+Curvature
3. Can AI-Generated Text be Reliably Detected? — Vinu Sankar Sadasivan, Aounon Kumar, Sriram Balasubramanian, Wenxiao Wang, Soheil Feizi, 2023
https://scholar.google.com/scholar?q=Can+AI-Generated+Text+be+Reliably+Detected%3F
4. A Watermark for Large Language Models — John Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz, Ian Miers, Tom Goldstein, 2023
https://scholar.google.com/scholar?q=A+Watermark+for+Large+Language+Models
5. GLTR: Statistical Detection and Visualization of Generated Text — Sebastian Gehrmann, Hendrik Strobelt, Alexander M. Rush, 2019
https://scholar.google.com/scholar?q=GLTR%3A+Statistical+Detection+and+Visualization+of+Generated+Text
6. The Curse of Recursion: Training on Generated Data Makes Models Forget — Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Yarin Gal, Nicolas Papernot, Ross Anderson, 2024 (Nature, updated from 2023 preprint)
https://scholar.google.com/scholar?q=The+Curse+of+Recursion%3A+Training+on+Generated+Data+Makes+Models+Forget
7. Is the Internet Dead? Evaluating Claims about Web Homogenization from AI-Generated Content (representative of the broader 2024-2025 measurement literature, e.g. Muzumdar et al. and related environmental-scanning studies of Dead Internet Theory discourse) — Various (this cluster of work is cited in the paper as Muzumdar et al., 2025 and similar), 2025
https://scholar.google.com/scholar?q=Is+the+Internet+Dead%3F+Evaluating+Claims+about+Web+Homogenization+from+AI-Generated+Content+%28representative+of+the+broader+2024-2025+measurement+literature%2C+e.g.+Muzumdar+et+al.+and+related+environmental-scanning+studies+of+Dead+Internet+Theory+discourse%29
8. Studies on bot/inauthentic-content prevalence on specific platforms (e.g., La Cava et al. 2025 on social media, Matatov et al. 2024 on platform-specific AI content) — Lucio La Cava et al.; Jonathan Matatov et al. (representative platform-specific studies cited in the paper's related-work section), 2024-2025
https://scholar.google.com/scholar?q=Studies+on+bot%2Finauthentic-content+prevalence+on+specific+platforms+%28e.g.%2C+La+Cava+et+al.+2025+on+social+media%2C+Matatov+et+al.+2024+on+platform-specific+AI+content%29
9. Public Trust and Perceptions of Artificial Intelligence (Ipsos / Reuters Institute Digital News Report and Edelman Trust Barometer AI-focused editions) — Ipsos (various); Reuters Institute for the Study of Journalism (Nic Newman et al.); Edelman Trust Barometer team, 2023-2025 (recurring annual)
https://scholar.google.com/scholar?q=Public+Trust+and+Perceptions+of+Artificial+Intelligence+%28Ipsos+%2F+Reuters+Institute+Digital+News+Report+and+Edelman+Trust+Barometer+AI-focused+editions%29
10. Americans' Views of Artificial Intelligence (Pew Research Center recurring survey series) — Pew Research Center (Alec Tyson, Emma Kikuchi, and colleagues), 2023-2025 (recurring)
https://scholar.google.com/scholar?q=Americans%27+Views+of+Artificial+Intelligence+%28Pew+Research+Center+recurring+survey+series%29
11. The Perception Gap: Comparing Public Beliefs about Misinformation to Empirical Prevalence Estimates (representative of the risk-perception vs. measured-prevalence literature this paper's framing descends from, e.g. work following Duffy et al. and general misperception-of-misinformation-prevalence studies) — Andrew Guess, colleagues in the misinformation-prevalence research cluster (representative of this line, distinct from the AI-specific surveys above), 2019-2023 (foundational misinformation-perception literature)
https://scholar.google.com/scholar?q=The+Perception+Gap%3A+Comparing+Public+Beliefs+about+Misinformation+to+Empirical+Prevalence+Estimates+%28representative+of+the+risk-perception+vs.+measured-prevalence+literature+this+paper%27s+framing+descends+from%2C+e.g.+work+following+Duffy+et+al.+and+general+misperception-of-misinformation-prevalence+studies%29
12. Documenting the English Colossal Clean Crawled Corpus (C4) — Jesse Dodge, Maarten Sap, Ana Marasović, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, Matt Gardner, 2021
https://scholar.google.com/scholar?q=Documenting+the+English+Colossal+Clean+Crawled+Corpus+%28C4%29
13. The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only — Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Alessandro Cappelli, Hamza Alobeidli, Baptiste Pannier, Ebtesam Almazrouei, Julien Launay, 2023
https://scholar.google.com/scholar?q=The+RefinedWeb+Dataset+for+Falcon+LLM%3A+Outperforming+Curated+Corpora+with+Web+Data%2C+and+Web+Data+Only
14. Quantifying Memorization Across Neural Language Models / broader Common Crawl representativeness and bias studies (e.g., work on Common Crawl's domain and language skew) — Various (Common Crawl bias/representativeness literature, e.g. work by Luccioni & Viviano on Common Crawl content quality, and follow-on studies), 2021-2023
https://scholar.google.com/scholar?q=Quantifying+Memorization+Across+Neural+Language+Models+%2F+broader+Common+Crawl+representativeness+and+bias+studies+%28e.g.%2C+work+on+Common+Crawl%27s+domain+and+language+skew%29
15. Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research — Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, et al. (Allen Institute for AI), 2024
https://scholar.google.com/scholar?q=Dolma%3A+an+Open+Corpus+of+Three+Trillion+Tokens+for+Language+Model+Pretraining+Research
16. AI models collapse when trained on recursively generated data — I. Shumailov, Z. Shumaylov, Y. Zhao, N. Papernot, R. Anderson, Y. Gal, 2024
https://scholar.google.com/scholar?q=AI+models+collapse+when+trained+on+recursively+generated+data
17. Can AI-generated text be reliably detected? Stress testing AI text detectors under various attacks — V. S. Sadasivan, A. Kumar, S. Balasubramanian, W. Wang, S. Feizi, 2025
https://scholar.google.com/scholar?q=Can+AI-generated+text+be+reliably+detected%3F+Stress+testing+AI+text+detectors+under+various+attacks
18. RAID: A shared benchmark for robust evaluation of machine-generated text detectors — L. Dugan, A. Hwang, F. Trhlik, A. Zhu, J. M. Ludan, H. Xu, D. Ippolito, C. Callison-Burch, 2024
https://scholar.google.com/scholar?q=RAID%3A+A+shared+benchmark+for+robust+evaluation+of+machine-generated+text+detectors
19. When incentives backfire, data stops being human — S. Santy, P. Bhattacharya, M. H. Ribeiro, K. Allen, S. Oh, 2025
https://scholar.google.com/scholar?q=When+incentives+backfire%2C+data+stops+being+human
20. Longitudinal sampling of URLs from the Wayback Machine — K. Garg, S. Alam, D. Ayala, M. Graham, M. C. Weigle, M. L. Nelson, 2025
https://scholar.google.com/scholar?q=Longitudinal+sampling+of+URLs+from+the+Wayback+Machine
21. Verbalized sampling: How to mitigate mode collapse and unlock LLM diversity — J. Zhang, S. Yu, D. Chong, A. Sicilia, M. R. Tomz, C. D. Manning, W. Shi, 2025
https://scholar.google.com/scholar?q=Verbalized+sampling%3A+How+to+mitigate+mode+collapse+and+unlock+LLM+diversity
Interactive Visualization: AI-Generated Text and the Death of the Open Web

This episode explores "Approaching Shannon Bound with Lossless LLM Weight Compression," which argues that model weights stored in formats like bf16 carry far less real information than their bit-width implies—entropy measurements across six models and seven numeric formats show gaps of several bits per weight that can be recovered without any change to the underlying values. The discussion covers why memory capacity and bandwidth, not raw compute, are the real bottleneck in GPU inference, and why generic compressors like gzip fail on IEEE-754 floating point data. The hosts dig into Asymmetric Numeral Systems (ANS) as the key engineering breakthrough, since it decodes fast enough and in a tile-parallel enough fashion to run inside a live GPU kernel without becoming a new bottleneck itself. Listeners interested in the intersection of information theory and practical LLM serving will find the walkthrough of how lossless compression differs fundamentally from quantization methods like int4 or AWQ particularly compelling.

Sources:
1. Approaching Shannon Bound with Lossless LLM Weight Compression — Hongshi Tan, Yao Chen, Gustavo Alonso, Weng-Fai Wong, Bingsheng He, 2026
http://arxiv.org/abs/2606.15789
2. 70% Size, 100% Accuracy: Lossless LLM Compression for Efficient GPU Inference via Dynamic-Length Float — Tianyi Zhang, Yang Sui, Shaochen (Henry) Zhong, et al., 2025
https://scholar.google.com/scholar?q=70%25+Size%2C+100%25+Accuracy%3A+Lossless+LLM+Compression+for+Efficient+GPU+Inference+via+Dynamic-Length+Float
3. NeuZip: Memory-Efficient Training and Inference with Dynamic Compression of Neural Networks — Yongchang Hao, Yanshuai Cao, Lili Mou, 2024
https://scholar.google.com/scholar?q=NeuZip%3A+Memory-Efficient+Training+and+Inference+with+Dynamic+Compression+of+Neural+Networks
4. ZipNN: Lossless Compression for AI Models — Moshik Hershcovitch, Andrew Wood, Leshem Choshen, et al. (IBM Research), 2024
https://scholar.google.com/scholar?q=ZipNN%3A+Lossless+Compression+for+AI+Models
5. Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding — Song Han, Huizi Mao, William J. Dally, 2016
https://scholar.google.com/scholar?q=Deep+Compression%3A+Compressing+Deep+Neural+Networks+with+Pruning%2C+Trained+Quantization+and+Huffman+Coding
6. Asymmetric Numeral Systems: Entropy Coding Combining Speed of Huffman Coding with Compression Rate of Arithmetic Coding — Jarek Duda, 2013
https://scholar.google.com/scholar?q=Asymmetric+Numeral+Systems%3A+Entropy+Coding+Combining+Speed+of+Huffman+Coding+with+Compression+Rate+of+Arithmetic+Coding
7. The Use of Asymmetric Numeral Systems as an Accurate Replacement for Huffman Coding — Jarek Duda, Khalid Tahboub, Neeraj J. Gadgil, Edward J. Delp, 2015
https://scholar.google.com/scholar?q=The+Use+of+Asymmetric+Numeral+Systems+as+an+Accurate+Replacement+for+Huffman+Coding
8. Variational Image Compression with a Scale Hyperprior — Johannes Ballé, David Minnen, Saurabh Singh, Sung Jin Hwang, Nick Johnston, 2018
https://scholar.google.com/scholar?q=Variational+Image+Compression+with+a+Scale+Hyperprior
9. Zstandard Compression and the application/zstd Media Type (RFC 8878) — Yann Collet, Murray Kucherawy (eds.), 2020
https://scholar.google.com/scholar?q=Zstandard+Compression+and+the+application%2Fzstd+Media+Type+%28RFC+8878%29
10. GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLMs — Hao Kang, Qingru Zhang, Souvik Kundu, Geonhwa Jeong, Zaoxing Liu, Tushar Krishna, Tuo Zhao, 2024
https://scholar.google.com/scholar?q=GEAR%3A+An+Efficient+KV+Cache+Compression+Recipe+for+Near-Lossless+Generative+Inference+of+LLMs
11. S-LoRA: Serving Thousands of Concurrent LoRA Adapters — Ying Sheng, Shiyi Cao, Dacheng Li, Coleman Hooper, Nicholas Lee, Shuo Yang, Christopher Chou, Banghua Zhu, Lianmin Zheng, Kurt Keutzer, Joseph E. Gonzalez, Ion Stoica, 2023
https://scholar.google.com/scholar?q=S-LoRA%3A+Serving+Thousands+of+Concurrent+LoRA+Adapters
12. QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving — Yujun Lin, Haotian Tang, Shang Yang, Zhekai Zhang, Guangxuan Xiao, Chuang Gan, Song Han, 2024
https://scholar.google.com/scholar?q=QServe%3A+W4A8KV4+Quantization+and+System+Co-design+for+Efficient+LLM+Serving
13. Efficient LLM Inference: Bandwidth, Compute, Synchronization, and Capacity are All You Need — M. Davies, N. Crago, K. Sankaralingam, C. Kozyrakis, 2025
https://scholar.google.com/scholar?q=Efficient+LLM+Inference%3A+Bandwidth%2C+Compute%2C+Synchronization%2C+and+Capacity+are+All+You+Need
Interactive Visualization: Approaching Shannon Bound: Lossless LLM Weight Compression

This episode explores Tent, a 2021 method for adapting a frozen classifier to shifted test data using only entropy minimization on unlabeled target inputs — no retraining, no labels, and no access to the original source dataset. It contrasts this "fully test-time adaptation" setting against classical domain adaptation, which still requires the source data on hand during adjustment, a constraint that's often impractical for vendors shipping models under privacy or bandwidth limits. The discussion digs into the mechanism: Tent re-estimates BatchNorm statistics on incoming test batches and tunes only the tiny per-channel scale-and-shift parameters (under 1% of the network), repurposing existing training infrastructure for adaptation. The hosts also interrogate the core intuition — that confident predictions tend to be correct — pressing on whether that assumption holds when a model's decision boundaries are already unreliable, without fully resolving the tension before turning to results. Listeners interested in low-cost deployment fixes for distribution shift, or skeptical of self-referential confidence-based methods, will find the back-and-forth pushback especially engaging.

Sources:
1. Tent: Fully Test-time Adaptation by Entropy Minimization — Dequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno Olshausen, Trevor Darrell, 2020
http://arxiv.org/abs/2006.10726
2. Semi-Supervised Learning by Entropy Minimization — Yves Grandvalet, Yoshua Bengio, 2004
https://scholar.google.com/scholar?q=Semi-Supervised+Learning+by+Entropy+Minimization
3. A DIRT-T Approach to Unsupervised Domain Adaptation — Rui Shu, Hung Bui, Hirokazu Narui, Stefano Ermon, 2018
https://scholar.google.com/scholar?q=A+DIRT-T+Approach+to+Unsupervised+Domain+Adaptation
4. Test-Time Training with Self-Supervision for Generalization under Distribution Shifts — Yu Sun, Xiaolong Wang, Zhuang Liu, John Miller, Alexei A. Efros, Moritz Hardt, 2020
https://scholar.google.com/scholar?q=Test-Time+Training+with+Self-Supervision+for+Generalization+under+Distribution+Shifts
5. Unsupervised Domain Adaptation by Backpropagation — Yaroslav Ganin, Victor Lempitsky, 2015
https://scholar.google.com/scholar?q=Unsupervised+Domain+Adaptation+by+Backpropagation
6. Learning from Synthetic Data: Addressing Domain Shift for Semantic Segmentation — Swami Sankaranarayanan, Yogesh Balaji, Arpit Jain, Ser Nam Lim, Rama Chellappa, 2018
https://scholar.google.com/scholar?q=Learning+from+Synthetic+Data%3A+Addressing+Domain+Shift+for+Semantic+Segmentation
7. Revisiting Batch Normalization For Practical Domain Adaptation — Yanghao Li, Naiyan Wang, Jianping Shi, Jiaying Liu, Xiaodi Hou, 2016
https://scholar.google.com/scholar?q=Revisiting+Batch+Normalization+For+Practical+Domain+Adaptation
8. CyCADA: Cycle-Consistent Adversarial Domain Adaptation — Judy Hoffman, Eric Tzeng, Taesung Park, Jun-Yan Zhu, Phillip Isola, Kate Saenko, Alexei A. Efros, Trevor Darrell, 2018
https://scholar.google.com/scholar?q=CyCADA%3A+Cycle-Consistent+Adversarial+Domain+Adaptation
9. Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift — Sergey Ioffe, Christian Szegedy, 2015
https://scholar.google.com/scholar?q=Batch+Normalization%3A+Accelerating+Deep+Network+Training+by+Reducing+Internal+Covariate+Shift
10. Evaluating Prediction-Time Batch Normalization for Robustness to Covariate Shift — Zachary Nado, Shreyas Padhy, D. Sculley, Alexander D'Amour, Balaji Lakshminarayanan, Jasper Snoek, 2020
https://scholar.google.com/scholar?q=Evaluating+Prediction-Time+Batch+Normalization+for+Robustness+to+Covariate+Shift
11. Improving Robustness Against Common Corruptions by Covariate Shift Adaptation — Steffen Schneider, Evgenia Rusak, Luisa Eck, Oliver Bringmann, Wieland Brendel, Matthias Bethge, 2020
https://scholar.google.com/scholar?q=Improving+Robustness+Against+Common+Corruptions+by+Covariate+Shift+Adaptation
12. Test-Time Training for Out-of-Distribution Generalization — Yu Sun, Xiaolong Wang, Zhuang Liu, John Miller, Alexei A. Efros, Moritz Hardt, 2019
https://scholar.google.com/scholar?q=Test-Time+Training+for+Out-of-Distribution+Generalization
13. Do We Really Need to Access the Source Data? Source Hypothesis Transfer for Unsupervised Domain Adaptation (SHOT) — Jian Liang, Dapeng Hu, Jiashi Feng, 2020
https://scholar.google.com/scholar?q=Do+We+Really+Need+to+Access+the+Source+Data%3F+Source+Hypothesis+Transfer+for+Unsupervised+Domain+Adaptation+%28SHOT%29
14. Do CIFAR-10 Classifiers Generalize to CIFAR-10? / Do ImageNet Classifiers Generalize to ImageNet? — Benjamin Recht, Rebecca Roelofs, Ludwig Schmidt, Vaishaal Shankar, 2018 / 2019
https://scholar.google.com/scholar?q=Do+CIFAR-10+Classifiers+Generalize+to+CIFAR-10%3F+%2F+Do+ImageNet+Classifiers+Generalize+to+ImageNet%3F
15. A Simple Way to Make Neural Networks Robust Against Diverse Image Corruptions (ANT) — Evgenia Rusak, Lukas Schott, Roland S. Zimmermann, Julian Bitterwolf, Oliver Bringmann, Matthias Bethge, Wieland Brendel, 2020
https://scholar.google.com/scholar?q=A+Simple+Way+to+Make+Neural+Networks+Robust+Against+Diverse+Image+Corruptions+%28ANT%29
Interactive Visualization: Test-Time Adaptation Through Entropy Minimization

This episode explores SOAP, a new optimizer from a Harvard/Kempner Institute team that fuses Shampoo's second-order preconditioning with Adam's update mechanics. The hosts trace the lineage from Adagrad's mathematically ideal but computationally infeasible full preconditioner matrix, through Adam's cheap diagonal approximation, to Shampoo's middle-ground Kronecker-product approach using two smaller per-dimension preconditioners. The core theoretical result discussed is a proof that Shampoo run with the one-half power is mathematically equivalent to running Adafactor inside the eigenbasis Shampoo's own preconditioner defines — which motivates simply swapping in full Adam within that same rotated basis, adding just one new hyperparameter (preconditioning frequency) over standard AdamW. The discussion highlights the paper's striking efficiency claims — over 40% fewer training iterations and 35% less wall-clock time versus AdamW, and roughly 20% better than Shampoo itself — while noting these numbers deserve scrutiny given real-world context like Shampoo's AlgoPerf benchmark win and its use in training Gemini 1.5 Flash. Listeners interested in the mechanics behind large-scale training efficiency will get a clear breakdown of why optimizer choice translates directly into cluster-scale compute costs and calendar time.

Sources:
1. SOAP: Improving and Stabilizing Shampoo using Adam — Nikhil Vyas, Depen Morwani, Rosie Zhao, Mujin Kwun, Itai Shapira, David Brandfonbrener, Lucas Janson, Sham Kakade, 2024
http://arxiv.org/abs/2409.11321
2. Muon: An optimizer for hidden layers in neural networks — Keller Jordan et al., 2024
https://scholar.google.com/scholar?q=Muon%3A+An+optimizer+for+hidden+layers+in+neural+networks
3. 4-bit Shampoo for Memory-Efficient Network Training — Sike Wang, Jia Li, Pan Zhou, Hua Huang, 2024
https://scholar.google.com/scholar?q=4-bit+Shampoo+for+Memory-Efficient+Network+Training
4. Combining axes preconditioners through Kronecker approximation for deep learning — Sai Surya Duvvuri, Fnu Devvrit, Rohan Anil, Cho-Jui Hsieh, Inderjit S. Dhillon, 2024
https://scholar.google.com/scholar?q=Combining+axes+preconditioners+through+Kronecker+approximation+for+deep+learning
5. No train no gain: Revisiting efficient training algorithms for transformer-based language models — Jean Kaddour, Oscar Key, Piotr Nawrot, Pasquale Minervini, Matt J. Kusner, 2023
https://scholar.google.com/scholar?q=No+train+no+gain%3A+Revisiting+efficient+training+algorithms+for+transformer-based+language+models
Interactive Visualization: SOAP: Stabilizing Shampoo's Second-Order Optimizer with Adam

This episode explores a distributed-systems paper from Meta, Mila/University of Montreal, and NVIDIA researchers that tackles a long-standing tradeoff in neural network training: second-order-style optimizers like Shampoo converge better than the ubiquitous diagonal methods (AdaGrad, RMSProp, Adam), but were assumed too computationally expensive to use at scale. The discussion traces the lineage from diagonal AdaGrad through the impractical "full-matrix" AdaGrad to Shampoo's key innovation — approximating each layer's preconditioner as a Kronecker product of two much smaller matrices, drawing a parallel to Martens and Grosse's independently-derived KFAC method. The hosts debate whether Shampoo counts as a true second-order method or something distinct rooted in online convex optimization theory, and highlight the paper's headline systems result: a distributed PyTorch implementation that keeps the wall-clock overhead of this matrix-based optimizer to roughly 10% per step. Listeners interested in the practical engineering behind making theoretically superior optimizers actually usable at billion-parameter scale will find the breakdown of the memory and compute tradeoffs especially compelling.

Sources:
1. A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale — Hao-Jun Michael Shi, Tsung-Hsien Lee, Shintaro Iwasaki, Jose Gallego-Posada, Zhijing Li, Kaushik Rangadurai, Dheevatsa Mudigere, Michael Rabbat, 2023
http://arxiv.org/abs/2309.06497
2. Adaptive Subgradient Methods for Online Learning and Stochastic Optimization — John Duchi, Elad Hazan, Yoram Singer, 2011
https://scholar.google.com/scholar?q=Adaptive+Subgradient+Methods+for+Online+Learning+and+Stochastic+Optimization
3. Shampoo: Preconditioned Stochastic Tensor Optimization — Vineet Gupta, Tomer Koren, Yoram Singer, 2018
https://scholar.google.com/scholar?q=Shampoo%3A+Preconditioned+Stochastic+Tensor+Optimization
4. Scalable Second Order Optimization for Deep Learning — Rohan Anil, Vineet Gupta, Tomer Koren, Kevin Regan, Yoram Singer (Google), 2020
https://scholar.google.com/scholar?q=Scalable+Second+Order+Optimization+for+Deep+Learning
5. Optimizing Neural Networks with Kronecker-factored Approximate Curvature — James Martens, Roger Grosse, 2015
https://scholar.google.com/scholar?q=Optimizing+Neural+Networks+with+Kronecker-factored+Approximate+Curvature
6. Optimizing Neural Networks with Kronecker-factored Approximate Curvature (K-FAC) — James Martens, Roger Grosse, 2015
https://scholar.google.com/scholar?q=Optimizing+Neural+Networks+with+Kronecker-factored+Approximate+Curvature+%28K-FAC%29
7. On the Factory Floor: ML Engineering for Industrial-Scale Ads Recommendation Models — Rohan Anil, Sandra Gadanho, Da Huang, et al., 2022
https://scholar.google.com/scholar?q=On+the+Factory+Floor%3A+ML+Engineering+for+Industrial-Scale+Ads+Recommendation+Models
8. Adafactor: Adaptive Learning Rates with Sublinear Memory Cost — Noam Shazeer, Mitchell Stern, 2018
https://scholar.google.com/scholar?q=Adafactor%3A+Adaptive+Learning+Rates+with+Sublinear+Memory+Cost
9. SOAP: Improving and Stabilizing Shampoo using Adam — Nikhil Vyas, Depen Morwani, Rosie Zhao, et al., 2024
https://scholar.google.com/scholar?q=SOAP%3A+Improving+and+Stabilizing+Shampoo+using+Adam
Interactive Visualization: Distributed Shampoo: Making Second-Order Optimization Practical at Scale

This episode explores the paper "Scalable Second Order Optimization for Deep Learning" and its introduction of Distributed Shampoo, a Kronecker-factored second-order optimizer that cuts training steps in half compared to a well-tuned Adam baseline on WMT'14 English-to-French translation, with further gains shown on BERT, Criteo click-through-rate modeling, and ResNet-50. The discussion traces the lineage from Newton's method and full-matrix AdaGrad's prohibitive O(N²)/O(N³) costs through Shampoo's Kronecker-product approximation, which replaces one massive preconditioner with smaller per-dimension matrices to make second-order optimization tractable at scale. Rival factored approaches, K-FAC and K-BFGS, are positioned as points of comparison throughout the paper. The conversation is aimed at listeners who use Adam daily but have never unpacked the mechanical distinction between first-order and second-order optimization, or the meaning of "preconditioning" itself. It's a compelling listen because second-order methods have long been dismissed as theoretically superior but practically unscalable, and this paper demonstrates real wall-clock wins across four distinct production-scale workloads rather than a single cherry-picked benchmark.

Sources:
1. Scalable Second Order Optimization for Deep Learning — Rohan Anil, Vineet Gupta, Tomer Koren, Kevin Regan, Yoram Singer, 2020
http://arxiv.org/abs/2002.09018
2. Shampoo: Preconditioned Stochastic Tensor Optimization — Vineet Gupta, Tomer Koren, Yoram Singer, 2018
https://scholar.google.com/scholar?q=Shampoo%3A+Preconditioned+Stochastic+Tensor+Optimization
3. Optimizing Neural Networks with Kronecker-factored Approximate Curvature (K-FAC) — James Martens, Roger Grosse, 2015
https://scholar.google.com/scholar?q=Optimizing+Neural+Networks+with+Kronecker-factored+Approximate+Curvature+%28K-FAC%29
4. Adam: A Method for Stochastic Optimization — Diederik P. Kingma, Jimmy Ba, 2014
https://scholar.google.com/scholar?q=Adam%3A+A+Method+for+Stochastic+Optimization
5. Adaptive Subgradient Methods for Online Learning and Stochastic Optimization (AdaGrad) — John Duchi, Elad Hazan, Yoram Singer, 2011
https://scholar.google.com/scholar?q=Adaptive+Subgradient+Methods+for+Online+Learning+and+Stochastic+Optimization+%28AdaGrad%29
6. Optimizing Neural Networks with Kronecker-factored Approximate Curvature — J. Martens, R. Grosse, 2015
https://scholar.google.com/scholar?q=Optimizing+Neural+Networks+with+Kronecker-factored+Approximate+Curvature
7. Distributed Second-Order Optimization using Kronecker-Factored Approximations — J. Ba, J. Martens, R. Grosse, 2017
https://scholar.google.com/scholar?q=Distributed+Second-Order+Optimization+using+Kronecker-Factored+Approximations
8. Large-Scale Distributed Second-Order Optimization Using Kronecker-Factored Approximate Curvature for Deep Convolutional Neural Networks — K. Osawa, Y. Tsuji, Y. Ueno, A. Naruse, R. Yokota, S. Matsuoka, 2019
https://scholar.google.com/scholar?q=Large-Scale+Distributed+Second-Order+Optimization+Using+Kronecker-Factored+Approximate+Curvature+for+Deep+Convolutional+Neural+Networks
9. Large Batch Optimization for Deep Learning: Training BERT in 76 Minutes — Y. You, J. Li, S. Reddi, et al. (LAMB), 2019
https://scholar.google.com/scholar?q=Large+Batch+Optimization+for+Deep+Learning%3A+Training+BERT+in+76+Minutes
10. A Large Batch Optimizer Reality Check: Traditional, Generic Optimizers Suffice Across Batch Sizes — Z. Nado, J. Gilmer, C. Shallue, R. Anil, G. Dahl, 2021
https://scholar.google.com/scholar?q=A+Large+Batch+Optimizer+Reality+Check%3A+Traditional%2C+Generic+Optimizers+Suffice+Across+Batch+Sizes
11. Limitations of the Empirical Fisher Approximation for Natural Gradient Descent — F. Kunstner, P. Hennig, L. Balles, 2019
https://scholar.google.com/scholar?q=Limitations+of+the+Empirical+Fisher+Approximation+for+Natural+Gradient+Descent
Interactive Visualization: Scalable Second-Order Optimization: Shampoo at Scale

This episode explores Shampoo, a 2018 optimization algorithm from Google Brain and Princeton that brings second-order curvature information to neural network training without the prohibitive cost of a full preconditioner. The discussion traces the lineage from Newton's method through AdaGrad's diagonal and full-matrix variants, explaining how Shampoo exploits the natural tensor structure of neural network weights — keeping separate small preconditioners per dimension and combining them via Kronecker products rather than materializing an impossibly large matrix. The hosts debate how much weight to put on the algorithm's convergence proofs, which rely on online convex optimization theory even though real network training is highly non-convex, concluding that the theory offers a sanity check rather than a guarantee and that empirical performance does the real work of justification. Along the way, they place Shampoo alongside K-FAC as one of the major structure-aware curvature approximations in the optimization literature. Listeners interested in the tradeoffs between cheap first-order methods and expensive second-order ones will find a clear walkthrough of how Shampoo threads that needle.

Sources:
1. Shampoo: Preconditioned Stochastic Tensor Optimization — Vineet Gupta, Tomer Koren, Yoram Singer, 2018
http://arxiv.org/abs/1802.09568
2. Adaptive Subgradient Methods for Online Learning and Stochastic Optimization — John Duchi, Elad Hazan, Yoram Singer, 2011
https://scholar.google.com/scholar?q=Adaptive+Subgradient+Methods+for+Online+Learning+and+Stochastic+Optimization
3. Scalable Second Order Optimization for Deep Learning (Distributed Shampoo) — Rohan Anil, Vineet Gupta, Tomer Koren, Kevin Regan, Yoram Singer, 2021
https://scholar.google.com/scholar?q=Scalable+Second+Order+Optimization+for+Deep+Learning+%28Distributed+Shampoo%29
4. SOAP: Improving and Stabilizing Shampoo using Adam — Nikhil Vyas, Depen Morwani, Rosie Zhao, Itai Shapira, David Brandfonbrener, Sham Kakade, 2024
https://scholar.google.com/scholar?q=SOAP%3A+Improving+and+Stabilizing+Shampoo+using+Adam
5. Muon: An optimizer for hidden layers in neural networks — Keller Jordan et al., 2024
https://scholar.google.com/scholar?q=Muon%3A+An+optimizer+for+hidden+layers+in+neural+networks
6. Adam: A Method for Stochastic Optimization — Diederik Kingma, Jimmy Ba, 2015
https://scholar.google.com/scholar?q=Adam%3A+A+Method+for+Stochastic+Optimization
7. Decoupled Weight Decay Regularization — Ilya Loshchilov, Frank Hutter, 2017
https://scholar.google.com/scholar?q=Decoupled+Weight+Decay+Regularization
8. Optimizing Neural Networks with Kronecker-factored Approximate Curvature — James Martens, Roger Grosse, 2015
https://scholar.google.com/scholar?q=Optimizing+Neural+Networks+with+Kronecker-factored+Approximate+Curvature
9. A Stochastic Quasi-Newton Method for Large-Scale Optimization — Richard Byrd, Samantha Hansen, Jorge Nocedal, Yoran Singer, 2016
https://scholar.google.com/scholar?q=A+Stochastic+Quasi-Newton+Method+for+Large-Scale+Optimization
10. Shampoo: Preconditioned Stochastic Tensor Optimization — Vineet Gupta, Tomer Koren, Yoram Singer, 2018
https://scholar.google.com/scholar?q=Shampoo%3A+Preconditioned+Stochastic+Tensor+Optimization
11. Online Convex Programming and Generalized Infinitesimal Gradient Ascent — Martin Zinkevich, 2003
https://scholar.google.com/scholar?q=Online+Convex+Programming+and+Generalized+Infinitesimal+Gradient+Ascent
12. Online Learning and Online Convex Optimization — Shai Shalev-Shwartz, 2012
https://scholar.google.com/scholar?q=Online+Learning+and+Online+Convex+Optimization
13. Attention Is All You Need — A. Vaswani, N. Shazeer, N. Parmar, et al., 2017
https://scholar.google.com/scholar?q=Attention+Is+All+You+Need
14. A Unified Approach to Adaptive Regularization in Online and Stochastic Optimization — V. Gupta, T. Koren, Y. Singer, 2017
https://scholar.google.com/scholar?q=A+Unified+Approach+to+Adaptive+Regularization+in+Online+and+Stochastic+Optimization
15. Understanding Deep Learning Requires Rethinking Generalization — C. Zhang, S. Bengio, M. Hardt, B. Recht, O. Vinyals, 2017
https://scholar.google.com/scholar?q=Understanding+Deep+Learning+Requires+Rethinking+Generalization
Interactive Visualization: Preconditioned Optimization Without the Full Matrix Cost

This episode explores a paper proposing SCALENET, a hypernetwork-based fix for unsupervised test-time adaptation (TTA) in large language models. It examines why naive per-prompt gradient updates are unstable — a 70-billion-parameter Llama model's negative log-likelihood balloons from 2.21 to 11.49 after just five adaptation steps — and traces the problem to high-variance single-sample gradients that can't average out the way batch training does. The discussion covers the constrained "adapt-and-reset" setup used in real deployment, where models take a few unsupervised gradient steps on LoRA attention matrices per prompt before discarding the update, and explains why a single global learning rate can't work when small rates do nothing and large ones destroy the model. Listeners interested in the mechanics of on-the-fly model adaptation, LoRA-based efficient tuning, and the control-theory-like challenge of stabilizing per-layer, per-step learning rates will find the breakdown of the failure modes and the proposed hypernetwork solution especially compelling.

Sources:
1. Unsupervised Layer-Wise Dynamic Test Time Adaptation for LLMs — Longhuan Xu, Cunjian Chen, Feng Yin, 2026
http://arxiv.org/abs/2602.09719
2. Test-Time Training with Self-Supervision for Generalization under Distribution Shifts — Yu Sun, Xiaolong Wang, Zhuang Liu, John Miller, Alexei Efros, Moritz Hardt, 2020
https://scholar.google.com/scholar?q=Test-Time+Training+with+Self-Supervision+for+Generalization+under+Distribution+Shifts
3. Tent: Fully Test-Time Adaptation by Entropy Minimization — Dequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno Olshausen, Trevor Darrell, 2021
https://scholar.google.com/scholar?q=Tent%3A+Fully+Test-Time+Adaptation+by+Entropy+Minimization
4. Test-Time Training on Nearest Neighbors for Large Language Models — Moritz Hardt, Yu Sun, 2024
https://scholar.google.com/scholar?q=Test-Time+Training+on+Nearest+Neighbors+for+Large+Language+Models
5. The Surprising Effectiveness of Test-Time Training for Abstract Reasoning — Ekin Akyürek, Mehul Damani, Linlu Qiu, Han Guo, Yoon Kim, Jacob Andreas, 2024
https://scholar.google.com/scholar?q=The+Surprising+Effectiveness+of+Test-Time+Training+for+Abstract+Reasoning
6. Test-time Learning for Large Language Models — Hu, J., Zhang, Z., Chen, G., Wen, X., Shuai, C., Luo, W., Xiao, B., Li, Y., Tan, M., 2025
https://scholar.google.com/scholar?q=Test-time+Learning+for+Large+Language+Models
7. SLOT: Sample-specific Language Model Optimization at Test-time — Hu, Y., Zhang, X., Fang, X., Chen, Z., Wang, X., Zhang, H., Qi, G., 2025
https://scholar.google.com/scholar?q=SLOT%3A+Sample-specific+Language+Model+Optimization+at+Test-time
8. COME: Test-time Adaption by Conservatively Minimizing Entropy — Zhang, Q., Bian, Y., Kong, X., Zhao, P., Zhang, C., 2024
https://scholar.google.com/scholar?q=COME%3A+Test-time+Adaption+by+Conservatively+Minimizing+Entropy
9. Revisiting Dynamic Evaluation: Online Adaptation for Large Language Models — Rannen-Triki, A., Bornschein, J., Pascanu, R., Hutter, M., et al., 2024
https://scholar.google.com/scholar?q=Revisiting+Dynamic+Evaluation%3A+Online+Adaptation+for+Large+Language+Models
10. Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks (MAML) — Finn, C., Abbeel, P., Levine, S., 2017
https://scholar.google.com/scholar?q=Model-Agnostic+Meta-Learning+for+Fast+Adaptation+of+Deep+Networks+%28MAML%29
Interactive Visualization: Naive Test-Time Adaptation Destabilizes LLM Predictions

This episode revisits Teuvo Kohonen's 1972 paper "Correlation Matrix Memories," which reframes associative memory as a hardware fault-tolerance problem rather than a representation-learning one. Kohonen builds a memory from outer-product sums of key and data vectors, then shows mathematically how much recall quality degrades when connections are randomly dropped (an "incomplete" correlation matrix memory) rather than fully wired. The discussion traces the paper's lineage against optical holography models and Steinbuch's Lernmatrix, and unpacks concepts like crosstalk and graceful degradation as information gets smeared additively across the matrix instead of stored in one fragile spot. A tangent draws — and partly disputes — a comparison between Kohonen's outer-product accumulation and the mechanics underlying modern attention, debating whether the resemblance is structural or purely coincidental given the total absence of learning or gradients in the original scheme. Listeners interested in the deep history of neural memory models and how old hardware constraints shaped ideas that echo in today's architectures will find plenty to chew on.

Sources:
1. Kohonen's 1972 Correlation Matrix Memory, Decades Before Attention
https://lucidar.me/fr/neural-networks/files/1972-correlation-matrix-memories.pdf
2. Neural networks and physical systems with emergent collective computational abilities — John J. Hopfield, 1982
https://scholar.google.com/scholar?q=Neural+networks+and+physical+systems+with+emergent+collective+computational+abilities
3. Non-Holographic Associative Memory — David Willshaw, O. P. Buneman, H. Christopher Longuet-Higgins, 1969
https://scholar.google.com/scholar?q=Non-Holographic+Associative+Memory
4. Hopfield Networks is All You Need — Hubert Ramsauer, Bernhard Schäfl, Johannes Lehner, et al., 2020
https://scholar.google.com/scholar?q=Hopfield+Networks+is+All+You+Need
5. Linear Transformers Are Secretly Fast Weight Programmers — Imanol Schlag, Kazuki Irie, Jürgen Schmidhuber, 2021
https://scholar.google.com/scholar?q=Linear+Transformers+Are+Secretly+Fast+Weight+Programmers
6. A Simple Neural Network Generating an Interactive Memory — James A. Anderson, 1972
https://scholar.google.com/scholar?q=A+Simple+Neural+Network+Generating+an+Interactive+Memory
7. Representation of Associated Data by Matrix Operators — Teuvo Kohonen, Matti Ruohonen, 1973
https://scholar.google.com/scholar?q=Representation+of+Associated+Data+by+Matrix+Operators
8. Sparse Distributed Memory — Pentti Kanerva, 1988
https://scholar.google.com/scholar?q=Sparse+Distributed+Memory
9. Die Lernmatrix — K. Steinbuch, 1961
https://scholar.google.com/scholar?q=Die+Lernmatrix
10. Associative holographic memories — D. Gabor, 1969
https://scholar.google.com/scholar?q=Associative+holographic+memories
11. A class of randomly organized associative memories — T. Kohonen, 1971
https://scholar.google.com/scholar?q=A+class+of+randomly+organized+associative+memories
Interactive Visualization: Kohonen's 1972 Correlation Matrix Memory, Decades Before Attention

This episode explores Model-Agnostic Meta-Learning (MAML), the 2017 approach from Chelsea Finn, Pieter Abbeel, and Sergey Levine that trains a single, architecture-agnostic initialization capable of fast adaptation across image classification, regression, and reinforcement learning. Rather than learning a task-specific update rule like earlier recurrent meta-learners, MAML optimizes the starting weights themselves so that a few steps of ordinary gradient descent adapt them well to a brand-new task from minimal data, tested through Omniglot and MiniImagenet few-shot classification, sinusoid regression, and MuJoCo/2D navigation RL. The discussion breaks down the inner-loop/outer-loop structure, the second-order gradient-through-gradient math (Hessian-vector products) needed to backpropagate through the adaptation step, and how finite-difference approximations sidestep the third-derivative problem when TRPO is used as the RL meta-optimizer. Listeners get a clear walkthrough of N-way K-shot learning and why one image per class is such an extreme test of generalization, plus a grounded comparison to the more familiar pretrain-then-fine-tune workflow. It's a good listen for anyone curious how a deceptively simple idea — learn to be easy to fine-tune — unified meta-learning across problem types that previously required separate specialized systems.

Sources:
1. Model-Agnostic Meta-Learning for Fast Task Adaptation
https://proceedings.mlr.press/v70/finn17a/finn17a.pdf
2. Optimization as a Model for Few-Shot Learning — Sachin Ravi, Hugo Larochelle, 2017
https://scholar.google.com/scholar?q=Optimization+as+a+Model+for+Few-Shot+Learning
3. Meta-Learning with Memory-Augmented Neural Networks — Adam Santoro, Sergey Bartunov, Matthew Botvinick, Daan Wierstra, Timothy Lillicrap, 2016
https://scholar.google.com/scholar?q=Meta-Learning+with+Memory-Augmented+Neural+Networks
4. Matching Networks for One Shot Learning — Oriol Vinyals, Charles Blundell, Tim Lillicrap, Daan Wierstra, 2016
https://scholar.google.com/scholar?q=Matching+Networks+for+One+Shot+Learning
5. Learning to Learn by Gradient Descent by Gradient Descent — Marcin Andrychowicz et al., 2016
https://scholar.google.com/scholar?q=Learning+to+Learn+by+Gradient+Descent+by+Gradient+Descent
6. Trust Region Policy Optimization — John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, Philipp Moritz, 2015
https://scholar.google.com/scholar?q=Trust+Region+Policy+Optimization
7. RL2: Fast Reinforcement Learning via Slow Reinforcement Learning — Yan Duan, John Schulman, Xi Chen, Peter L. Bartlett, Ilya Sutskever, Pieter Abbeel, 2016
https://scholar.google.com/scholar?q=RL2%3A+Fast+Reinforcement+Learning+via+Slow+Reinforcement+Learning
Interactive Visualization: Model-Agnostic Meta-Learning for Fast Task Adaptation

This episode explores a new test-time training method called In-Place TTT, which repurposes the down-projection matrix inside a model's existing gated MLP as adaptable "fast weights," letting a pretrained model keep learning during inference without any architectural changes. A key innovation is replacing the reconstruction-style training target used in prior TTT approaches with an LM-aligned target built from a causal convolution over token embeddings, which the authors prove (via an induction-head theorem) actually raises the probability of the correct next token. The discussion covers how a context-parallel scan preserves causality while enabling parallel computation of these updates, and walks through benchmark results showing the method trailing a baseline at short context but pulling substantially ahead as sequence length grows, tested across Qwen3-4B, LLaMA-3.1-8B, and Qwen3-14B. The hosts also dig into an ablation showing that mid-sized chunk sizes outperform larger ones — a counterintuitive result tied to how often the fast weights get to update rather than raw parallelism — plus efficiency data showing the approach barely affects throughput or memory. It's a concrete look at how far you can push adaptive inference-time learning while reusing a model's own existing structure.

Sources:
1. In-Place Test-Time Training — Guhao Feng, Shengjie Luo, Kai Hua, Ge Zhang, Di He, Wenhao Huang, Tianle Cai, 2026
http://arxiv.org/abs/2604.06169
2. Learning to (Learn at Test Time): RNNs with Expressive Hidden States — Yu Sun, Xinhao Li, Karan Dalal, et al., 2024
https://scholar.google.com/scholar?q=Learning+to+%28Learn+at+Test+Time%29%3A+RNNs+with+Expressive+Hidden+States
3. Test-Time Training Done Right (LaCT) — Tianyuan Zhang, Sai Bi, Yicong Hong, et al., 2025
https://scholar.google.com/scholar?q=Test-Time+Training+Done+Right+%28LaCT%29
4. Titans: Learning to Memorize at Test Time — Ali Behrouz, Peilin Zhong, Vahab Mirrokni, 2024
https://scholar.google.com/scholar?q=Titans%3A+Learning+to+Memorize+at+Test+Time
5. Transformer Feed-Forward Layers Are Key-Value Memories — Mor Geva, Roei Schuster, Jonathan Berant, Omer Levy, 2020
https://scholar.google.com/scholar?q=Transformer+Feed-Forward+Layers+Are+Key-Value+Memories
6. Locating and Editing Factual Associations in GPT (ROME) — Kevin Meng, David Bau, Alex Andonian, Yonatan Belinkov, 2022
https://scholar.google.com/scholar?q=Locating+and+Editing+Factual+Associations+in+GPT+%28ROME%29
7. LoRA: Low-Rank Adaptation of Large Language Models — Edward Hu, Yelong Shen, Phillip Wallis, et al., 2022
https://scholar.google.com/scholar?q=LoRA%3A+Low-Rank+Adaptation+of+Large+Language+Models
Interactive Visualization: In-Place Test-Time Training Turns Fast Weights Into Online Memory

This episode explores a paper examining what happens to continual learning problems when LLM agents shift from parametric updates to memory-augmented architectures. Rather than accepting the industry assumption that external memory sidesteps catastrophic forgetting entirely, the researchers run classic continual-learning protocols on memory-based agents and find the same core problem resurfaces in a new form — shifting from parameter capacity to context-window retrieval capacity. They identify three specific failure modes: retrieval pollution (irrelevant memories crowding the prompt), context competition (useful memories getting displaced by other retrieved items), and memory dilution (relevant material becoming harder to surface as the memory store grows). The discussion traces this argument against the history of catastrophic forgetting and prior mitigation techniques like Elastic Weight Consolidation and Gradient Episodic Memory, then explains how the paper reframes the stability-plasticity dilemma for retrieval-based systems. Listeners interested in agent design, RAG architectures, or the assumptions underlying memory-augmented LLMs will find the paper's reframing — that memory doesn't eliminate the bottleneck, it just relocates it — a useful corrective to a widely repeated industry pitch.

Sources:
1. Hal:
Memory-Augmented LLM Agents Still Hit Continual Learning's Wall
https://arxiv.org/pdf/2604.27003
2. A-Mem: Agentic Memory for LLM Agents — Wujiang Xu, Zujie Liang, Kai Mei, Hang Gao, Juntao Tan, Yongfeng Zhang, 2025
https://scholar.google.com/scholar?q=A-Mem%3A+Agentic+Memory+for+LLM+Agents
3. MemGPT: Towards LLMs as Operating Systems — Charles Packer, Vivian Fang, Shishir G. Patil, Kevin Lin, Sarah Wooders, Joseph E. Gonzalez, 2023
https://scholar.google.com/scholar?q=MemGPT%3A+Towards+LLMs+as+Operating+Systems
4. How Memory Management Impacts LLM Agents: An Empirical Study of Experience-Following Behavior — Zidi Xiong, Yuping Lin, Wenya Xie, Pengfei He, Zirui Liu, Jiliang Tang, Himabindu Lakkaraju, Zhen Xiang, 2025
https://scholar.google.com/scholar?q=How+Memory+Management+Impacts+LLM+Agents%3A+An+Empirical+Study+of+Experience-Following+Behavior
5. From RAG to Memory: Non-Parametric Continual Learning for Large Language Models — Bernal Jiménez Gutiérrez, Yiheng Shu, Weijian Qi, Sizhe Zhou, Yu Su, 2025
https://scholar.google.com/scholar?q=From+RAG+to+Memory%3A+Non-Parametric+Continual+Learning+for+Large+Language+Models
6. The Probabilistic Relevance Framework: BM25 and Beyond — Stephen Robertson, Hugo Zaragoza, 2009
https://scholar.google.com/scholar?q=The+Probabilistic+Relevance+Framework%3A+BM25+and+Beyond
Interactive Visualization: Hal:
Memory-Augmented LLM Agents Still Hit Continual Learning's Wall

This episode explores a survey on continual learning in large language models, examining how models can be updated after pretraining without the prohibitive cost of full retraining or the risk of catastrophic forgetting — the phenomenon where new training quietly degrades performance on tasks a model previously handled well. The discussion breaks down the problem across three distinct LLM training stages (pretraining, fine-tuning, and alignment) and maps them onto three classical mitigation strategies: rehearsal-based methods that replay old data, regularization-based methods that penalize changes to critical parameters, and architecture-based methods that add task-specific capacity like adapters or LoRA modules while freezing the rest. The hosts debate the survey's core organizational claim — that structuring the literature by mechanism rather than by application domain (medical, legal, financial) offers a more useful lens for practitioners trying to borrow a specific forgetting-mitigation technique. Listeners interested in the practical tradeoffs of keeping frontier models current — especially around data that can never legally enter a pretraining corpus, like medical or financial records — will find this a grounded framing of a problem every deployed LLM eventually faces.

Sources:
1. Continual Learning in Large Language Models: Methods, Challenges, and Opportunities — Hongyang Chen, Zhongwu Sun, Hongfei Ye, Kunchi Li, Xuemin Lin, 2026
http://arxiv.org/abs/2603.12658
2. Don't Stop Pretraining: Adapt Language Models to Domains and Tasks — Suchin Gururangan, Ana Marasović, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, Noah A. Smith, 2020
https://scholar.google.com/scholar?q=Don%27t+Stop+Pretraining%3A+Adapt+Language+Models+to+Domains+and+Tasks
3. Simple and Scalable Strategies to Continually Pre-train Large Language Models — Adam Ibrahim, Benjamin Thérien, Kshitij Gupta, Mats L. Richter, Quentin Anthony, Timothée Lesort, Eugene Belilovsky, Irina Rish, 2024
https://scholar.google.com/scholar?q=Simple+and+Scalable+Strategies+to+Continually+Pre-train+Large+Language+Models
4. LLaMA Pro: Progressive LLaMA with Block Expansion — Chengyue Wu, Yukang Gan, Yixiao Ge, Zeyu Lu, Jiahao Wang, Ye Feng, Ying Shan, Ping Luo, 2024
https://scholar.google.com/scholar?q=LLaMA+Pro%3A+Progressive+LLaMA+with+Block+Expansion
5. Code Llama: Open Foundation Models for Code — Baptiste Rozière, Jonas Gehring, Fabian Gloeckle, and the Code Llama team at Meta AI, 2023
https://scholar.google.com/scholar?q=Code+Llama%3A+Open+Foundation+Models+for+Code
6. Editing Models with Task Arithmetic — Gabriel Ilharco, Marco Tulio Ribeiro, Mitchell Wortsman, Suchin Gururangan, Ludwig Schmidt, Hannaneh Hajishirzi, Ali Farhadi, 2023
https://scholar.google.com/scholar?q=Editing+Models+with+Task+Arithmetic
7. TIES-Merging: Resolving Interference When Merging Models — Prateek Yadav, Derek Tam, Leshem Choshen, Colin Raffel, Mohit Bansal, 2023
https://scholar.google.com/scholar?q=TIES-Merging%3A+Resolving+Interference+When+Merging+Models
8. Overcoming Catastrophic Forgetting in Neural Networks (EWC) — James Kirkpatrick et al., 2017
https://scholar.google.com/scholar?q=Overcoming+Catastrophic+Forgetting+in+Neural+Networks+%28EWC%29
Interactive Visualization: Continual Learning in LLMs: Beyond Catastrophic Forgetting

This episode explores catastrophic forgetting and plasticity loss in RL-trained language models, and introduces "Fast-Slow Training," a method combining slow weight updates (RLVR) with fast in-context learning to address both. The hosts unpack the distinction between RLVR's automatic, verifiable rewards and traditional RLHF, then dig into two separate failure modes of pure RL post-training: models forgetting general competence while chasing a narrow reward signal, and a subtler loss of plasticity where updates leave models increasingly unable to absorb new tasks. Framing the two training channels as a System 1/System 2 split, the discussion centers on the paper's headline result — combining both channels reaches RL's peak accuracy with up to three times fewer samples, drifts up to seventy percent less from the base model, and preserves the capacity to learn subsequent tasks where pure RL stalls. Listeners interested in the mechanics and tradeoffs of continual learning in large language models will find a grounded walkthrough of why prompting alone hits a ceiling and why weight updates alone come with hidden costs.

Sources:
1. Learning, Fast and Slow: Towards LLMs That Adapt Continually — Rishabh Tiwari, Kusha Sareen, Lakshya A Agrawal, Joseph E. Gonzalez, Matei Zaharia, Kurt Keutzer, Inderjit S Dhillon, Rishabh Agarwal, Devvrit Khatri, 2026
http://arxiv.org/abs/2605.12484
2. Loss of Plasticity in Deep Continual Learning — Shibhansh Dohare, J. Fernando Hernandez-Garcia, Qingfeng Lan, A. Rupam Mahmood, Richard S. Sutton, et al., 2024 (Nature; preprint circulated as 'Maintaining Plasticity via Continual Backprop' from 2021)
https://scholar.google.com/scholar?q=Loss+of+Plasticity+in+Deep+Continual+Learning
3. On Warm-Starting Neural Network Training — Jordan T. Ash, Ryan P. Adams, 2020 (NeurIPS)
https://scholar.google.com/scholar?q=On+Warm-Starting+Neural+Network+Training
4. The Primacy Bias in Deep Reinforcement Learning — Evgenii Nikishin, Max Schwarzer, Pierluca D'Oro, Pierre-Luc Bacon, Aaron Courville, 2022 (ICML)
https://scholar.google.com/scholar?q=The+Primacy+Bias+in+Deep+Reinforcement+Learning
5. Understanding Plasticity in Neural Networks — Clare Lyle, Zeyu Zheng, Evgenii Nikishin, Bernardo Avila Pires, Razvan Pascanu, Will Dabney, 2023 (ICML)
https://scholar.google.com/scholar?q=Understanding+Plasticity+in+Neural+Networks
6. RL's razor: Why online reinforcement learning forgets less — Idan Shenfeld, Jyothish Pari, Pulkit Agrawal, 2025
https://scholar.google.com/scholar?q=RL%27s+razor%3A+Why+online+reinforcement+learning+forgets+less
7. The Art of Scaling Reinforcement Learning Compute for LLMs (ScaleRL) — Devvrit Khatri, Lovish Madaan, Rishabh Tiwari, Rachit Bansal, Sai Surya Duvvuri, Manzil Zaheer, Inderjit S. Dhillon, David Brandfonbrener, Rishabh Agarwal, 2025
https://scholar.google.com/scholar?q=The+Art+of+Scaling+Reinforcement+Learning+Compute+for+LLMs+%28ScaleRL%29
8. Fine-tuning and prompt optimization: Two great steps that work better together (BetterTogether) — Dilara Soylu, Christopher Potts, Omar Khattab, 2024
https://scholar.google.com/scholar?q=Fine-tuning+and+prompt+optimization%3A+Two+great+steps+that+work+better+together+%28BetterTogether%29
9. Mitigating plasticity loss in continual reinforcement learning by reducing churn — Hongyao Tang, Johan Obando-Ceron, Pablo Samuel Castro, Aaron Courville, Glen Berseth, 2025
https://scholar.google.com/scholar?q=Mitigating+plasticity+loss+in+continual+reinforcement+learning+by+reducing+churn
10. What can you do when you have zero rewards during RL? — Jatin Prakash, Anirudh Buvanesh, 2025
https://scholar.google.com/scholar?q=What+can+you+do+when+you+have+zero+rewards+during+RL%3F
Interactive Visualization: Learning, Fast and Slow: LLMs That Adapt Without Forgetting

This episode explores Kohei Honda's tutorial and survey "Model Predictive Control via Probabilistic Inference," which unifies two decades of scattered research—path integral control, reinforcement learning theory, and variational inference—into a single coherent framework called PI-MPC. The discussion traces why classical gradient- and Hessian-based MPC solvers break down on contact-rich robotics, learned neural dynamics, or discontinuous costs, and why the resulting fallback to naive random-shooting sampling collapses under the curse of dimensionality. The core argument is that reframing sampling-based MPC as inference over a distribution of good control sequences—rather than search for a single optimum—yields dramatic gains in sample efficiency and parallelizability, with MPPI's Boltzmann-weighted, temperature-controlled posterior serving as the paper's central worked example. Along the way, the hosts debate whether "inference" is meaningfully different from optimization, tracing how entropy terms in algorithms like Soft Actor-Critic emerge naturally from the probabilistic framing rather than being added as an exploration hack. Listeners interested in robotics, control theory, or the mathematical bridges between classical control and modern probabilistic ML will find the episode's account of why this synthesis only became practical with GPU-scale parallel rollouts particularly compelling.

Sources:
1. Model Predictive Control via Probabilistic Inference: A Tutorial and Survey — Kohei Honda, 2025
http://arxiv.org/abs/2511.08019v4
2. Constrained Model Predictive Control: Stability and Optimality — D. Q. Mayne, J. B. Rawlings, C. V. Rao, P. O. M. Scokaert, 2000
https://scholar.google.com/scholar?q=Constrained+Model+Predictive+Control%3A+Stability+and+Optimality
3. A Survey of Industrial Model Predictive Control Technology — S. Joe Qin, Thomas A. Badgwell, 2003
https://scholar.google.com/scholar?q=A+Survey+of+Industrial+Model+Predictive+Control+Technology
4. Model Predictive Control: Theory and Practice — A Survey — Carlos E. Garcia, David M. Prett, Manfred Morari, 1989
https://scholar.google.com/scholar?q=Model+Predictive+Control%3A+Theory+and+Practice+%E2%80%94+A+Survey
5. Model Predictive Path Integral Control using Covariance Variable Importance Sampling — Grady Williams, Andrew Aldrich, Evangelos A. Theodorou, 2015
https://scholar.google.com/scholar?q=Model+Predictive+Path+Integral+Control+using+Covariance+Variable+Importance+Sampling
6. Information Theoretic MPC for Model-Based Reinforcement Learning — Grady Williams, Nolan Wagener, Brian Goldfain, Paul Drews, James M. Rehg, Byron Boots, Evangelos A. Theodorou, 2017
https://scholar.google.com/scholar?q=Information+Theoretic+MPC+for+Model-Based+Reinforcement+Learning
7. Robust Sampling Based Model Predictive Control with Sparse Objective Information — Grady Williams, Brian Goldfain, Paul Drews, Kamil Saigol, James M. Rehg, Evangelos A. Theodorou, 2018
https://scholar.google.com/scholar?q=Robust+Sampling+Based+Model+Predictive+Control+with+Sparse+Objective+Information
8. Reinforcement Learning and Control as Probabilistic Inference: Tutorial and Review — Sergey Levine, 2018
https://scholar.google.com/scholar?q=Reinforcement+Learning+and+Control+as+Probabilistic+Inference%3A+Tutorial+and+Review
9. Robot Trajectory Optimization using Approximate Inference — Marc Toussaint, 2009
https://scholar.google.com/scholar?q=Robot+Trajectory+Optimization+using+Approximate+Inference
10. Optimal Control as a Graphical Model Inference Problem — Hilbert J. Kappen, Vicenç Gómez, Manfred Opper, 2012
https://scholar.google.com/scholar?q=Optimal+Control+as+a+Graphical+Model+Inference+Problem
11. Variational Inference: A Review for Statisticians — David M. Blei, Alp Kucukelbir, Jon D. McAuliffe, 2017
https://scholar.google.com/scholar?q=Variational+Inference%3A+A+Review+for+Statisticians
12. Auto-Encoding Variational Bayes — Diederik P. Kingma, Max Welling, 2013
https://scholar.google.com/scholar?q=Auto-Encoding+Variational+Bayes
13. An Introduction to Variational Methods for Graphical Models — Michael I. Jordan, Zoubin Ghahramani, Tommi S. Jaakkola, Lawrence K. Saul, 1999
https://scholar.google.com/scholar?q=An+Introduction+to+Variational+Methods+for+Graphical+Models
14. Predictive Sampling: Real-Time Behaviour Synthesis with MuJoCo — Taylor Howell, Nimrod Gileadi, Saran Tunyasuvunakool, Kevin Zakka, Tom Erez, Yuval Tassa, 2022
https://scholar.google.com/scholar?q=Predictive+Sampling%3A+Real-Time+Behaviour+Synthesis+with+MuJoCo
15. STORM: An Integrated Framework for Fast Joint-Space Model-Predictive Control for Reactive Manipulation — Mohak Bhardwaj, Balakumar Sundaralingam, Arsalan Mousavian, Nathan D. Ratliff, Dieter Fox, Fabio Ramos, Byron Boots, 2021
https://scholar.google.com/scholar?q=STORM%3A+An+Integrated+Framework+for+Fast+Joint-Space+Model-Predictive+Control+for+Reactive+Manipulation
16. Information-Theoretic Model Predictive Control: Theory and Applications to Autonomous Driving — Grady Williams, Paul Drews, Brian Goldfain, James M. Rehg, Evangelos A. Theodorou, 2018
https://scholar.google.com/scholar?q=Information-Theoretic+Model+Predictive+Control%3A+Theory+and+Applications+to+Autonomous+Driving
17. Model-Based Diffusion for Trajectory Optimization — Chaoyi Pan, Zeji Yi, Guanya Shi, Guannan Qu, 2024
https://scholar.google.com/scholar?q=Model-Based+Diffusion+for+Trajectory+Optimization
18. TD-MPC2: Scalable, Robust World Models for Continuous Control — Nicklas Hansen, Hao Su, Xiaolong Wang, 2023
https://scholar.google.com/scholar?q=TD-MPC2%3A+Scalable%2C+Robust+World+Models+for+Continuous+Control
19. Recent Advances in Path Integral Control for Trajectory Optimization: An Overview in Theoretical and Algorithmic Perspectives — Muhammad Kazim, Jungee Hong, Min-Gyeom Kim, Kwang-Ki K. Kim, 2024
https://scholar.google.com/scholar?q=Recent+Advances+in+Path+Integral+Control+for+Trajectory+Optimization%3A+An+Overview+in+Theoretical+and+Algorithmic+Perspectives
Interactive Visualization: Volatility Optimization Is Actually Bayesian Inference

This episode explores TwinQuant, a 4-bit post-training quantization method for large language models that challenges a core assumption behind prior techniques like SVDQuant: that a weight matrix's important information can be captured in a small, fixed set of directions. The hosts explain how LLM weight outliers turn out to be spread across hundreds of directions rather than concentrated, forcing earlier low-rank decomposition approaches into an unwinnable tradeoff between speed and accuracy. They unpack TwinQuant's solution — learning the low-rank split itself via manifold optimization, using a true orthogonal (Stiefel manifold) rotation that folds cleanly into RMSNorm layers alongside a more flexible invertible (general linear) transform for layer-specific residual handling — plus a fused kernel designed to keep the approach fast at inference. Along the way, the conversation walks through foundational quantization vocabulary (PTQ, WxAy notation, mixed-precision splits) for listeners newer to the topic. It's a compelling listen for anyone tracking how far LLMs can be compressed without sacrificing accuracy, and why the math behind "which parts of a weight matrix matter" is more complicated than earlier compression work assumed.

Sources:
1. TwinQuant: Learnable Subspace Decomposition for 4-Bit LLM Quantization — Haodong Wang, Junjie Liu, Zicong Hong, Qianli Liu, Jian Lin, Song Guo, Xu Chen, 2026
http://arxiv.org/abs/2606.01556
2. Optimization Algorithms on Matrix Manifolds — P.-A. Absil, R. Mahony, R. Sepulchre, 2008
https://scholar.google.com/scholar?q=Optimization+Algorithms+on+Matrix+Manifolds
3. SpinQuant: LLM Quantization with Learned Rotations — Zechun Liu, Changsheng Zhao, Igor Fedorov, et al. (Meta AI), 2024
https://scholar.google.com/scholar?q=SpinQuant%3A+LLM+Quantization+with+Learned+Rotations
4. QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs — Saleh Ashkboos, Amirkeivan Mohtashami, Maximilian L. Croci, et al., 2024
https://scholar.google.com/scholar?q=QuaRot%3A+Outlier-Free+4-Bit+Inference+in+Rotated+LLMs
5. Efficient Riemannian Optimization on the Stiefel Manifold via the Cayley Transform — Jun Li, Fuxin Li, Sinisa Todorovic, 2020
https://scholar.google.com/scholar?q=Efficient+Riemannian+Optimization+on+the+Stiefel+Manifold+via+the+Cayley+Transform
6. SVDQuant: Absorbing Outliers by Low-Rank Components for 4-Bit Diffusion Models — Li, M., Lin, Y., Zhang, Z., Cai, T., Li, X., Guo, J., Xie, E., Meng, C., Zhu, J.-Y., Han, S., 2025
https://scholar.google.com/scholar?q=SVDQuant%3A+Absorbing+Outliers+by+Low-Rank+Components+for+4-Bit+Diffusion+Models
7. FlatQuant: Flatness Matters for LLM Quantization — Sun, Y., Liu, R., Bai, H., Bao, H., Zhao, K., Li, Y., Yu, X., Hou, L., Yuan, C., Jiang, X., et al., 2025
https://scholar.google.com/scholar?q=FlatQuant%3A+Flatness+Matters+for+LLM+Quantization
8. OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models — Shao, W., Chen, M., Zhang, Z., Xu, P., Zhao, L., Li, Z., Zhang, K., Gao, P., Qiao, Y., Luo, P., 2024
https://scholar.google.com/scholar?q=OmniQuant%3A+Omnidirectionally+Calibrated+Quantization+for+Large+Language+Models
Interactive Visualization: TwinQuant: Manifold-Constrained Low-Rank Decomposition for 4-Bit Quantization

This episode explores TFGN, an architectural approach to continual pre-training of large language models that claims to solve catastrophic forgetting without four common crutches: replay buffers, task identifiers, small-scale toy benchmarks, and external penalty terms like Fisher-information regularization. The hosts trace the lineage of the forgetting problem back to 1989, explain why popular fixes like LoRA-based parameter-efficient fine-tuning don't actually address forgetting (they just shrink the blast radius), and why classic regularization methods like Elastic Weight Consolidation break down at billion-parameter scale. They also clarify why long-context windows and prompt-based knowledge aren't a substitute for genuinely updating model weights on massive, unbounded corpora like full codebases or legal archives. The conversation lays out TFGN's core mechanism as a dense, input-conditioned overlay operating inside each transformer block, contrasting it with sparse mixture-of-experts routing, and sets up backward transfer as the key metric for measuring whether old knowledge survives new training. Listeners interested in how production LLMs might eventually absorb new domains without expensive retraining or fragile adapter stacking will find the framing of this open problem sharply drawn.

Sources:
1. TFGN: Task-Free, Replay-Free Continual Pre-Training Without Catastrophic Forgetting at LLM Scale — Anurup Ganguli, 2026
http://arxiv.org/abs/2605.15053
2. Overcoming catastrophic forgetting in neural networks (EWC) — J. Kirkpatrick et al., 2017
https://scholar.google.com/scholar?q=Overcoming+catastrophic+forgetting+in+neural+networks+%28EWC%29
3. Loss of plasticity in deep continual learning — S. Dohare et al., 2024, Nature
https://scholar.google.com/scholar?q=Loss+of+plasticity+in+deep+continual+learning
4. Mechanistic Analysis of Catastrophic Forgetting in Large Language Models During Continual Fine-Tuning — O. Y. L. Imanov, 2026, arXiv:2601.18699
https://scholar.google.com/scholar?q=Mechanistic+Analysis+of+Catastrophic+Forgetting+in+Large+Language+Models+During+Continual+Fine-Tuning
5. Examining Forgetting in Continual Pre-training of Aligned Large Language Models — C.-A. Li and H.-Y. Lee, 2024, arXiv:2401.03129
https://scholar.google.com/scholar?q=Examining+Forgetting+in+Continual+Pre-training+of+Aligned+Large+Language+Models
6. Revisiting Replay and Gradient Alignment for Continual Pre-Training of Large Language Models — I. Abbes, G. Subbaraj, M. Riemer, et al., 2025, arXiv:2508.01908
https://scholar.google.com/scholar?q=Revisiting+Replay+and+Gradient+Alignment+for+Continual+Pre-Training+of+Large+Language+Models
7. Do Latent Tokens Think? A Causal and Adversarial Analysis of Chain-of-Continuous-Thought — Y. Zhang, B. Tang, T. Ju, S. Duan, G. Liu, 2025, arXiv:2512.21711
https://scholar.google.com/scholar?q=Do+Latent+Tokens+Think%3F+A+Causal+and+Adversarial+Analysis+of+Chain-of-Continuous-Thought
Interactive Visualization: TFGN: Replay-Free, Task-Free Continual Pre-Training at Scale

This episode explores SnapStream, a technique from SambaNova Systems for compressing KV caches during long-sequence LLM decoding on dataflow accelerators, demonstrated at production scale with a 671-billion-parameter DeepSeek-R1 deployment running 128K-token context at over 1,800 tokens per second. The discussion covers why established training-free KV cache eviction methods like SnapKV and StreamingLLM have struggled to reach real deployments despite promising accuracy results: continuous batching makes it unclear when to trigger compression across requests at different lifecycle stages, and static-graph compilers used by dataflow accelerators can't easily accommodate the dynamic, variable-shaped operations that standard compression implementations rely on. It explains how SnapStream fuses SnapKV's attention-based token selection with StreamingLLM's sink-plus-sliding-window approach into a single fixed-size cache, splitting sequences into sink tokens, recent tokens, and a compressed middle section during prefill. The conversation is grounded in fundamentals—clarifying the prefill/decode split, why decode is memory-bound, and what makes dataflow accelerators architecturally different from GPUs—making it accessible to listeners unfamiliar with KV cache mechanics while still delivering a specific, hardware-grounded engineering story rather than a purely algorithmic one.

Sources:
1. SnapStream: Efficient Long Sequence Decoding on Dataflow Accelerators — Jonathan Li, Nasim Farahini, Evgenii Iuliugin, Magnus Vesterlund, Christian Häggström, Guangtao Wang, Shubhangi Upasani, Ayush Sachdeva, Rui Li, Faline Fu, Chen Wu, Ayesha Siddiqua, John Long, Tuowen Zhao, Matheen Musaddiq, Håkan Zeffer, Yun Du, Mingran Wang, Qinghua Li, Bo Li, Urmish Thakker, Raghu Prabhakar, 2025
http://arxiv.org/abs/2511.03092
2. Plasticine: A Reconfigurable Architecture For Parallel Patterns — Raghu Prabhakar, Yaqi Zhang, David Koeplinger, Matt Feldman, Tian Zhao, Stefan Hadjis, Ardavan Pedram, Christos Kozyrakis, Kunle Olukotun, 2017
https://scholar.google.com/scholar?q=Plasticine%3A+A+Reconfigurable+Architecture+For+Parallel+Patterns
3. SN40L: Scaling the AI Memory Wall with Dataflow and Composition of Experts — Raghu Prabhakar and SambaNova Systems architecture team, 2024
https://scholar.google.com/scholar?q=SN40L%3A+Scaling+the+AI+Memory+Wall+with+Dataflow+and+Composition+of+Experts
4. Think Fast: A Tensor Streaming Processor (TSP) for Accelerating Deep Learning Workloads — Dennis Abts, Jonathan Ross, Jonathan Sparling, Mark Wong-VanHaren, et al. (Groq), 2020
https://scholar.google.com/scholar?q=Think+Fast%3A+A+Tensor+Streaming+Processor+%28TSP%29+for+Accelerating+Deep+Learning+Workloads
5. In-Datacenter Performance Analysis of a Tensor Processing Unit — Norman P. Jouppi et al. (Google), 2017
https://scholar.google.com/scholar?q=In-Datacenter+Performance+Analysis+of+a+Tensor+Processing+Unit
6. DistServe: Disaggregating Prefill and Decoding for Goodput-Optimized LLM Serving — Y. Zhong, S. Liu, J. Chen, J. Hu, Y. Zhu, X. Liu, X. Jin, H. Zhang, 2024
https://scholar.google.com/scholar?q=DistServe%3A+Disaggregating+Prefill+and+Decoding+for+Goodput-Optimized+LLM+Serving
7. Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference — J. Tang, Y. Zhao, K. Zhu, G. Xiao, B. Kasikci, S. Han, 2024
https://scholar.google.com/scholar?q=Quest%3A+Query-Aware+Sparsity+for+Efficient+Long-Context+LLM+Inference
8. InfLLM: Training-Free Long-Context Extrapolation with an Efficient Context Memory — C. Xiao, P. Zhang, X. Han, G. Xiao, Y. Lin, Z. Zhang, Z. Liu, M. Sun, 2024
https://scholar.google.com/scholar?q=InfLLM%3A+Training-Free+Long-Context+Extrapolation+with+an+Efficient+Context+Memory
9. DeepSeek-V3.2-Exp: Boosting Long-Context Efficiency with DeepSeek Sparse Attention — DeepSeek-AI, 2025
https://scholar.google.com/scholar?q=DeepSeek-V3.2-Exp%3A+Boosting+Long-Context+Efficiency+with+DeepSeek+Sparse+Attention
10. Cartridges: Lightweight and General-Purpose Long Context Representations via Self-Study — S. Eyuboglu, R. Ehrlich, S. Arora, N. Guha, D. Zinsley, E. Liu, W. Tennien, A. Rudra, J. Zou, A. Mirhoseini, C. Re, 2025
https://scholar.google.com/scholar?q=Cartridges%3A+Lightweight+and+General-Purpose+Long+Context+Representations+via+Self-Study
11. LLMs Know What to Drop: Self-Attention Guided KV Cache Eviction (SAGE-KV) — G. Wang, S. Upasani, C. Wu, D. Gandhi, J. Li, C. Hu, B. Li, U. Thakker, 2025
https://scholar.google.com/scholar?q=LLMs+Know+What+to+Drop%3A+Self-Attention+Guided+KV+Cache+Eviction+%28SAGE-KV%29
Interactive Visualization: SnapStream: Taming KV-Cache Eviction for Dataflow Accelerators

This episode explores StrataCL, a fabric-native communication library from researchers at Peking University, ICT-CAS, UCAS, Shanghai Jiao Tong University, and Huawei, tested on Huawei's CloudMatrix384 supernode. The discussion centers on how communication overhead — which the paper puts at 30-45% of end-to-end time in distributed LLM training and up to 50% at scale — can be cut by giving collectives and MoE routing true zero-copy access to application buffers on unified-memory fabrics, without breaking compatibility with frameworks like PyTorch and SGLang. A key insight is why buffer-centric libraries like NCCL and HCCL fall short even on fast unified-address fabrics, and how MoE dispatch/combine traffic exposes the limits of naive redesigns. The core technical contribution is registration-on-allocation: exploiting the multi-second gap between physical memory allocation and first use by a communication operator to move registration off the critical path entirely, asynchronously, the moment memory is mapped. The result is a 1.4x iteration-time speedup on a 512-die production training run with no changes to the model, optimizer, or data — pure systems engineering payoff.

Sources:
1. StrataCL: Fabric-Native Communication Library for Production Supernodes — Tiancheng Hu, Jin Qin, Yuzheng Wang, Ke Liu, TangShengsheng Li, Sheng Wang, Zhongzhe Hu, Tianlun Hu, Wei Wang, Lijun Li, Jingbin Zhou, Xiaoming Bao, Hongwei Sun, Jieru Zhao, Huimin Cui, Tao Xie, Chenxi Wang, 2026
http://arxiv.org/abs/2607.26444
2. U-Net: A User-Level Network Interface for Parallel and Distributed Computing — Thorsten von Eicken, Anindya Basu, Vineet Buch, Werner Vogels, 1995
https://scholar.google.com/scholar?q=U-Net%3A+A+User-Level+Network+Interface+for+Parallel+and+Distributed+Computing
3. Design Guidelines for High Performance RDMA Systems — Anuj Kalia, Michael Kaminsky, David G. Andersen, 2016
https://scholar.google.com/scholar?q=Design+Guidelines+for+High+Performance+RDMA+Systems
4. Blink: Fast and Generic Collectives for Distributed ML — Guanhua Wang, Shivaram Venkataraman, Amar Phanishayee, Jorgen Thelin, Nikhil Devanur, Ion Stoica, 2020
https://scholar.google.com/scholar?q=Blink%3A+Fast+and+Generic+Collectives+for+Distributed+ML
5. NVSHMEM (GPU-initiated, PGAS-style one-sided communication library) — NVIDIA (library/runtime, not a single academic paper), 2016 (initial release, iterated since)
https://scholar.google.com/scholar?q=NVSHMEM+%28GPU-initiated%2C+PGAS-style+one-sided+communication+library%29
6. Collective Communication for 100k+ GPUs — Min Si, Pavan Balaji, Yongzhou Chen, et al., 2025
https://scholar.google.com/scholar?q=Collective+Communication+for+100k%2B+GPUs
7. SwiftEP: Accelerating MoE Inference with Buffer Fusion and TMA Offloading — Xingyi Li, Yadong Liu, Xiaojie Huang, et al., 2026 (NSDI 26)
https://scholar.google.com/scholar?q=SwiftEP%3A+Accelerating+MoE+Inference+with+Buffer+Fusion+and+TMA+Offloading
8. PyTorch Symmetric Memory / NVSHMEM-style same-VA mirrored buffers — PyTorch Team / NVIDIA (NVSHMEM), 2024-2025
https://scholar.google.com/scholar?q=PyTorch+Symmetric+Memory+%2F+NVSHMEM-style+same-VA+mirrored+buffers
9. Efficient Memory Management for Large Language Model Serving with PagedAttention — Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, et al., 2023 (SOSP 23)
https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention
Interactive Visualization: StrataCL: Fabric-Native Zero-Copy Comms for AI Supernodes

This episode explores a challenge to conventional wisdom in parameter-efficient fine-tuning, examining a method called MiCA that inverts the logic behind LoRA (Low-Rank Adaptation). Rather than letting trainable weight-update matrices drift freely, as standard LoRA does, MiCA deliberately anchors one matrix to the minor singular-value directions of a weight matrix — the low-energy, rarely-used "corners" that classical compression theory says to discard — leaving those directions free for new knowledge rather than overwriting the dominant, pretrained-heavy subspace. The discussion traces the technique's lineage through SVD, the Eckart-Young-Mirsky theorem, PiSSA's SVD-based initialization, and Minor Component Analysis, framing MiCA's core bet: catastrophic forgetting during fine-tuning may stem from cramming new information into already-saturated high-energy directions. Listeners interested in the mechanics of efficient model adaptation, knowledge editing, and where the field's assumptions about "useless" weight-matrix structure might be wrong will find the debate over whether this is a genuine architectural insight or a narrower refinement of existing PEFT ideas especially engaging.

Sources:
1. MiCA: Mining Minor Singular Directions for Knowledge Injection Beyond LoRA
https://arxiv.org/pdf/2604.01694
2. The Approximation of One Matrix by Another of Lower Rank — Carl Eckart, Gale Young, 1936
https://scholar.google.com/scholar?q=The+Approximation+of+One+Matrix+by+Another+of+Lower+Rank
3. LoRA: Low-Rank Adaptation of Large Language Models — Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, 2021
https://scholar.google.com/scholar?q=LoRA%3A+Low-Rank+Adaptation+of+Large+Language+Models
4. PiSSA: Principal Singular Values and Singular Vectors Adaptation of Large Language Models — Fanxu Meng, Zhaohui Wang, Muhan Zhang, 2024
https://scholar.google.com/scholar?q=PiSSA%3A+Principal+Singular+Values+and+Singular+Vectors+Adaptation+of+Large+Language+Models
5. AdaLoRA: Adaptive Budget Allocation for Parameter-Efficient Fine-Tuning — Qingru Zhang, Minshuo Chen, Alexander Bukharin, Pengcheng He, Yu Cheng, Weizhu Chen, Tuo Zhao, 2023
https://scholar.google.com/scholar?q=AdaLoRA%3A+Adaptive+Budget+Allocation+for+Parameter-Efficient+Fine-Tuning
6. SOMA: Singular Value Decomposed Minor Components Adaptation for Domain Generalizable Representation Learning — Seokju Yun, Seunghye Chae, Dongheon Lee, Youngmin Ro, 2025
https://scholar.google.com/scholar?q=SOMA%3A+Singular+Value+Decomposed+Minor+Components+Adaptation+for+Domain+Generalizable+Representation+Learning
7. Learning Rate Matters: Vanilla LoRA May Suffice for LLM Fine-Tuning — Yu-Ang Lee, Ching-Yun Ko, Pin-Yu Chen, Mi-Yen Yeh, 2026
https://scholar.google.com/scholar?q=Learning+Rate+Matters%3A+Vanilla+LoRA+May+Suffice+for+LLM+Fine-Tuning
8. DoRA: Weight-Decomposed Low-Rank Adaptation — Shih-Yang Liu, Chien-Yi Wang, Hongxu Yin, Pavlo Molchanov, Yu-Chiang Frank Wang, Kwang-Ting Cheng, Min-Hung Chen, 2024
https://scholar.google.com/scholar?q=DoRA%3A+Weight-Decomposed+Low-Rank+Adaptation
9. Locating and Editing Factual Associations in GPT (ROME) — Kevin Meng, David Bau, Alex Andonian, Yonatan Belinkov, 2022
https://scholar.google.com/scholar?q=Locating+and+Editing+Factual+Associations+in+GPT+%28ROME%29
10. Editing Models with Task Arithmetic — Gabriel Ilharco, Marco Tulio Ribeiro, Mitchell Wortsman, Suchin Gururangan, Ludwig Schmidt, Hannaneh Hajishirzi, Ali Farhadi, 2023
https://scholar.google.com/scholar?q=Editing+Models+with+Task+Arithmetic
Interactive Visualization: MiCA: Mining Minor Singular Directions for Knowledge Injection Beyond LoRA

This episode explores cross-instance attention in disaggregated LLM serving, focusing on the surprising size inversion created by Multi-head Latent Attention: a routed decoding query shrinks to roughly a kilobyte while the cache chunk it must read can balloon to 61 megabytes across layers, upending the old assumption that query and cache are comparably sized. The discussion traces why this scenario is becoming routine — providers sharing precomputed caches for large corpora that outgrow a single GPU's memory, and agentic workloads where many sub-agents query one oversized shared prefix — and lays out the three possible strategies (route, fetch, or recompute locally) for handling the mismatch, including how sparse indexers further shrink the routable unit to scattered top-k blocks. A key thread examines device-initiated RDMA via IBGDA, challenging the intuition that skipping the CPU proxy is automatically faster: prior work on tiny mixture-of-experts messages actually found IBGDA slower, but the paper's controlled test on kilobyte-scale attention traffic shows the CPU-proxy path is 40% slower at the median and over 50% slower at steady state. Listeners interested in GPU networking, KV-cache architecture, or the practical plumbing behind large-scale LLM inference will find the paper's empirical resolution of a previously untested assumption particularly compelling.

Sources:
1. Move the Query, Not the Cache: Characterizing Cross-Instance Latent Attention Redistribution Across GPU Fabrics — Bole Ma, Jan Eitzinger, Harald Köstler, Gerhard Wellein, 2026
http://arxiv.org/abs/2606.01502
2. DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model — DeepSeek-AI (research team), 2024
https://scholar.google.com/scholar?q=DeepSeek-V2%3A+A+Strong%2C+Economical%2C+and+Efficient+Mixture-of-Experts+Language+Model
3. DeepSeek-V3 Technical Report — DeepSeek-AI (research team), 2024
https://scholar.google.com/scholar?q=DeepSeek-V3+Technical+Report
4. GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints — Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebrón, Sumit Sanghai, 2023
https://scholar.google.com/scholar?q=GQA%3A+Training+Generalized+Multi-Query+Transformer+Models+from+Multi-Head+Checkpoints
5. TransMLA: Multi-Head Latent Attention Is All You Need — Fanxu Meng, Zengwei Yao, Muhan Zhang, et al., 2025
https://scholar.google.com/scholar?q=TransMLA%3A+Multi-Head+Latent+Attention+Is+All+You+Need
6. Improving Network Performance of HPC Systems Using NVIDIA Magnum IO NVSHMEM and GPUDirect Async — NVIDIA (NVSHMEM / Magnum IO engineering team), 2023
https://scholar.google.com/scholar?q=Improving+Network+Performance+of+HPC+Systems+Using+NVIDIA+Magnum+IO+NVSHMEM+and+GPUDirect+Async
7. Efficient Inter-node MPI Communication using GPUDirect RDMA for InfiniBand Clusters with NVIDIA GPUs — Sreeram Potluri, Khaled Hamidouche, Akshay Venkatesh, Devendar Bureddy, Dhabaleswar K. Panda, 2013
https://scholar.google.com/scholar?q=Efficient+Inter-node+MPI+Communication+using+GPUDirect+RDMA+for+InfiniBand+Clusters+with+NVIDIA+GPUs
8. DeepEP: an efficient expert-parallel communication library (and related DeepSeek-V3 Technical Report communication sections) — DeepSeek-AI (research/infra team), 2025 / 2024
https://scholar.google.com/scholar?q=DeepEP%3A+an+efficient+expert-parallel+communication+library+%28and+related+DeepSeek-V3+Technical+Report+communication+sections%29
9. Introducing OpenSHMEM: SHMEM for the PGAS Community — Barbara Chapman, Tony Curtis, Swaroop Pophale, et al., 2010
https://scholar.google.com/scholar?q=Introducing+OpenSHMEM%3A+SHMEM+for+the+PGAS+Community
10. Preble: Efficient Distributed Prompt Scheduling for LLM Serving — Vikranth Srivatsa, Zijian He, Reyna Abhyankar, Dongming Li, Yiying Zhang, 2024
https://scholar.google.com/scholar?q=Preble%3A+Efficient+Distributed+Prompt+Scheduling+for+LLM+Serving
11. MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool — Cunchen Hu, Heyang Huang, Junhao Hu, Jiang Xu, Xusheng Chen, Tao Xie, Chenxi Wang, Sa Wang, Yungang Bao, Ninghui Sun, Yizhou Shan, 2024
https://scholar.google.com/scholar?q=MemServe%3A+Context+Caching+for+Disaggregated+LLM+Serving+with+Elastic+Memory+Pool
12. H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models — Zhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen, Lianmin Zheng, Ruisi Cai, Zhao Song, Yuandong Tian, Christopher Ré, Clark Barrett, Zhangyang Wang, Beidi Chen, 2023
https://scholar.google.com/scholar?q=H2O%3A+Heavy-Hitter+Oracle+for+Efficient+Generative+Inference+of+Large+Language+Models
Interactive Visualization: Move the Query, Not the Cache: MLA Rewrites GPU Fabric Attention Routing

This episode examines "Silicon Showdown," a study comparing Nvidia discrete-GPU and Apple unified-memory architectures for running large language models on consumer hardware, tested across model sizes from 1.5 billion to 80 billion parameters. It explains why Nvidia's VRAM Wall forces a stark trade-off between quantizing models down or offloading to slower system RAM across a PCIe bottleneck, while Apple's unified memory pool lets large models load fully without that penalty, at the cost of slower per-byte bandwidth. The discussion breaks down the competing software stacks—Nvidia's TensorRT-LLM with its new NVFP4 format and split-backend behavior, Apple's compilation-free MLX, and the cross-platform GGUF fallback from llama.cpp—and how each shapes real-world performance on metrics like time-to-first-token and tokens per joule. The episode highlights a gap in existing benchmarks like MLPerf and vLLM research, which focus on data-center throughput rather than the moment a model outgrows a single consumer GPU's memory. Listeners interested in running frontier open-weight models like Llama-3.3-70B or Qwen3-Next-80B on their own hardware will find a grounded, hardware-specific account of where each platform's approach breaks down.

Sources:
1. Silicon Showdown: Performance, Efficiency, and Ecosystem Barriers in Consumer-Grade LLM Inference — Abdurrahman Javat, Allan Kazakov, 2026
http://arxiv.org/abs/2605.00519
2. GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers — Frantar, Ashkboos, Hoefler, Alistarh, 2022
https://scholar.google.com/scholar?q=GPTQ%3A+Accurate+Post-Training+Quantization+for+Generative+Pre-trained+Transformers
3. AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration — Lin, Tang, Tang, Yang, Dang, Han, 2023
https://scholar.google.com/scholar?q=AWQ%3A+Activation-aware+Weight+Quantization+for+LLM+Compression+and+Acceleration
4. LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale — Dettmers, Lewis, Belkada, Zettlemoyer, 2022
https://scholar.google.com/scholar?q=LLM.int8%28%29%3A+8-bit+Matrix+Multiplication+for+Transformers+at+Scale
5. Mixtral of Experts — Jiang et al. (Mistral AI), 2024
https://scholar.google.com/scholar?q=Mixtral+of+Experts
6. Fast Inference from Transformers via Speculative Decoding — Leviathan, Kalman, Matias, 2023
https://scholar.google.com/scholar?q=Fast+Inference+from+Transformers+via+Speculative+Decoding
7. Efficient Memory Management for Large Language Model Serving with PagedAttention — Kwon et al., 2023
https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention
Interactive Visualization: Silicon Showdown: GPU vs Apple Silicon LLM Inference Limits

Sitting down with a two-author-plus-one chemical-engineering-rooted textbook on Model Predictive Control, this episode unpacks why the field treats real-time feasibility as a hard constraint rather than a nice-to-have — walking through how MPC re-solves an optimization problem from scratch every control cycle using a known dynamics model, with no learning or reward signal involved. The discussion centers on the structural trick that makes this tractable on embedded hardware: exploiting the block-banded, time-local coupling of the problem via Riccati recursion or condensing to cut a naive O(N³) solve down to O(N), and how Diehl, Bock, and Schlöder's 2005 real-time iteration scheme turned this from a lab curiosity into something a drone or engine controller can rerun dozens of times per second. It also covers moving horizon estimation as the optimization-based counterpart to the Kalman filter, explaining why MHE can enforce physical constraints a Kalman filter can't, and why the book cuts particle filtering from its main text once state dimensionality climbs past five. Listeners get a clear picture of why the same machinery underlies powered-descent guidance, automotive control, and legged robotics — not as a trend, but as the only approach that reliably meets millisecond-scale deadlines.

Sources:
1. Model Predictive Control's Real-Time Structure, from Chapter to Cockpit
https://sites.engineering.ucsb.edu/~jbraw/mpc/MPC-book-2nd-edition-4th-printing.pdf
2. A Real-Time Iteration Scheme for Nonlinear Optimization in Optimal Feedback Control — Moritz Diehl, Hans Georg Bock, Johannes P. Schlöder, 2005
https://scholar.google.com/scholar?q=A+Real-Time+Iteration+Scheme+for+Nonlinear+Optimization+in+Optimal+Feedback+Control
3. CasADi: A Software Framework for Nonlinear Optimization and Optimal Control — Joel A. E. Andersson, Joris Gillis, Greg Horn, James B. Rawlings, Moritz Diehl, 2019
https://scholar.google.com/scholar?q=CasADi%3A+A+Software+Framework+for+Nonlinear+Optimization+and+Optimal+Control
4. acados: A Modular Open-Source Framework for Fast Embedded Optimal Control — Robin Verschueren, Gianluca Frison, Dimitris Kouzoupis, Jonathan Frey, Niels van Duijkeren, Andrea Zanelli, Branimir Novoselnik, Thivaharan Albin, Rien Quirynen, Moritz Diehl, 2022
https://scholar.google.com/scholar?q=acados%3A+A+Modular+Open-Source+Framework+for+Fast+Embedded+Optimal+Control
5. On the Implementation of an Interior-Point Filter Line-Search Algorithm for Large-Scale Nonlinear Programming — Andreas Wächter, Lorenz T. Biegler, 2006
https://scholar.google.com/scholar?q=On+the+Implementation+of+an+Interior-Point+Filter+Line-Search+Algorithm+for+Large-Scale+Nonlinear+Programming
6. Constrained Linear State Estimation — A Moving Horizon Approach — Christopher V. Rao, James B. Rawlings, Jay H. Lee, 2001
https://scholar.google.com/scholar?q=Constrained+Linear+State+Estimation+%25E2%2580%2594+A+Moving+Horizon+Approach
7. Constrained State Estimation for Nonlinear Discrete-Time Systems: Stability and Moving Horizon Approximations — Christopher V. Rao, James B. Rawlings, David Q. Mayne, 2003
https://scholar.google.com/scholar?q=Constrained+State+Estimation+for+Nonlinear+Discrete-Time+Systems%3A+Stability+and+Moving+Horizon+Approximations
8. Moving-Horizon State Estimation for Nonlinear Discrete-Time Systems: New Stability Results and Approximation Schemes — Angelo Alessandri, Marco Baglietto, Giorgio Battistelli, 2008
https://scholar.google.com/scholar?q=Moving-Horizon+State+Estimation+for+Nonlinear+Discrete-Time+Systems%3A+New+Stability+Results+and+Approximation+Schemes
9. Stochastic Model Predictive Control: An Overview and Perspectives for Future Research — Ali Mesbah, 2016
https://scholar.google.com/scholar?q=Stochastic+Model+Predictive+Control%3A+An+Overview+and+Perspectives+for+Future+Research
10. Stochastic Linear Model Predictive Control with Chance Constraints — A Review — Marcello Farina, Luca Giulioni, Riccardo Scattolini, 2016
https://scholar.google.com/scholar?q=Stochastic+Linear+Model+Predictive+Control+with+Chance+Constraints+%25E2%2580%2594+A+Review
11. Learning-Based Model Predictive Control: Toward Safe Learning in Control — Lukas Hewing, Kim P. Wabersich, Marcel Menner, Melanie N. Zeilinger, 2020
https://scholar.google.com/scholar?q=Learning-Based+Model+Predictive+Control%3A+Toward+Safe+Learning+in+Control
12. Architectures for Distributed and Hierarchical Model Predictive Control — A Review — Riccardo Scattolini, 2009
https://scholar.google.com/scholar?q=Architectures+for+Distributed+and+Hierarchical+Model+Predictive+Control+%25E2%2580%2594+A+Review
13. Distributed MPC Strategies with Application to Power System Automatic Generation Control — Aswin N. Venkat, Ian A. Hiskens, James B. Rawlings, Stephen J. Wright, 2008
https://scholar.google.com/scholar?q=Distributed+MPC+Strategies+with+Application+to+Power+System+Automatic+Generation+Control
14. Distributed Model Predictive Control: A Tutorial Review and Future Research Directions — Panagiotis D. Christofides, Riccardo Scattolini, David Muñoz de la Peña, Jinfeng Liu, 2013
https://scholar.google.com/scholar?q=Distributed+Model+Predictive+Control%3A+A+Tutorial+Review+and+Future+Research+Directions
15. acados — a modular open-source framework for fast embedded optimal control — R. Verschueren, G. Frison, D. Kouzoupis, N. van Duijkeren, A. Zanelli, B. Novoselnik, T. Albin, R. Quirynen, M. Diehl, 2022
https://scholar.google.com/scholar?q=acados+%25E2%2580%2594+a+modular+open-source+framework+for+fast+embedded+optimal+control
16. The scenario approach to robust control design — G.C. Calafiore, M.C. Campi, 2006
https://scholar.google.com/scholar?q=The+scenario+approach+to+robust+control+design
17. Stability of nonstationary receding horizon control (foundational stability result for stochastic MPC) — D. Chatterjee, J. Lygeros, 2015
https://scholar.google.com/scholar?q=Stability+of+nonstationary+receding+horizon+control+%28foundational+stability+result+for+stochastic+MPC%29
18. Robust MPC and dissipativity-based analysis for stochastic constrained systems (source of Assumption 3.22, stochastic MPC Version 2) — D.Q. Mayne, P. Falugi, 2019
https://scholar.google.com/scholar?q=Robust+MPC+and+dissipativity-based+analysis+for+stochastic+constrained+systems+%28source+of+Assumption+3.22%2C+stochastic+MPC+Version+2%29
19. A model predictive control framework for industrial turbodiesel engine control (source system for the nonlinear distributed MPC example) — B.T. Stewart, A.N. Venkat, J.B. Rawlings, S.J. Wright, G. Pannocchia (2011 IEEE CDC paper referenced as Stewart et al. 2011), 2011
https://scholar.google.com/scholar?q=A+model+predictive+control+framework+for+industrial+turbodiesel+engine+control+%28source+system+for+the+nonlinear+distributed+MPC+example%29
Interactive Visualization: Model Predictive Control's Real-Time Structure, from Chapter to Cockpit

This episode explores adaptive verification for speculative decoding when the target model is a sparse Mixture-of-Experts (MoE) system rather than a dense transformer, focusing on the paper "Making Every Verified Token Count." The discussion traces the lineage from Leviathan et al.'s original speculative decoding through tree-based drafting methods like Medusa and EAGLE-3, then explains why MoE architectures break a core assumption: since different draft-tree branches can route to entirely different experts, verifying a tree means loading every expert any branch touched. Drawing on the paper's benchmarks across three MoE models (including Qwen3-30B-A3B), the hosts unpack the striking finding that verification alone consumes 79-89% of per-iteration decoding latency once trees grow past thirty nodes — flipping the "verification is nearly free" pitch that made speculative decoding attractive in the first place. Listeners interested in LLM inference serving, GPU memory-bandwidth bottlenecks, or the practical tradeoffs of deploying sparse MoE models will find the episode's breakdown of why dense-model intuition fails on MoE targets especially clarifying.

Sources:
1. Making Every Verified Token Count: Adaptive Verification for MoE Speculative Decoding — Lehan Pan, Ziyang Tao, Ruoyu Pang, Xiao Wang, Jianjun Zhao, Yanyong Zhang, 2026
http://arxiv.org/abs/2605.00342
2. Fast Inference from Transformers via Speculative Decoding — Yaniv Leviathan, Matan Kalman, Yossi Matias, 2023
https://scholar.google.com/scholar?q=Fast+Inference+from+Transformers+via+Speculative+Decoding
3. Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads — Tianle Cai, Yuhong Li, Zhengyang Geng, Hongwu Peng, Jason D. Lee, Deming Chen, Tri Dao, 2024
https://scholar.google.com/scholar?q=Medusa%3A+Simple+LLM+Inference+Acceleration+Framework+with+Multiple+Decoding+Heads
4. EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test — Yuhui Li, Fangyun Wei, Chao Zhang, Hongyang Zhang, 2025
https://scholar.google.com/scholar?q=EAGLE-3%3A+Scaling+up+Inference+Acceleration+of+Large+Language+Models+via+Training-Time+Test
5. Mixtral of Experts — Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, and the Mistral AI team, 2024
https://scholar.google.com/scholar?q=Mixtral+of+Experts
6. Utility-driven speculative decoding for mixture-of-experts — Anish Saxena, Po-An Tsai, Hritvik Taneja, Aamer Jaleel, Moinuddin Qureshi, 2025
https://scholar.google.com/scholar?q=Utility-driven+speculative+decoding+for+mixture-of-experts
7. MoE-Spec: Expert budgeting for efficient speculative decoding — Bradley McDanel, Steven Li, Sruthikesh Surineni, Harshit Khaitan, 2026
https://scholar.google.com/scholar?q=MoE-Spec%3A+Expert+budgeting+for+efficient+speculative+decoding
8. ECHO: Elastic speculative decoding with sparse gating for high-concurrency scenarios — Xinyi Hu, Yuhao Shen, Baolin Zhang, Hengxin Zhang, Jun Dai, Shuang Ge, Lei Chen, Yue Li, Mingcheng Wan, 2026
https://scholar.google.com/scholar?q=ECHO%3A+Elastic+speculative+decoding+with+sparse+gating+for+high-concurrency+scenarios
9. MoESD: Unveil speculative decoding's potential for accelerating sparse MoE — Zongle Huang, Lei Zhu, Zongyuan Zhan, Ting Hu, Weikai Mao, Xianzhi Yu, Yongpan Liu, Tianyu Zhang, 2025
https://scholar.google.com/scholar?q=MoESD%3A+Unveil+speculative+decoding%27s+potential+for+accelerating+sparse+MoE
Interactive Visualization: Making Every Verified Token Count in MoE Speculative Decoding

This episode explores FreeAct, a new approach to quantizing large language models down to 4-bit weights and activations (W4A4), presented by researchers from the National University of Singapore, Huawei Technology, and Central South University. The discussion traces how prior methods like QuaRot and FlatQuant rely on a rigid one-to-one pairing between a rotation matrix applied to activations and its exact inverse applied to weights — an assumption that breaks down for diffusion language models, where masked and unmasked tokens have different statistical profiles, and for multimodal models mixing vision and text tokens through the same layers. The hosts unpack the outlier-channel problem that makes activation quantization so much harder than weight quantization, tracing it back to Dettmers' LLM.int8 findings, and explain how FreeAct exploits a linear-algebra insight — dubbed Proposition 1 — showing that rank-deficient activation matrices allow a whole family of transformations rather than a single exact inverse, enabling different token types to use different activation-side matrices while keeping one shared weight-side transform. It's a compelling listen for anyone tracking how quantization techniques are adapting to increasingly heterogeneous token streams in modern AI systems.

Sources:
1. FreeAct: Freeing Activations for LLM Quantization — Xiaohao Liu, Xiaobo Xia, Manyi Zhang, Ji-Fu Li, Xianzhi Yu, Fei Shen, Xiu Su, See-Kiong Ng, Tat-Seng Chua, 2026
http://arxiv.org/abs/2603.01776
2. LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale — Tim Dettmers, Mike Lewis, Younes Belkada, Luke Zettlemoyer, 2022
https://scholar.google.com/scholar?q=LLM.int8%28%29%3A+8-bit+Matrix+Multiplication+for+Transformers+at+Scale
3. SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models — Guangxuan Xiao, Ji Lin, Mickael Seznec, Julien Demouth, Song Han, 2023
https://scholar.google.com/scholar?q=SmoothQuant%3A+Accurate+and+Efficient+Post-Training+Quantization+for+Large+Language+Models
4. Atom: Low-bit Quantization for Efficient and Accurate LLM Serving — Yilong Zhao, Chien-Yu Lin, Kan Zhu, et al., 2024
https://scholar.google.com/scholar?q=Atom%3A+Low-bit+Quantization+for+Efficient+and+Accurate+LLM+Serving
5. QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs — Saleh Ashkboos, Amirkeivan Mohtashami, Maximilian Croci, et al., 2024
https://scholar.google.com/scholar?q=QuaRot%3A+Outlier-Free+4-Bit+Inference+in+Rotated+LLMs
6. FlatQuant: Flatness Matters for LLM Quantization — Yuxuan Sun, et al., 2025
https://scholar.google.com/scholar?q=FlatQuant%3A+Flatness+Matters+for+LLM+Quantization
7. SpinQuant: LLM Quantization with Learned Rotations — Zechun Liu, Changsheng Zhao, et al. (Meta AI), 2024
https://scholar.google.com/scholar?q=SpinQuant%3A+LLM+Quantization+with+Learned+Rotations
8. QuIP: 2-Bit Quantization of Large Language Models With Guarantees — Jerry Chee, Yaohui Cai, Volodymyr Kuleshov, Christopher De Sa, 2023
https://scholar.google.com/scholar?q=QuIP%3A+2-Bit+Quantization+of+Large+Language+Models+With+Guarantees
9. MQuant: Unleashing the Inference Potential of Multimodal Large Language Models via Static Quantization — Jiangyong Yu, Sifan Zhou, Dawei Yang, et al., 2025
https://scholar.google.com/scholar?q=MQuant%3A+Unleashing+the+Inference+Potential+of+Multimodal+Large+Language+Models+via+Static+Quantization
10. DLLMQuant: Quantizing Diffusion-based Large Language Models — Chen Xu, Dan Yang, 2025
https://scholar.google.com/scholar?q=DLLMQuant%3A+Quantizing+Diffusion-based+Large+Language+Models
11. Quant-dLLM: Post-Training Extreme Low-Bit Quantization for Diffusion Large Language Models — Tianao Zhang, Zhiteng Li, Xianglong Yan, Haotong Qin, Yong Guo, Yulun Zhang, 2025
https://scholar.google.com/scholar?q=Quant-dLLM%3A+Post-Training+Extreme+Low-Bit+Quantization+for+Diffusion+Large+Language+Models
12. DuQuant: Distributing Outliers via Dual Transformation Makes Stronger Quantized LLMs — Haokun Lin, Haobo Xu, Yichen Wu, et al., 2024
https://scholar.google.com/scholar?q=DuQuant%3A+Distributing+Outliers+via+Dual+Transformation+Makes+Stronger+Quantized+LLMs
Interactive Visualization: FreeAct: Rethinking One-to-One Transforms for LLM Quantization

This episode surveys how large language model serving systems manage the key-value cache — the memory storing every token's key and value vectors — as it has grown from a disposable per-request tensor into a resource actively managed, moved, and contended for across GPUs, nodes, and storage tiers. Drawing on a Texas Tech University paper classifying over thirty existing systems, the hosts unpack the arithmetic behind why KV cache footprint balloons with long context windows (reaching roughly 40 gigabytes for a single 128K-token request on a 70-billion-parameter model) and why bandwidth, not just capacity, becomes the real bottleneck during decode. They trace the field's foundational shift back to PagedAttention, the vLLM technique that introduced OS-style paging for KV memory, and explain how nearly every later system builds on its block-table abstraction. The conversation then turns to a four-dimensional taxonomy — locality, lifetime, ownership, and transport — used to organize the design space, highlighting a striking gap where two of five lifetime categories contain zero real-world systems. Listeners interested in LLM infrastructure, memory hierarchies, or the practical limits of long-context and agentic serving will find a clear framework for reasoning about a problem that's easy to underestimate with a single "the cache grows" intuition.

Sources:
1. From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving — Jie Li, Tongyang Wang, Yong Chen, 2026
http://arxiv.org/abs/2607.02574
2. DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model — DeepSeek-AI, 2024
https://scholar.google.com/scholar?q=DeepSeek-V2%3A+A+Strong%2C+Economical%2C+and+Efficient+Mixture-of-Experts+Language+Model
3. Efficient Memory Management for Large Language Model Serving with PagedAttention — Woosuk Kwon et al., 2023
https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention
4. Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving — Ruoyu Qin et al., 2024
https://scholar.google.com/scholar?q=Mooncake%3A+A+KVCache-centric+Disaggregated+Architecture+for+LLM+Serving
5. Medusa / EAGLE speculative decoding work (Cai et al. 2024; Li et al. 2024) — Tianle Cai et al.; Yuhui Li et al., 2024
https://scholar.google.com/scholar?q=Medusa+%2F+EAGLE+speculative+decoding+work+%28Cai+et+al.+2024%3B+Li+et+al.+2024%29
6. S-LoRA: Serving Thousands of Concurrent LoRA Adapters — Ying Sheng et al., 2023/2024
https://scholar.google.com/scholar?q=S-LoRA%3A+Serving+Thousands+of+Concurrent+LoRA+Adapters
Interactive Visualization: Global Memory Bloat in Long-Context LLM Serving

This episode covers "DualDecoder: Accelerate Long Context LLM Inference by Predictive Prefetch," which tackles a hidden cost in dynamic sparse KV-cache systems: the GPU-resident bookkeeping state (landmarks, reconstructed keys) used to make host-memory offloading fast can itself consume up to 64% of GPU memory — 8.5 times larger than the actual sparse KV entries it's meant to retrieve. Drawing on the lineage from H2O's heavy-hitter observation to ShadowKV's landmark-based retrieval, the discussion explains how this auxiliary overhead quietly erodes the memory savings these systems promise, with ShadowKV reaching only 6.7% of its idealized batch-size capacity on a 32-billion-parameter model. DualDecoder's proposed fix is predictive prefetching: rather than permanently parking retrieval-support state on the GPU, it predicts the next decoding step's needs one step ahead and pulls entries from host memory just in time. Listeners interested in LLM inference efficiency will find a concrete, measured account of how a fix for one memory wall can quietly build a smaller one right next to it — and a proposed way out.

Sources:
1. DualDecoder: Accelerate Long Context LLM Inference by Predictive Prefetch — Zuning Liang, Zhiyi Yao, Qi Chen, Yuedong Xu, Hao Dai, Zhiqiang Ding, Tongkai Yang, Jinlong Hou, Yuan Cheng, 2026
http://arxiv.org/abs/2607.26475
2. SpeContext: Enabling Efficient Long-context Reasoning with Speculative Context Sparsity in LLMs — Jiaming Xu, Jiayi Pan, Hanzhen Wang, Yongkang Zhou, Jiancai Ye, Yu Wang, Guohao Dai, 2026
https://scholar.google.com/scholar?q=SpeContext%3A+Enabling+Efficient+Long-context+Reasoning+with+Speculative+Context+Sparsity+in+LLMs
3. H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models — Zhenyu Zhang, Ying Sheng, Tianyi Zhou, et al., 2023
https://scholar.google.com/scholar?q=H2O%3A+Heavy-Hitter+Oracle+for+Efficient+Generative+Inference+of+Large+Language+Models
4. RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval — Di Liu, Meng Chen, Baotong Lu, et al., 2024
https://scholar.google.com/scholar?q=RetrievalAttention%3A+Accelerating+Long-Context+LLM+Inference+via+Vector+Retrieval
5. MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool — Cunchen Hu, Heyang Huang, Junhao Hu, et al., 2024
https://scholar.google.com/scholar?q=MemServe%3A+Context+Caching+for+Disaggregated+LLM+Serving+with+Elastic+Memory+Pool
6. Efficient Memory Management for Large Language Model Serving with PagedAttention (vLLM) — Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, et al., 2023
https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention+%28vLLM%29
Interactive Visualization: DualDecoder: Fixing GPU Memory Bloat in Long-Context KV Cache Offloading

This episode explores why open-weight LLMs like Llama 3.1, Gemma3, Qwen3, and Olmo3 systematically lose 11–39% relative accuracy on facts from 2023–2024 compared to facts from 2020–2021, even though the more recent data falls within their training window. Drawing on Kyutai's paper "Understanding Data Temporality Impact on Large Language Models Pre-training," the discussion traces this "knowledge horizon gap" to a design choice baked into standard pretraining: corpora from many years are pooled and globally shuffled before training, erasing any timestamp signal and letting older, more frequently re-crawled data dominate. The hosts connect this to learning-rate decay schedules, arguing that data seen late in training — when updates are small and durable — gets imprinted far more strongly than data seen early, so chronological ordering (feeding snapshots 2018 through 2025 in sequence) could exploit that same mechanism to anchor recent facts instead of losing them. They situate the work against Zhao et al.'s "Set the Clock" research and Bengio's foundational curriculum-learning ideas, framing chronological training as a strikingly cheap intervention — same tokens, same compute, same architecture — for a problem the field has largely ignored. It's a compelling listen for anyone puzzling over why "knowledge cutoff" claims don't match what models actually seem to know.

Sources:
1. Understanding Data Temporality Impact on Large Language Models Pre-training — Hippolyte Pilchen, Romain Fabre, Franck Signe Talla, Patrick Perez, Edouard Grave, 2026
http://arxiv.org/abs/2605.22769
2. Set the Clock: Temporal Alignment of Pretrained Language Models — Bowen Zhao, Zander Brumbaugh, Yizhong Wang, Hannaneh Hajishirzi, Noah A. Smith, 2024
https://scholar.google.com/scholar?q=Set+the+Clock%3A+Temporal+Alignment+of+Pretrained+Language+Models
3. Time-Aware Language Models as Temporal Knowledge Bases — Bhuwan Dhingra, Jeremy R. Cole, Julian Martin Eisenschlos, Daniel Gillick, Jacob Eisenstein, William W. Cohen, 2022
https://scholar.google.com/scholar?q=Time-Aware+Language+Models+as+Temporal+Knowledge+Bases
4. TemporalWiki: A Lifelong Benchmark for Training and Evaluating Ever-Evolving Language Models — Joel Jang, Seonghyeon Ye, Changho Lee, Sohee Yang, Joongbo Shin, Janghoon Han, Gyeonghun Kim, Minjoon Seo, 2022
https://scholar.google.com/scholar?q=TemporalWiki%3A+A+Lifelong+Benchmark+for+Training+and+Evaluating+Ever-Evolving+Language+Models
5. RealTime QA: What's the Answer Right Now? — Jungo Kasai, Keisuke Sakaguchi, Yoichi Takahashi, Yutaro Yamada, Deqing Fu, Tushar Khot, Ashish Sabharwal, Rik Koncel-Kedziorski, Yejin Choi, Noah A. Smith, Kentaro Inui, 2022
https://scholar.google.com/scholar?q=RealTime+QA%3A+What%27s+the+Answer+Right+Now%3F
6. Curriculum Learning — Yoshua Bengio, Jérôme Louradour, Ronan Collobert, Jason Weston, 2009
https://scholar.google.com/scholar?q=Curriculum+Learning
7. TimeLMs: Diachronic Language Models from Twitter — Daniel Loureiro, Francesco Barbieri, Leonardo Neves, Luis Espinosa Anke, Jose Camacho-Collados, 2022
https://scholar.google.com/scholar?q=TimeLMs%3A+Diachronic+Language+Models+from+Twitter
8. In-Context Pretraining: Language Modeling Beyond Document Boundaries — Weijia Shi, Sewon Min, Maria Lomeli, Chunting Zhou, Margaret Li, Xi Victoria Lin, Noah A. Smith, Luke Zettlemoyer, Wen-tau Yih, Mike Lewis, 2023
https://scholar.google.com/scholar?q=In-Context+Pretraining%3A+Language+Modeling+Beyond+Document+Boundaries
9. Towards Continual Knowledge Learning of Language Models — Joel Jang, Seonghyeon Ye, Sohee Yang, Joongbo Shin, Janghoon Han, Gyeonghun Kim, Stanley Jungkyu Choi, Minjoon Seo, 2022
https://scholar.google.com/scholar?q=Towards+Continual+Knowledge+Learning+of+Language+Models
10. TiC-LM: A web-scale benchmark for time-continual LLM pretraining — Li, J., Armandpour, M., Mirzadeh, I., Mehta, S., Shankar, V., Vemulapalli, R., Bengio, S., Tuzel, O., Farajtabar, M., Pouransari, H., Faghri, F., 2025
https://scholar.google.com/scholar?q=TiC-LM%3A+A+web-scale+benchmark+for+time-continual+LLM+pretraining
11. How do language models learn facts? Dynamics, curricula and hallucinations — Zucchet, N., Bornschein, J., Chan, S. C., Lampinen, A. K., Pascanu, R., De, S., 2025
https://scholar.google.com/scholar?q=How+do+language+models+learn+facts%3F+Dynamics%2C+curricula+and+hallucinations
12. Data mixing can induce phase transitions in knowledge acquisition — Gu, X., Lyu, K., Li, J., Zhang, J., 2026
https://scholar.google.com/scholar?q=Data+mixing+can+induce+phase+transitions+in+knowledge+acquisition
13. TiMoE: Time-aware mixture of language experts — Faro, R., Fan, D., Alphaidze, T., Jaggi, M., 2025
https://scholar.google.com/scholar?q=TiMoE%3A+Time-aware+mixture+of+language+experts
14. Does your data spark joy? Performance gains from domain upsampling at the end of training — Blakeney, C., Paul, M., Larsen, B. W., Owen, S., Frankle, J., 2024
https://scholar.google.com/scholar?q=Does+your+data+spark+joy%3F+Performance+gains+from+domain+upsampling+at+the+end+of+training
Interactive Visualization: Data Temporality's Hidden Impact on LLM Pretraining

This episode explores "Lifelong Learning of Large Language Model based Agents: A Roadmap," a survey examining how AI agents can continuously adapt to changing environments without losing prior knowledge. The discussion centers on the stability-plasticity dilemma—the tension between preserving learned capabilities and remaining flexible enough to absorb new information—and how this classical problem from connectionist neuroscience resurfaces in a new form for modern agents that rarely fine-tune their underlying weights. Key arguments include the concept of "functional forgetting," where information technically persists in vector stores but becomes practically inaccessible if retrieval or context limits fail to surface it, and a four-part memory taxonomy spanning working, episodic, semantic, and parametric memory. The hosts also trace how this survey synthesizes and extends two separate research lineages—internal-knowledge-focused LLM surveys and agent-architecture surveys—into a unified framework modeled as a goal-conditioned POMDP. Listeners interested in why coding assistants, web-browsing agents, and other AI tools degrade over time as their environments shift will find concrete framing for that problem here.

Sources:
1. Lifelong Learning of Large Language Model based Agents: A Roadmap — Junhao Zheng, Chengming Shi, Xidi Cai, Qiuke Li, Duzhen Zhang, Chenxing Li, Dong Yu, Qianli Ma, 2025
http://arxiv.org/abs/2501.07278
2. Overcoming Catastrophic Forgetting in Neural Networks — James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, et al. (DeepMind), 2017
https://scholar.google.com/scholar?q=Overcoming+Catastrophic+Forgetting+in+Neural+Networks
3. Continual Lifelong Learning with Neural Networks: A Review — German I. Parisi, Ronald Kemker, Jose L. Part, Christopher Kanan, Stefan Wermter, 2019
https://scholar.google.com/scholar?q=Continual+Lifelong+Learning+with+Neural+Networks%3A+A+Review
4. Generative Agents: Interactive Simulacra of Human Behavior — Joon Sung Park, Joseph O'Brien, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, Michael S. Bernstein (Stanford / Google), 2023
https://scholar.google.com/scholar?q=Generative+Agents%3A+Interactive+Simulacra+of+Human+Behavior
5. Voyager: An Open-Ended Embodied Agent with Large Language Models — Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, Anima Anandkumar (NVIDIA, Caltech, UT Austin), 2023
https://scholar.google.com/scholar?q=Voyager%3A+An+Open-Ended+Embodied+Agent+with+Large+Language+Models
6. Towards Lifelong Learning of Large Language Models: A Survey — J. Zheng, S. Qiu, C. Shi, Q. Ma, 2024
https://scholar.google.com/scholar?q=Towards+Lifelong+Learning+of+Large+Language+Models%3A+A+Survey
7. Loss of Plasticity in Deep Continual Learning — S. Dohare, J. F. Hernandez-Garcia, Q. Lan, P. Rahman, A. R. Mahmood, R. S. Sutton, 2024
https://scholar.google.com/scholar?q=Loss+of+Plasticity+in+Deep+Continual+Learning
8. A Survey on Large Language Model Based Autonomous Agents — L. Wang, C. Ma, X. Feng, et al., 2024
https://scholar.google.com/scholar?q=A+Survey+on+Large+Language+Model+Based+Autonomous+Agents
9. WISE: Rethinking the Knowledge Memory for Lifelong Model Editing of Large Language Models — P. Wang, Z. Li, N. Zhang, Z. Xu, Y. Yao, Y. Jiang, P. Xie, F. Huang, H. Chen, 2024
https://scholar.google.com/scholar?q=WISE%3A+Rethinking+the+Knowledge+Memory+for+Lifelong+Model+Editing+of+Large+Language+Models
10. Catastrophic Interference in Connectionist Networks: The Sequential Learning Problem — M. McCloskey, N. J. Cohen, 1989
https://scholar.google.com/scholar?q=Catastrophic+Interference+in+Connectionist+Networks%3A+The+Sequential+Learning+Problem
Interactive Visualization: Adapting Without Forgetting: A Lifelong Learning Roadmap for LLM Agents

This episode explores DWDP (Distributed Weight Data Parallelism), a new NVIDIA-authored approach to LLM inference on NVL72 systems that targets a subtle but costly inefficiency: GPUs sitting idle while they wait to synchronize with slower peers. The hosts unpack how existing model-parallelism strategies—expert, tensor, and pipeline parallelism—all share a hidden flaw, forcing every GPU to hit a synchronization barrier at each layer boundary, which the paper's own baseline shows can waste around twelve percent of total inference time even under ordinary workload imbalance. They explain why smarter scheduling alone (cache-aware or load-aware routing) can't fix this, since it only shrinks the imbalance feeding into the wait rather than eliminating the wait itself. The discussion then turns to DWDP's core idea: keeping GPUs fully data-parallel while having each one asynchronously prefetch missing expert weights from peers on demand, timed to hide the fetch behind ongoing compute. Listeners interested in the mechanics of large-scale MoE inference, GPU synchronization bottlenecks, and practical systems-level solutions to straggler problems will find the technical walkthrough especially rewarding.

Sources:
1. DWDP: Distributed Weight Data Parallelism for High-Performance LLM Inference on NVL72 — Wanqian Li, Jintao Peng, Zongfei Jing, Tianyu Zhang, Ze Long, Xianjie Qiao, Xiaoming Chen, Dongxu Yang, Kefeng Duan, June Yang, 2026
http://arxiv.org/abs/2604.01621
2. DeepSeek-V3 Technical Report — DeepSeek-AI, Aixin Liu, Bei Feng, et al., 2024
https://scholar.google.com/scholar?q=DeepSeek-V3+Technical+Report
3. Mooncake: Trading More Storage for Less Computation — A KVCache-centric Architecture for Serving LLM Chatbot — Ruoyu Qin, Zheming Li, Weiran He, Mingxing Zhang, et al., 2025
https://scholar.google.com/scholar?q=Mooncake%3A+Trading+More+Storage+for+Less+Computation+%E2%80%94+A+KVCache-centric+Architecture+for+Serving+LLM+Chatbot
4. DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving — Yinmin Zhong, Shengyu Liu, Junda Chen, et al., 2024
https://scholar.google.com/scholar?q=DistServe%3A+Disaggregating+Prefill+and+Decoding+for+Goodput-optimized+Large+Language+Model+Serving
5. Splitwise: Efficient Generative LLM Inference Using Phase Splitting — Pratyush Patel, Esha Choukse, Chaojie Zhang, et al., 2024
https://scholar.google.com/scholar?q=Splitwise%3A+Efficient+Generative+LLM+Inference+Using+Phase+Splitting
6. Tutel: Adaptive Mixture-of-Experts at Scale — Changho Hwang, Wei Cui, Yifan Xiong, et al., 2023
https://scholar.google.com/scholar?q=Tutel%3A+Adaptive+Mixture-of-Experts+at+Scale
7. DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale — Samyam Rajbhandari, Conglong Li, Zhewei Yao, et al., 2022
https://scholar.google.com/scholar?q=DeepSpeed-MoE%3A+Advancing+Mixture-of-Experts+Inference+and+Training+to+Power+Next-Generation+AI+Scale
Interactive Visualization: Distributed Weight Data Parallelism Cuts LLM Inference Stalls

This episode explores cross-family speculative prefill, a technique for cutting long-context inference latency by using a small "draft" model to identify which parts of a lengthy prompt matter before a much larger target model processes it. The hosts unpack why this is a hard problem in principle — draft and target models often use completely different tokenizers and architectures, meaning attention-based importance signals shouldn't obviously transfer between them — and trace the lineage from speculative decoding through the original same-family Speculative Prefill work to this paper's cross-family generalization. They highlight the practical motivation: models like DeepSeek and Kimi-K2 have no smaller sibling in their own family, so a technique that only works with matched draft/target pairs is a dead end for real deployments. Key results discussed include an 18x reduction in time-to-first-token, and the episode weighs supporting evidence from prior work on attention sinks against the stronger, less obvious claim that a full salience ranking over a 100,000-token document can transfer across unrelated architectures. Listeners interested in practical LLM efficiency techniques and the mechanics of long-context inference will find the back-and-forth skepticism over whether the method should even work, given the tokenizer mismatch, particularly engaging.

Sources:
1. Cross-Family Speculative Prefill: Training-Free Long-Context Compression with Small Draft Models — Shubhangi Upasani, Ravi Shanker Raju, Bo Li, Mengmeng Ji, John Long, Chen Wu, Urmish Thakker, Guangtao Wang, 2026
http://arxiv.org/abs/2603.02631
2. Speculative Prefill — Liu et al., 2025
https://scholar.google.com/scholar?q=Speculative+Prefill
3. Fast Inference from Transformers via Speculative Decoding — Yaniv Leviathan, Matan Kalman, Yossi Matias, 2023
https://scholar.google.com/scholar?q=Fast+Inference+from+Transformers+via+Speculative+Decoding
4. Efficient Streaming Language Models with Attention Sinks — Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, Mike Lewis, 2023 (ICLR 2024)
https://scholar.google.com/scholar?q=Efficient+Streaming+Language+Models+with+Attention+Sinks
5. SnapKV: LLM Knows What You Are Looking For Before Generation — Yuhong Li, Yingbing Huang, Bowen Yang, Bharat Venkitesh, Acyr Locatelli, Hanchen Ye, Tianle Cai, Patrick Lewis, Deming Chen, 2024
https://scholar.google.com/scholar?q=SnapKV%3A+LLM+Knows+What+You+Are+Looking+For+Before+Generation
6. LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models — Huiqiang Jiang, Qianhui Wu, Chin-Yew Lin, Yuqing Yang, Lili Qiu, 2023 (EMNLP 2023)
https://scholar.google.com/scholar?q=LLMLingua%3A+Compressing+Prompts+for+Accelerated+Inference+of+Large+Language+Models
7. LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression — Huiqiang Jiang, Qianhui Wu, Xufang Luo, Dongsheng Li, Chin-Yew Lin, Yuqing Yang, Lili Qiu, 2024 (ACL 2024)
https://scholar.google.com/scholar?q=LongLLMLingua%3A+Accelerating+and+Enhancing+LLMs+in+Long+Context+Scenarios+via+Prompt+Compression
8. LLMLingua-2: Data Distillation for Efficient and Faithful Task-Agnostic Prompt Compression — Zhuoshi Pan, Qianhui Wu, Huiqiang Jiang, Menglin Xia, Xufang Luo, Jue Zhang, Qingwei Lin, Victor Ruhle, Yuqing Yang, Chin-Yew Lin, H. Vicky Zhao, Lili Qiu, Dongmei Zhang, 2024 (ACL Findings 2024)
https://scholar.google.com/scholar?q=LLMLingua-2%3A+Data+Distillation+for+Efficient+and+Faithful+Task-Agnostic+Prompt+Compression
9. Learning to Compress Prompts with Gist Tokens — Jesse Mu, Xiang Lisa Li, Noah Goodman, 2023 (NeurIPS 2023)
https://scholar.google.com/scholar?q=Learning+to+Compress+Prompts+with+Gist+Tokens
10. Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation — Jingyu Liu, Beidi Chen, Ce Zhang, 2025
https://scholar.google.com/scholar?q=Speculative+Prefill%3A+Turbocharging+TTFT+with+Lightweight+and+Training-Free+Token+Importance+Estimation
11. SnapStream: Efficient Long Sequence Decoding on Dataflow Accelerators — Jonathan Li, Nasim Farahini, et al. (SambaNova), 2025
https://scholar.google.com/scholar?q=SnapStream%3A+Efficient+Long+Sequence+Decoding+on+Dataflow+Accelerators
12. LLMs Know What to Drop: Self-Attention Guided KV Cache Eviction for Efficient Long-Context Inference — Guangtao Wang, Shubhangi Upasani, Chen Wu, et al. (SambaNova), 2025
https://scholar.google.com/scholar?q=LLMs+Know+What+to+Drop%3A+Self-Attention+Guided+KV+Cache+Eviction+for+Efficient+Long-Context+Inference
13. MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention — Huiqiang Jiang, Yucheng Li, Chengruidong Zhang, et al., 2024
https://scholar.google.com/scholar?q=MInference+1.0%3A+Accelerating+Pre-filling+for+Long-Context+LLMs+via+Dynamic+Sparse+Attention
Interactive Visualization: Cross-Family Speculative Prefill Cuts Long-Context Latency

This episode explores AdaJEPA, an adaptive latent world model that challenges the standard "train once, freeze forever" assumption behind robot planning systems. The hosts trace the technical lineage from Yann LeCun's Joint-Embedding Predictive Architecture concept through model predictive control's decades-old roots in process engineering and rocket landing, showing how these pieces combine to let a deployed robot keep updating its internal model using only the consequences of its own actions — no new labels, demonstrations, or retraining pipeline required. Central to the discussion is how distribution shift causes small prediction errors to compound across multi-step planning horizons, and how test-time adaptation, borrowed from image classification and paralleled to cerebellar motor learning, closes that loop by treating each observed transition as a live training example. The conversation grounds abstract control theory in concrete deployment scenarios, from unfamiliar object shapes to shifting friction and lighting. Listeners interested in robotics, control theory, or self-supervised learning will find a clear walkthrough of why frozen world models fail in the wild and what it means for a model to keep learning after "training" officially ends.

Sources:
1. AdaJEPA: An Adaptive Latent World Model — Ying Wang, Oumayma Bounou, Yann LeCun, Mengye Ren, 2026
http://arxiv.org/abs/2606.32026
2. Model Predictive Control: Theory and Practice — A Survey — Carlos E. García, David M. Prett, Manfred Morari, 1989
https://scholar.google.com/scholar?q=Model+Predictive+Control%3A+Theory+and+Practice+%E2%80%94+A+Survey
3. Model Predictive Control: Classical, Robust and Stochastic — Basil Kouvaritakis, Mark Cannon, 2016
https://scholar.google.com/scholar?q=Model+Predictive+Control%3A+Classical%2C+Robust+and+Stochastic
4. Deep Reinforcement Learning in a Handful of Trials using Probabilistic Dynamics Models (PETS) — Kurtland Chua, Roberto Calandra, Rowan McAllister, Sergey Levine, 2018
https://scholar.google.com/scholar?q=Deep+Reinforcement+Learning+in+a+Handful+of+Trials+using+Probabilistic+Dynamics+Models+%28PETS%29
5. Learning Latent Dynamics for Planning from Pixels (PlaNet) — Danijar Hafner, Timothy Lillicrap, Ian Fischer, Ruben Villegas, David Ha, Honglak Lee, James Davidson, 2019
https://scholar.google.com/scholar?q=Learning+Latent+Dynamics+for+Planning+from+Pixels+%28PlaNet%29
6. Dino-wm: World models on pre-trained visual features enable zero-shot planning — Zhou, G., Pan, H., LeCun, Y., and Pinto, L., 2025
https://scholar.google.com/scholar?q=Dino-wm%3A+World+models+on+pre-trained+visual+features+enable+zero-shot+planning
7. Temporal straightening for latent planning — Wang, Y., Bounou, O., Zhou, G., Balestriero, R., Rudner, T. G., LeCun, Y., and Ren, M., 2026
https://scholar.google.com/scholar?q=Temporal+straightening+for+latent+planning
8. Closing the train-test gap in world models for gradient-based planning — Parthasarathy, A., Kalra, N., Agrawal, R., LeCun, Y., Bounou, O., Izmailov, P., and Goldblum, M., 2025
https://scholar.google.com/scholar?q=Closing+the+train-test+gap+in+world+models+for+gradient-based+planning
9. Td-mpc2: Scalable, robust world models for continuous control — Hansen, N., Su, H., and Wang, X., 2024
https://scholar.google.com/scholar?q=Td-mpc2%3A+Scalable%2C+robust+world+models+for+continuous+control
10. Adawm: Adaptive world model based planning for autonomous driving — Wang, H., Ye, X., Tao, F., Pan, C., Mallik, A., Yaman, B., Ren, L., and Zhang, J., 2025
https://scholar.google.com/scholar?q=Adawm%3A+Adaptive+world+model+based+planning+for+autonomous+driving
11. Test-time training with self-supervision for generalization under distribution shifts — Sun, Y., Wang, X., Liu, Z., Miller, J., Efros, A. A., and Hardt, M., 2020
https://scholar.google.com/scholar?q=Test-time+training+with+self-supervision+for+generalization+under+distribution+shifts
Interactive Visualization: AdaJEPA: Self-Adapting Latent World Models via Test-Time MPC

This episode explores Alpha-RTL, a framework applying test-time training to RTL hardware optimization, where an LLM updates its own weights live for each chip design using real EDA toolchain feedback rather than a static, pre-trained policy. The discussion contrasts this approach with two existing camps: agentic search methods (like REvolution) that iterate over a frozen model and discard synthesis feedback after each run, and training-time reinforcement learning (like ChipSeek) that learns once offline and only samples at inference. It unpacks why functional correctness in Verilog is a weak proxy for what chip teams actually optimize — PPA, the area-delay-power product measured only after synthesis — and traces the paper's core techniques back to their origins: test-time training from Sun et al.'s 2020 UC Berkeley work, and PUCT search from Kocsis and Szepesvári's 2006 UCT paper, extended here into a persistent state pool of Verilog candidates refined over gradient updates rather than resampled from scratch. Listeners interested in the mechanics of closing the loop between LLM code generation and physical design constraints — and the unusual tradeoff of burning GPU-hours to fine-tune a model for a single, disposable hardware block — will find the episode's breakdown of RLVR-style staged verification (compile, simulate, synthesize) particularly useful.

Sources:
1. Alpha-RTL: Test-Time Training for RTL Hardware Optimization — Peilong Zhou, Zhirong Chen, Cangyuan Li, Haoyu Gao, Kaiyan Chang, Ziming Qu, Ying Wang, 2026
http://arxiv.org/abs/2606.05253
2. Bandit based Monte-Carlo Planning — Levente Kocsis, Csaba Szepesvári, 2006
https://scholar.google.com/scholar?q=Bandit+based+Monte-Carlo+Planning
3. A Survey of Monte Carlo Tree Search Methods — Cameron Browne, Edward Powley, Daniel Whitehouse, Simon Lucas, Peter Cowling, Philipp Rohlfshagen, Stephen Tavener, Diego Perez, Spyridon Samothrakis, Simon Colton, 2012
https://scholar.google.com/scholar?q=A+Survey+of+Monte+Carlo+Tree+Search+Methods
4. Mastering the game of Go with deep neural networks and tree search — David Silver, Aja Huang, Chris J. Maddison, Arthur Guez, Laurent Sifre, et al. (DeepMind), 2016
https://scholar.google.com/scholar?q=Mastering+the+game+of+Go+with+deep+neural+networks+and+tree+search
5. A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play — David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, et al. (DeepMind), 2018
https://scholar.google.com/scholar?q=A+general+reinforcement+learning+algorithm+that+masters+chess%2C+shogi%2C+and+Go+through+self-play
6. Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm (AlphaZero) — David Silver et al., 2017/2018
https://scholar.google.com/scholar?q=Mastering+Chess+and+Shogi+by+Self-Play+with+a+General+Reinforcement+Learning+Algorithm+%28AlphaZero%29
7. A Graph Placement Methodology for Fast Chip Design (AlphaChip) — Azalia Mirhoseini, Anna Goldie, Mustafa Yazgan, et al., 2021
https://scholar.google.com/scholar?q=A+Graph+Placement+Methodology+for+Fast+Chip+Design+%28AlphaChip%29
8. Data-Driven Offline Optimization for Architecting Hardware Accelerators (PRIME) — Aviral Kumar, Amir Yazdanbakhsh, Milad Hashemi, Kevin Swersky, Sergey Levine, 2021/2022
https://scholar.google.com/scholar?q=Data-Driven+Offline+Optimization+for+Architecting+Hardware+Accelerators+%28PRIME%29
9. SymbiYosys / eqy (formal equivalence checking for Yosys-based flows) — YosysHQ / Claire Wolf and contributors, ongoing
https://scholar.google.com/scholar?q=SymbiYosys+%2F+eqy+%28formal+equivalence+checking+for+Yosys-based+flows%29

Browse by month

Sep 2026 Aug 2026 Jul 2026 Jun 2026 May 2026 Apr 2026 Mar 2026 Feb 2026 Jan 2026 Dec 2025 Nov 2025 Oct 2025 Sep 2025 Aug 2025