This episode looks at RelayS2S, a dual-path design for real-time voice dialogue. A small, fast speech model speaks the first five words of a reply while a stronger text-based pipeline writes the rest. The discussion sets the target at roughly 200 milliseconds, the average gap between turns in human conversation. It contrasts slow but smart ASR→LLM→TTS cascades with fast but weaker full-duplex models like Moshi. The key empirical seed is the authors' finding that 82.5 to 95 percent of five-word prefixes from a weak speech model were contextually appropriate even when the full answer was not. The episode compares the approach to speculative decoding and stresses that it is not lossless. The draft comes from a different model family, the gate is a learned classifier that can be wrong, and spoken words can't be retracted. The only safeguards are the gate and a fallback to the plain cascade. Listeners get a clear account of why the five-word buffer sets a latency floor, how it relates to older work on incremental speech generation and self-repair, and why the choice of where to start the latency clock matters.

Sources:
1. RelayS2S: A Dual-Path Speculative Generation for Real-Time Dialogue — Long Mai, Junli Liang, 2026
http://arxiv.org/abs/2603.23346
2. Fast Inference from Transformers via Speculative Decoding — Yaniv Leviathan, Matan Kalman, Yossi Matias, 2023 (ICML)
https://scholar.google.com/scholar?q=Fast+Inference+from+Transformers+via+Speculative+Decoding
3. Accelerating Large Language Model Decoding with Speculative Sampling — Charlie Chen, Sebastian Borgeaud, Geoffrey Irving, Jean-Baptiste Lespiau, Laurent Sifre, John Jumper, 2023
https://scholar.google.com/scholar?q=Accelerating+Large+Language+Model+Decoding+with+Speculative+Sampling
4. Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads — Tianle Cai, Yuhong Li, Zhengyang Geng, Hongwu Peng, Jason D. Lee, Deming Chen, Tri Dao, 2024
https://scholar.google.com/scholar?q=Medusa%3A+Simple+LLM+Inference+Acceleration+Framework+with+Multiple+Decoding+Heads
5. Moshi: a speech-text foundation model for real-time dialogue — Alexandre Défossez, Laurent Mazaré, Manu Orsini, Amélie Royer, Patrick Pérez, Hervé Jégou, Edouard Grave, Neil Zeghidour, 2024
https://scholar.google.com/scholar?q=Moshi%3A+a+speech-text+foundation+model+for+real-time+dialogue
6. Generative Spoken Dialogue Language Modeling (dGSLM) — Tu Anh Nguyen, Eugene Kharitonov, Jade Copet, Yossi Adi, Wei-Ning Hsu, Ali Elkahky, Paden Tomasello, Robin Algayres, Benoit Sagot, Abdelrahman Mohamed, Emmanuel Dupoux, 2022/2023 (TACL)
https://scholar.google.com/scholar?q=Generative+Spoken+Dialogue+Language+Modeling+%28dGSLM%29
7. Language Model Can Listen While Speaking (LSLM) — Ziyang Ma, Yakun Song, Chenpeng Du, Jian Cong, Zhuo Chen, Yuping Wang, Yuxuan Wang, Xie Chen, 2024
https://scholar.google.com/scholar?q=Language+Model+Can+Listen+While+Speaking+%28LSLM%29
8. Timing in turn-taking and its implications for processing models of language — Stephen C. Levinson, Francisco Torreira, 2015
https://scholar.google.com/scholar?q=Timing+in+turn-taking+and+its+implications+for+processing+models+of+language
9. Robust Speech Recognition via Large-Scale Weak Supervision (Whisper) — Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, Ilya Sutskever, 2022
https://scholar.google.com/scholar?q=Robust+Speech+Recognition+via+Large-Scale+Weak+Supervision+%28Whisper%29
10. Talking Turns: Benchmarking Audio Foundation Models on Turn-Taking Dynamics — Siddhant Arora, Zhiyun Lu, Chung-Cheng Chiu, Ruoming Pang, Shinji Watanabe, 2025
https://scholar.google.com/scholar?q=Talking+Turns%3A+Benchmarking+Audio+Foundation+Models+on+Turn-Taking+Dynamics
11. Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities — Guan-Ting Lin, Jiachen Lian, Tingle Li, Qirui Wang, Gopala Anumanchipalli, Alexander H. Liu, Hung-yi Lee, 2025
https://scholar.google.com/scholar?q=Full-Duplex-Bench%3A+A+Benchmark+to+Evaluate+Full-duplex+Spoken+Dialogue+Models+on+Turn-taking+Capabilities
12. Universals and cultural variation in turn-taking in conversation — Tanya Stivers, N. J. Enfield, Penelope Brown, Christina Englert, Makoto Hayashi, Trine Heinemann, Gertie Hoymann, Federico Rossano, Jan Peter de Ruiter, Kyung-Eun Yoon, Stephen C. Levinson, 2009
https://scholar.google.com/scholar?q=Universals+and+cultural+variation+in+turn-taking+in+conversation
13. Beyond Turn-Based Interfaces: Synchronous LLMs as Full-Duplex Dialogue Agents — Bandhav Veluri, Benjamin N. Peloquin, Bokai Yu, Hongyu Gong, Shyamnath Gollakota, 2024
https://scholar.google.com/scholar?q=Beyond+Turn-Based+Interfaces%3A+Synchronous+LLMs+as+Full-Duplex+Dialogue+Agents
14. LLaMA-Omni: Seamless Speech Interaction with Large Language Models — Qingkai Fang, Shoutao Guo, Yan Zhou, Zhengrui Ma, Shaolei Zhang, Yang Feng, 2024
https://scholar.google.com/scholar?q=LLaMA-Omni%3A+Seamless+Speech+Interaction+with+Large+Language+Models
15. Freeze-Omni: A Smart and Low Latency Speech-to-Speech Dialogue Model with Frozen LLM — Xiong Wang, Yangze Li, Chaoyou Fu, et al., 2025
https://scholar.google.com/scholar?q=Freeze-Omni%3A+A+Smart+and+Low+Latency+Speech-to-Speech+Dialogue+Model+with+Frozen+LLM
16. KAME: Tandem Architecture for Enhancing Knowledge in Real-Time Speech-to-Speech Conversational AI — So Kuroki, Yotaro Kubo, Takuya Akiba, Yujin Tang, 2026
https://scholar.google.com/scholar?q=KAME%3A+Tandem+Architecture+for+Enhancing+Knowledge+in+Real-Time+Speech-to-Speech+Conversational+AI
17. Discourse-Aware Dual-Track Streaming Response for Low-Latency Spoken Dialogue Systems (DDTSR) — Siyuan Liu, Jiahui Xu, Feng Jiang, et al., 2026
https://scholar.google.com/scholar?q=Discourse-Aware+Dual-Track+Streaming+Response+for+Low-Latency+Spoken+Dialogue+Systems+%28DDTSR%29
18. PredGen: Accelerated Inference of Large Language Models through Input-Time Speculation for Real-Time Speech Interaction — Shufan Li, Aditya Grover, 2025
https://scholar.google.com/scholar?q=PredGen%3A+Accelerated+Inference+of+Large+Language+Models+through+Input-Time+Speculation+for+Real-Time+Speech+Interaction
19. ChipChat: Low-Latency Cascaded Conversational Agent in MLX — Tatiana Likhomanenko, Richard He Bai, Zijin Gu, et al., 2026
https://scholar.google.com/scholar?q=ChipChat%3A+Low-Latency+Cascaded+Conversational+Agent+in+MLX
20. SpeakStream: Streaming Text-to-Speech with Interleaved Data — Richard He Bai, Zijin Gu, Tatiana Likhomanenko, Navdeep Jaitly, 2025
https://scholar.google.com/scholar?q=SpeakStream%3A+Streaming+Text-to-Speech+with+Interleaved+Data
21. Full-duplex dialogue evaluation benchmarks (e.g., Full-Duplex-Bench) — Guan-Ting Lin et al., 2025
https://scholar.google.com/scholar?q=Full-duplex+dialogue+evaluation+benchmarks+%28e.g.%2C+Full-Duplex-Bench%29
22. Early-exit / confidence estimation for autoregressive models (calibration literature, e.g., selective prediction, 'Selective Classification for Deep Neural Networks') — Yonatan Geifman, Ran El-Yaniv, 2017
https://scholar.google.com/scholar?q=Early-exit+%2F+confidence+estimation+for+autoregressive+models+%28calibration+literature%2C+e.g.%2C+selective+prediction%2C+%27Selective+Classification+for+Deep+Neural+Networks%27%29
Interactive Visualization: RelayS2S: Dual-Path Speculative Generation for Real-Time Dialogue

This episode examines BitWhisper, a 2015 paper from Ben-Gurion University claiming that two air-gapped computers can exchange data using only CPU heat, with the receiving PC reading its own thermal sensors. It places the work against earlier covert channels (FM radio, ultrasound, electromagnetic and optical leakage), most of which can only leak data out, whereas this one claims half-duplex two-way signaling with no added hardware. The attack model is that malware already inside an isolated network has no way to receive commands, and the thermal link would supply one. The claimed range is 0 to 40 centimeters at 1 to 8 bits per hour, and the hosts flag the gap between "signals" and "bits" and the missing modulation and countermeasures sections in the draft. The discussion also covers the physics of workload-driven heating, the one-degree sensor resolution, and the two-room experiment setup with CPU and thermal-camera measurements. Listeners get a clear look at how far a physical side channel can be pushed, and at how narrow its practical limits are.

Sources:
1. BitWhisper: Covert Signaling Channel between Air-Gapped Computers using Thermal Manipulations — Mordechai Guri, Matan Monitz, Yisroel Mirski, Yuval Elovici, 2015
http://arxiv.org/abs/1503.07919
2. A Note on the Confinement Problem — Butler W. Lampson, 1973
https://scholar.google.com/scholar?q=A+Note+on+the+Confinement+Problem
3. AirHopper: Bridging the Air-Gap between Isolated Networks and Mobile Phones using Radio Frequencies — Mordechai Guri, Gabi Kedma, Assaf Kachlon, Yuval Elovici, 2014
https://scholar.google.com/scholar?q=AirHopper%3A+Bridging+the+Air-Gap+between+Isolated+Networks+and+Mobile+Phones+using+Radio+Frequencies
4. Ultrasonic Networking / On Covert Acoustical Mesh Networks in Air — Michael Hanspach, Michael Goetz, 2013
https://scholar.google.com/scholar?q=Ultrasonic+Networking+%2F+On+Covert+Acoustical+Mesh+Networks+in+Air
5. Air-Gap Security: Survey of Covert Channels (e.g., Air-Gap Covert Channels: A Survey and Taxonomy) — Mordechai Guri, Yuval Elovici and collaborators, 2018-2020
https://scholar.google.com/scholar?q=Air-Gap+Security%3A+Survey+of+Covert+Channels+%28e.g.%2C+Air-Gap+Covert+Channels%3A+A+Survey+and+Taxonomy%29
6. BitWhisper: Covert Signaling Channel between Air-Gapped Computers using Thermal Manipulations — Mordechai Guri, Matan Monitz, Yisroel Mirski, Yuval Elovici, 2015
https://scholar.google.com/scholar?q=BitWhisper%3A+Covert+Signaling+Channel+between+Air-Gapped+Computers+using+Thermal+Manipulations
7. Covert Channels through Thermal Emissions / Thermal Covert Channels on Multi-core Platforms — Ramya Jayaram Masti, Devendra Rai, Aanjhan Ranganathan, Christian Müller, Lothar Thiele, Srdjan Capkun, 2015
https://scholar.google.com/scholar?q=Covert+Channels+through+Thermal+Emissions+%2F+Thermal+Covert+Channels+on+Multi-core+Platforms
8. Hot Pixels / Frequency-based thermal channels (thermal side-channel work) — Various authors, 2015-2019
https://scholar.google.com/scholar?q=Hot+Pixels+%2F+Frequency-based+thermal+channels+%28thermal+side-channel+work%29
9. Stuxnet Dossier — Nicolas Falliere, Liam O Murchu, Eric Chien (Symantec), 2011
https://scholar.google.com/scholar?q=Stuxnet+Dossier
10. Exfiltration of Information from Air-Gapped Machines Using Monitor's LED Indicator / and related Guri et al. work (e.g., Air-Gap Exfiltration surveys) — Mordechai Guri, Boris Zadov, Yuval Elovici and others, 2014-2019
https://scholar.google.com/scholar?q=Exfiltration+of+Information+from+Air-Gapped+Machines+Using+Monitor%27s+LED+Indicator+%2F+and+related+Guri+et+al.+work+%28e.g.%2C+Air-Gap+Exfiltration+surveys%29
11. Data Exfiltration: A Review of External Attack Vectors and Countermeasures — Faheem Ullah, Matthew Edwards, Rajiv Ramdhany, Ruzanna Chitchyan, M. Ali Babar, Awais Rashid, 2018
https://scholar.google.com/scholar?q=Data+Exfiltration%3A+A+Review+of+External+Attack+Vectors+and+Countermeasures
12. A Mathematical Theory of Communication — Claude E. Shannon, 1948
https://scholar.google.com/scholar?q=A+Mathematical+Theory+of+Communication
13. An Introduction to Deep Learning for the Physical Layer — Timothy O'Shea, Jakob Hoydis, 2017
https://scholar.google.com/scholar?q=An+Introduction+to+Deep+Learning+for+the+Physical+Layer
14. Covert channel work on timing and protocol design (e.g., Cabuk, Brodley, Shields, 'IP Covert Timing Channels: Design and Detection') — Serdar Cabuk, Carla E. Brodley, Clay Shields, 2004
https://scholar.google.com/scholar?q=Covert+channel+work+on+timing+and+protocol+design+%28e.g.%2C+Cabuk%2C+Brodley%2C+Shields%2C+%27IP+Covert+Timing+Channels%3A+Design+and+Detection%27%29
15. Hot or Not: Revealing Hidden Services by their Clock Skew — Steven J. Murdoch, 2006
https://scholar.google.com/scholar?q=Hot+or+Not%3A+Revealing+Hidden+Services+by+their+Clock+Skew
16. On Covert Acoustical Mesh Networks in Air — Michael Hanspach, Michael Goetz, 2013
https://scholar.google.com/scholar?q=On+Covert+Acoustical+Mesh+Networks+in+Air
17. Soft Tempest: Hidden Data Transmission Using Electromagnetic Emanations — Markus G. Kuhn, Ross J. Anderson, 1998
https://scholar.google.com/scholar?q=Soft+Tempest%3A+Hidden+Data+Transmission+Using+Electromagnetic+Emanations
18. Whispers in the Hyper-space: High-speed Covert Channel Attacks in the Cloud — Zhenyu Wu, Zhang Xu, Haining Wang, 2012
https://scholar.google.com/scholar?q=Whispers+in+the+Hyper-space%3A+High-speed+Covert+Channel+Attacks+in+the+Cloud
19. Fansmitter: Acoustic Data Exfiltration from (Speakerless) Air-Gapped Computers — Mordechai Guri, Yosef Solewicz, Andrey Daidakulov, Yuval Elovici, 2016
https://scholar.google.com/scholar?q=Fansmitter%3A+Acoustic+Data+Exfiltration+from+%28Speakerless%29+Air-Gapped+Computers
20. Thermal covert channels on multi-core platforms — Ramya Jayaram Masti, Devendra Rai, Aanjhan Ranganathan, Christian Müller, Lothar Thiele, Srdjan Capkun, 2015
https://scholar.google.com/scholar?q=Thermal+covert+channels+on+multi-core+platforms
Interactive Visualization: BitWhisper: Covert Thermal Signaling Between Air-Gapped Computers

This episode examines "More Agents Is All You Need," a Tencent paper showing that sampling a single LLM multiple times and voting on the outputs can match the performance of a model roughly five times larger — no fine-tuning, no multi-agent debate, no specialized prompting required. The discussion traces this "Agent Forest" method back to classical ensemble techniques like Random Forests and connects it to inference-time compute scaling, contrasting it with precursors like self-consistency decoding (CoT-SC) and LLM-Debate. A key mechanism explored is similarity-weighted majority voting, which lets the same simple sampling-and-voting recipe generalize across very different tasks like math, multiple choice, and code generation. Listeners interested in cheap alternatives to scaling up model size, or in how far simple statistical tricks can push LLM performance, will find the systematic breakdown of when and why this scaling trend holds — and where it starts to fail — particularly compelling.

Sources:
1. More Agents Is All You Need — Junyou Li, Qin Zhang, Yangbin Yu, Qiang Fu, Deheng Ye, 2024
http://arxiv.org/abs/2402.05120
2. Self-Consistency Improves Chain of Thought Reasoning in Language Models — Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, Denny Zhou, 2022 (ICLR 2023)
https://scholar.google.com/scholar?q=Self-Consistency+Improves+Chain+of+Thought+Reasoning+in+Language+Models
3. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models — Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, Denny Zhou, 2022
https://scholar.google.com/scholar?q=Chain-of-Thought+Prompting+Elicits+Reasoning+in+Large+Language+Models
4. Rationale-Augmented Ensembles in Language Models — Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Denny Zhou, 2022
https://scholar.google.com/scholar?q=Rationale-Augmented+Ensembles+in+Language+Models
5. LLM-Debate: Improving Factuality and Reasoning in Language Models through Multiagent Debate — Yilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum, Igor Mordatch, 2023
https://scholar.google.com/scholar?q=LLM-Debate%3A+Improving+Factuality+and+Reasoning+in+Language+Models+through+Multiagent+Debate
6. AI Agents That Matter — Sayash Kapoor, Benedikt Stroebl, Zachary S. Siegel, Nitya Nadgir, Arvind Narayanan, 2024
https://scholar.google.com/scholar?q=AI+Agents+That+Matter
7. Improving Factuality and Reasoning in Language Models through Multiagent Debate — Yilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum, Igor Mordatch, 2023
https://scholar.google.com/scholar?q=Improving+Factuality+and+Reasoning+in+Language+Models+through+Multiagent+Debate
8. Blending Is All You Need: Cheaper, Better Alternative to Trillion-Parameters LLM — Xiaoding Lu, Adian Liusie, Vyas Raina, Yuwen Zhang, William Beauchamp, 2024
https://scholar.google.com/scholar?q=Blending+Is+All+You+Need%3A+Cheaper%2C+Better+Alternative+to+Trillion-Parameters+LLM
9. Knowledge Fusion of Large Language Models — Fanqi Wan, Xinting Huang, Deng Cai, Xiaojun Quan, Wei Bi, Shuming Shi, 2024
https://scholar.google.com/scholar?q=Knowledge+Fusion+of+Large+Language+Models
Interactive Visualization: More Agents Is All You Need: Sampling Beats Scaling

This episode examines "Language Modeling with Gated Convolutional Networks" (Dauphin et al., Facebook AI Research, ICML 2017), which replaces recurrent architectures with stacked causal convolutions for language modeling. The discussion covers why LSTMs dominated the field due to vanishing-gradient mitigation through gating, and contrasts that with the paper's Gated Linear Unit — a sigmoid-gated linear projection that preserves an undiminished gradient path, unlike the tanh-based gating in DeepMind's PixelCNN. A central debate weighs the RNN's theoretically unbounded context against the convolutional model's finite but better-conditioned receptive field, testing whether structured, hierarchical context outperforms raw sequential memory. The hosts walk through the model's pipeline — embeddings, bottlenecked causal convolution blocks with residual connections, and an adaptive softmax — and preview results on the Google Billion Word benchmark and the long-range WikiText-103 test. Listeners interested in the shift toward parallelizable, non-recurrent sequence models will find the tension between theoretical and practical memory capacity especially compelling.

Sources:
1. Language Modeling with Gated Convolutional Networks — Yann N. Dauphin, Angela Fan, Michael Auli, David Grangier, 2016
http://arxiv.org/abs/1612.08083
2. Exploring the Limits of Language Modeling — Jozefowicz, Vinyals, Schuster, Shazeer, Wu, 2016
https://scholar.google.com/scholar?q=Exploring+the+Limits+of+Language+Modeling
3. Conditional Image Generation with PixelCNN Decoders — van den Oord, Kalchbrenner, Vinyals, Espeholt, Graves, Kavukcuoglu, 2016
https://scholar.google.com/scholar?q=Conditional+Image+Generation+with+PixelCNN+Decoders
4. Neural Machine Translation in Linear Time — Kalchbrenner, Espeholt, Simonyan, van den Oord, Graves, Kavukcuoglu, 2016
https://scholar.google.com/scholar?q=Neural+Machine+Translation+in+Linear+Time
5. Improving Neural Language Models with a Continuous Cache — Grave, Joulin, Usunier, 2016
https://scholar.google.com/scholar?q=Improving+Neural+Language+Models+with+a+Continuous+Cache
6. Efficient softmax approximation for GPUs — Grave, Joulin, Cissé, Grangier, Jégou, 2016
https://scholar.google.com/scholar?q=Efficient+softmax+approximation+for+GPUs
7. Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer — Shazeer, Mirhoseini, Maziarz, Davis, Le, Hinton, Dean, 2017
https://scholar.google.com/scholar?q=Outrageously+Large+Neural+Networks%3A+The+Sparsely-Gated+Mixture-of-Experts+Layer
Interactive Visualization: Gated Convolutions Challenge RNNs in Language Modeling

This episode examines JustAskJev, a reinforcement-learning-trained detector that flags AI alignment failures using a single generic yes-or-no question asked in a zero-shot setting. It contrasts this approach with conventional generative LLM-as-judge and classifier-based detectors like Llama Guard, explaining how RLCD (reinforcement learning for calibrated decisions) lets one model process typed questions—yes/no, categorical, ordinal—against a shared "state" in a single call rather than requiring separate passes per failure type. The discussion covers the paper's headline result: a 0.886 median AUROC across ten distinct failure modes (sycophancy, jailbreaks, deception, prompt injection, hallucination, privacy violation, bias, reward hacking, concealed uncertainty, and power seeking) spanning forty-four benchmarks and five target models. It also unpacks calibration as a property distinct from accuracy, using sycophancy as a case study for why a model's stated confidence needs to track real-world correctness. Listeners interested in AI safety evaluation, model auditing costs, or the mechanics of confidence calibration will find the paper's claims about cheaper, unified failure detection a compelling departure from today's fragmented screening tools.

Sources:
1. Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures — Ruoqi Guo, Yi Liu, Gelei Deng, Yuekang Li, Lida Zhao, Yutao Wu, Simin Chen, Ying Zhang, Leo Yu Zhang, 2026
http://arxiv.org/abs/2609.29429
2. On Calibration of Modern Neural Networks — Chuan Guo, Geoff Pleiss, Yu Sun, Kilian Q. Weinberger, 2017
https://scholar.google.com/scholar?q=On+Calibration+of+Modern+Neural+Networks
3. Language Models (Mostly) Know What They Know — Saurav Kadavath, Tom Conerly, Amanda Askell, and others (Anthropic), 2022
https://scholar.google.com/scholar?q=Language+Models+%28Mostly%29+Know+What+They+Know
4. Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback — Katherine Tian, Eric Mitchell, Allan Zhou, Archit Sharma, Rafael Rafailov, Huaxiu Yao, Chelsea Finn, Christopher D. Manning, 2023
https://scholar.google.com/scholar?q=Just+Ask+for+Calibration%3A+Strategies+for+Eliciting+Calibrated+Confidence+Scores+from+Language+Models+Fine-Tuned+with+Human+Feedback
5. Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations — Hakan Inan, Kartikeya Upasani, and others (Meta), 2023
https://scholar.google.com/scholar?q=Llama+Guard%3A+LLM-based+Input-Output+Safeguard+for+Human-AI+Conversations
6. Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models — Carson Denison, Monte MacDiarmid, Fazl Barez, et al., 2024
https://scholar.google.com/scholar?q=Sycophancy+to+Subterfuge%3A+Investigating+Reward-Tampering+in+Large+Language+Models
7. The MASK Benchmark: Disentangling Honesty from Accuracy in AI Systems — Richard Ren, Arunim Agarwal, Mantas Mazeika, et al., 2025
https://scholar.google.com/scholar?q=The+MASK+Benchmark%3A+Disentangling+Honesty+from+Accuracy+in+AI+Systems
8. Confident Learning: Estimating Uncertainty in Dataset Labels — Curtis G. Northcutt, Lu Jiang, Isaac L. Chuang, 2021
https://scholar.google.com/scholar?q=Confident+Learning%3A+Estimating+Uncertainty+in+Dataset+Labels
Interactive Visualization: JustAskJev: One Question Detects Ten AI Failure Modes

This episode explores a paper arguing that KV cache compressibility is not an inherent property of an input, but a structural property of how a transformer's weights happen to represent a computation. Using a histogram-counting thought experiment, the authors prove that two transformers can compute the identical function while one is trivially compressible and the other is fundamentally not, exposing a blind spot in existing inference-time compression techniques like heavy-hitter eviction, attention-sink methods, and optimization-based approaches such as Cartridges and Attention Matching. The discussion surveys the broader landscape of memory-saving strategies, contrasting architectural fixes like linear attention and state space models against post-hoc interventions on standard transformers, and explains why long-context and agentic workloads make cache size a serious bottleneck. The hosts debate whether the paper's formal construction is just a toy example or a meaningful warning, concluding that it reframes compressibility as something that must be deliberately trained for rather than assumed to emerge naturally from ordinary pretraining. Listeners interested in efficient LLM serving will find this useful for understanding why some models resist compression no matter how sophisticated the algorithm applied to them.

Sources:
1. Training Transformers for KV Cache Compressibility — Yoav Gelberg, Yam Eitan, Michael Bronstein, Yarin Gal, Haggai Maron, 2026
http://arxiv.org/abs/2605.05971
2. GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints — Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebrón, Sumit Sanghai, 2023
https://scholar.google.com/scholar?q=GQA%3A+Training+Generalized+Multi-Query+Transformer+Models+from+Multi-Head+Checkpoints
3. H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models — Zhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen, Lianmin Zheng, Ruisi Cai, Zhao Song, Yuandong Tian, Christopher Ré, Clark Barrett, Zhangyang Wang, Beidi Chen, 2023
https://scholar.google.com/scholar?q=H2O%3A+Heavy-Hitter+Oracle+for+Efficient+Generative+Inference+of+Large+Language+Models
4. Efficient Streaming Language Models with Attention Sinks (StreamingLLM) — Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, Mike Lewis, 2023
https://scholar.google.com/scholar?q=Efficient+Streaming+Language+Models+with+Attention+Sinks+%28StreamingLLM%29
5. DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model — DeepSeek-AI (Aixin Liu et al.), 2024
https://scholar.google.com/scholar?q=DeepSeek-V2%3A+A+Strong%2C+Economical%2C+and+Efficient+Mixture-of-Experts+Language+Model
6. Cartridges: Lightweight and general-purpose long context representations via self-study — Sabri Eyuboglu, Ryan Ehrlich, Simran Arora, Neel Guha, Dylan Zinsley, Emily Liu, Will Tennien, Atri Rudra, James Zou, Azalia Mirhoseini, Christopher Ré, 2025
https://scholar.google.com/scholar?q=Cartridges%3A+Lightweight+and+general-purpose+long+context+representations+via+self-study
7. Fast KV compaction via attention matching — Adam Zweiger, Xinghong Fu, Han Guo, Yoon Kim, 2026
https://scholar.google.com/scholar?q=Fast+KV+compaction+via+attention+matching
8. Dynamic chunking for end-to-end hierarchical sequence modeling — Sukjun Hwang, Brandon Wang, Albert Gu, 2025
https://scholar.google.com/scholar?q=Dynamic+chunking+for+end-to-end+hierarchical+sequence+modeling
9. DuoAttention: Efficient long-context LLM inference with retrieval and streaming heads — Guangxuan Xiao, Jiaming Tang, Jingwei Zuo, Junxian Guo, Shang Yang, Haotian Tang, Yao Fu, Song Han, 2025
https://scholar.google.com/scholar?q=DuoAttention%3A+Efficient+long-context+LLM+inference+with+retrieval+and+streaming+heads
10. Lexico: Extreme KV cache compression via sparse coding over universal dictionaries — Junhyuck Kim, Jongho Park, Jaewoong Cho, Dimitris Papailiopoulos, 2025
https://scholar.google.com/scholar?q=Lexico%3A+Extreme+KV+cache+compression+via+sparse+coding+over+universal+dictionaries
Interactive Visualization: Training Transformers for Compressible KV Caches

This episode examines the Qwen team's design-rationale report for Qwen3.8-Flash-Next, a 125B-total, 6B-activated base model with 51B of n-gram embeddings held in host memory. It walks through the five components: a hybrid of Gated DeltaNet layers with periodic full attention, Qwen's sparse attention variant (QSA) as a response to the quadratic cost of DeepSeek-style indexers, a four-branch gated residual stream, a hashed n-gram lookup layer, and the Muon optimizer. The recurring question is whether loss, benchmark accuracy, cost, and stability agree. The paper reports that larger n-gram vocabularies always lower loss while accuracy stays flat. The hosts also note that most ablations are single runs with no reported seeds, and that the paper itself says top-setting margins likely fall within evaluation noise. The headline claim, matching or beating a 397B-A17B predecessor on eight of fourteen benchmarks at roughly a ninth of the training compute, is set against the ablation evidence behind it. The episode is useful for anyone who wants to see how the design choices trace to prior work (highway networks, Hyper-Connections, Engram, LongCat) and how much weight those choices can bear.

Sources:
1. On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability — Zihan Qiu, Zekun Wang, Xiao Li, Yanpeng Li, Yang Xu, Yixuan Wang, Huaqing Zhang, Rui Men, Bochao Mao, Chengruidong Zhang, Fan Zhou, Hao Luo, Haofeng Huang, Haoran Lian, Haoyan Huang, Hongqing Chen, Jianwei Zhang, Jing Xu, Junjie Wang, Langshi Chen, Liangyu Wang, Linlang Jiang, Man Yuan, Minmin Sun, Peng Jin, Siqi Zhang, Siyu Wang, Xingzhang Ren, Yakai Wang, Yi Zhang, Yiming Dong, Yizhong Cao, Yubo Ma, Yunfei Mao, Bo Zheng, Dayiheng Liu, 2026
http://arxiv.org/abs/2608.30320v1
2. Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models (Engram) — Xin Cheng et al. (DeepSeek-AI, with Peking University), 2026
https://scholar.google.com/scholar?q=Conditional+Memory+via+Scalable+Lookup%3A+A+New+Axis+of+Sparsity+for+Large+Language+Models+%28Engram%29
3. Scaling Embedding Layers in Language Models (SCONE) — Da Yu, Edith Cohen, Badih Ghazi, Yangsibo Huang, Pritish Kamath, Ravi Kumar, Daogao Liu, Chiyuan Zhang (Google), 2025
https://scholar.google.com/scholar?q=Scaling+Embedding+Layers+in+Language+Models+%28SCONE%29
4. Over-Tokenized Transformer: Vocabulary is Generally Worth Scaling — Hongzhi Huang, Defa Zhu, Banggu Wu, Yutao Zeng, Ya Wang, Qiyang Min, Xun Zhou (ByteDance Seed), 2025
https://scholar.google.com/scholar?q=Over-Tokenized+Transformer%3A+Vocabulary+is+Generally+Worth+Scaling
5. Memory Layers at Scale — Vincent-Pierre Berges, Barlas Oguz, Daniel Haziza, Wen-tau Yih, Luke Zettlemoyer, Gargi Ghosh (Meta FAIR), 2024
https://scholar.google.com/scholar?q=Memory+Layers+at+Scale
6. Scaling Embeddings Outperforms Scaling Experts in Language Models — Hong Liu, Jiaqi Zhang, Chao Wang, ... Xunliang Cai (Meituan LongCat), 2026
https://scholar.google.com/scholar?q=Scaling+Embeddings+Outperforms+Scaling+Experts+in+Language+Models
7. N-Grammer: Augmenting Transformers with Latent n-grams — Aurko Roy, Rohan Anil, Guangda Lai, et al., 2022
https://scholar.google.com/scholar?q=N-Grammer%3A+Augmenting+Transformers+with+Latent+n-grams
8. Scaling Embedding Layers in Language Models — Da Yu, Edith Cohen, Badih Ghazi, Yangsibo Huang, Pritish Kamath, Ravi Kumar, Daogao Liu, Chiyuan Zhang, 2025
https://scholar.google.com/scholar?q=Scaling+Embedding+Layers+in+Language+Models
9. L3: Large Lookup Layers — Albert Tseng, Christopher De Sa, 2026
https://scholar.google.com/scholar?q=L3%3A+Large+Lookup+Layers
10. STEM: Scaling Transformers with Embedding Modules — Ranajoy Sadhukhan, Sheng Cao, Harry Dong, Changsheng Zhao, Attiano Purpura-Pontoniere, Yuandong Tian, Zechun Liu, Beidi Chen, 2026
https://scholar.google.com/scholar?q=STEM%3A+Scaling+Transformers+with+Embedding+Modules
11. Large Memory Layers with Product Keys — Guillaume Lample, Alexandre Sablayrolles, Marc'Aurelio Ranzato, Ludovic Denoyer, Hervé Jégou, 2019
https://scholar.google.com/scholar?q=Large+Memory+Layers+with+Product+Keys
12. Mixture of A Million Experts (PEER) — Xu Owen He, 2024
https://scholar.google.com/scholar?q=Mixture+of+A+Million+Experts+%28PEER%29
13. Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws — Zeyuan Allen-Zhu, Yuanzhi Li, 2024
https://scholar.google.com/scholar?q=Physics+of+Language+Models%3A+Part+3.3%2C+Knowledge+Capacity+Scaling+Laws
14. Quantifying Memorization Across Neural Language Models — Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramèr, Chiyuan Zhang, 2023
https://scholar.google.com/scholar?q=Quantifying+Memorization+Across+Neural+Language+Models
15. Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention — Jingyang Yuan et al. (DeepSeek), 2025
https://scholar.google.com/scholar?q=Native+Sparse+Attention%3A+Hardware-Aligned+and+Natively+Trainable+Sparse+Attention
16. Kimi Linear: An Expressive, Efficient Attention Architecture — Kimi Team, 2025
https://scholar.google.com/scholar?q=Kimi+Linear%3A+An+Expressive%2C+Efficient+Attention+Architecture
17. Small-scale Proxies for Large-scale Transformer Training Instabilities — Mitchell Wortsman et al., 2023
https://scholar.google.com/scholar?q=Small-scale+Proxies+for+Large-scale+Transformer+Training+Instabilities
Interactive Visualization: Qwen3.8-Next Design: Hybrid Attention, Residuals, and N-gram Embeddings

This episode examines DeepSeek-AI's mHC: Manifold-Constrained Hyper-Connections, which tries to improve the transformer's residual connection, a piece of the architecture that has barely changed in a decade. It first explains why the plain skip path x + F(x) has been so hard to displace. It then covers ByteDance's Hyper-Connections, which widen the residual stream into four parallel streams with learnable read, write and mixing maps. The mixing matrix alone accounts for most of the reported loss gain. The discussion turns to the failure mode: the product of unconstrained mixing matrices across layers lets signal gain climb to roughly 3000 in a 27B model, alongside a loss surge and a memory-bandwidth bill. The fix constrains the mixing matrix to be doubly stochastic using the 1967 Sinkhorn-Knopp algorithm, which keeps the composite mapping bounded. The hosts also weigh whether a doubly stochastic matrix, which is not identity, can preserve the gradient-flow argument for identity skip paths. They cover the added kernel and pipeline engineering, which reportedly costs about 6.7% extra training time.

Sources:
1. mHC: Manifold-Constrained Hyper-Connections — Zhenda Xie, Yixuan Wei, Huanqi Cao, Chenggang Zhao, Chengqi Deng, Jiashi Li, Damai Dai, Huazuo Gao, Jiang Chang, Kuai Yu, Liang Zhao, Shangyan Zhou, Zhean Xu, Zhengyan Zhang, Wangding Zeng, Shengding Hu, Yuqing Wang, Jingyang Yuan, Lean Wang, Wenfeng Liang, 2025
http://arxiv.org/abs/2512.24880
2. Hyper-Connections — Defa Zhu, Hongzhi Huang, Zihao Huang, Yutao Zeng, Yunyao Mao, Banggu Wu, Qiyang Min, Xun Zhou, 2024
https://scholar.google.com/scholar?q=Hyper-Connections
3. Identity Mappings in Deep Residual Networks — Kaiming He, Xiangyu Zhang, Shaoqing Ren, Jian Sun, 2016
https://scholar.google.com/scholar?q=Identity+Mappings+in+Deep+Residual+Networks
4. Concerning nonnegative matrices and doubly stochastic matrices — Richard Sinkhorn, Paul Knopp, 1967
https://scholar.google.com/scholar?q=Concerning+nonnegative+matrices+and+doubly+stochastic+matrices
5. Sinkformers: Transformers with Doubly Stochastic Attention — Michael E. Sander, Pierre Ablin, Mathieu Blondel, Gabriel Peyré, 2022
https://scholar.google.com/scholar?q=Sinkformers%3A+Transformers+with+Doubly+Stochastic+Attention
6. Attention is Not All You Need: Pure Attention Loses Rank Doubly Exponentially with Depth — Yihe Dong, Jean-Baptiste Cordonnier, Andreas Loukas, 2021
https://scholar.google.com/scholar?q=Attention+is+Not+All+You+Need%3A+Pure+Attention+Loses+Rank+Doubly+Exponentially+with+Depth
7. Deeper Insights into Graph Convolutional Networks for Semi-Supervised Learning / Graph Neural Networks Exponentially Lose Expressive Power for Node Classification — Qimai Li, Zhichao Han, Xiao-Ming Wu; Kenta Oono, Taiji Suzuki, 2018 / 2020
https://scholar.google.com/scholar?q=Deeper+Insights+into+Graph+Convolutional+Networks+for+Semi-Supervised+Learning+%2F+Graph+Neural+Networks+Exponentially+Lose+Expressive+Power+for+Node+Classification
8. DenseFormer: Enhancing Information Flow in Transformers via Depth Weighted Averaging — Matteo Pagliardini, Amirkeivan Mohtashami, François Fleuret, Martin Jaggi, 2024
https://scholar.google.com/scholar?q=DenseFormer%3A+Enhancing+Information+Flow+in+Transformers+via+Depth+Weighted+Averaging
9. MUDDFormer: Breaking Residual Bottlenecks in Transformers via Multiway Dynamic Dense Connections — Da Xiao, Qingye Meng, Shengping Li, Xingyuan Yuan, 2025
https://scholar.google.com/scholar?q=MUDDFormer%3A+Breaking+Residual+Bottlenecks+in+Transformers+via+Multiway+Dynamic+Dense+Connections
10. Residual Matrix Transformers: Scaling the Size of the Residual Stream — Brian Mak, Jeffrey Flanigan, 2025
https://scholar.google.com/scholar?q=Residual+Matrix+Transformers%3A+Scaling+the+Size+of+the+Residual+Stream
11. DeepNet: Scaling Transformers to 1,000 Layers — Hongyu Wang, Shuming Ma, Li Dong, et al., 2022
https://scholar.google.com/scholar?q=DeepNet%3A+Scaling+Transformers+to+1%2C000+Layers
12. ReZero is All You Need: Fast Convergence at Large Depth — Thomas Bachlechner, Bodhisattwa Prasad Majumder, Huanru Henry Mao, Garrison W. Cottrell, Julian McAuley, 2020
https://scholar.google.com/scholar?q=ReZero+is+All+You+Need%3A+Fast+Convergence+at+Large+Depth
13. The Curse of Depth in Large Language Models — Wenfang Sun et al., 2025
https://scholar.google.com/scholar?q=The+Curse+of+Depth+in+Large+Language+Models
14. Value Residual Learning — Zhanchao Zhou et al., 2024
https://scholar.google.com/scholar?q=Value+Residual+Learning
15. Resurrecting the Sigmoid in Deep Learning through Dynamical Isometry / Deep Information Propagation — Jeffrey Pennington, Samuel Schoenholz, Surya Ganguli; Samuel Schoenholz et al., 2017
https://scholar.google.com/scholar?q=Resurrecting+the+Sigmoid+in+Deep+Learning+through+Dynamical+Isometry+%2F+Deep+Information+Propagation
16. DeepSeek-V3 Technical Report — DeepSeek-AI, 2024
https://scholar.google.com/scholar?q=DeepSeek-V3+Technical+Report
Interactive Visualization: mHC: Manifold-Constrained Hyper-Connections Stabilize Wider Residual Streams

This episode examines "You Only Cache Once," a decoder-decoder architecture from Microsoft Research and Tsinghua University that computes one global key-value cache and reuses it across every upper layer. The hosts explain why the KV cache becomes a deployment bottleneck at long context. They cite a 65B model needing about 86 GB of cache at 512K tokens even with grouped-query attention and 8-bit quantization. They then walk through the design: a bottom self-decoder uses constant-state efficient attention (sliding-window or gated retention), and a top cross-decoder cross-attends to the single shared cache. The result stays causal like a decoder-only model but cuts cache memory roughly L-fold and lets prefill skip the cross-decoder. The paper claims about an 80x smaller cache, 71.8x faster prefill at a million tokens, and 9.6x higher throughput at 512K with quality on par with a Transformer. The hosts also set out to test what those multipliers are measured against and whether every layer reading the same memory costs quality. It suits listeners interested in long-context serving costs and in how architecture design can address them.

Sources:
1. You Only Cache Once: Decoder-Decoder Architectures for Language Models — Yutao Sun, Li Dong, Yi Zhu, Shaohan Huang, Wenhui Wang, Shuming Ma, Quanlu Zhang, Jianyong Wang, Furu Wei, 2024
http://arxiv.org/abs/2405.05254v2
2. You Only Cache Once: Decoder-Decoder Architectures for Language Models — Yutao Sun, Li Dong, Yi Zhu, Shaohan Huang, Wenhui Wang, Shuming Ma, Quanlu Zhang, Jianyong Wang, Furu Wei, 2024
https://scholar.google.com/scholar?q=You+Only+Cache+Once%3A+Decoder-Decoder+Architectures+for+Language+Models
3. Attention Is All You Need — Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, Illia Polosukhin, 2017
https://scholar.google.com/scholar?q=Attention+Is+All+You+Need
4. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer (T5) — Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, Peter J. Liu, 2020
https://scholar.google.com/scholar?q=Exploring+the+Limits+of+Transfer+Learning+with+a+Unified+Text-to-Text+Transformer+%28T5%29
5. Decoder-Hybrid-Decoder Architecture for Efficient Reasoning with Long Generation (SambaY) — Liliang Ren, Congcong Chen, Haoran Xu, Young Jin Kim, Yelong Shen, Weizhu Chen, Jianfeng Gao and coauthors (Microsoft), 2025
https://scholar.google.com/scholar?q=Decoder-Hybrid-Decoder+Architecture+for+Efficient+Reasoning+with+Long+Generation+%28SambaY%29
6. Reducing Transformer Key-Value Cache Size with Cross-Layer Attention — William Brandon, Mayank Mishra, Aniruddha Nrusimha, Rameswar Panda, Jonathan Ragan-Kelley, 2024
https://scholar.google.com/scholar?q=Reducing+Transformer+Key-Value+Cache+Size+with+Cross-Layer+Attention
7. Layer-Condensed KV Cache for Efficient Inference of Large Language Models — Haoyi Wu, Kewei Tu, 2024
https://scholar.google.com/scholar?q=Layer-Condensed+KV+Cache+for+Efficient+Inference+of+Large+Language+Models
8. Fast Transformer Decoding: One Write-Head is All You Need, and GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints — Noam Shazeer (2019); Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebron, Sumit Sanghai (2023), 2019 and 2023
https://scholar.google.com/scholar?q=Fast+Transformer+Decoding%3A+One+Write-Head+is+All+You+Need%2C+and+GQA%3A+Training+Generalized+Multi-Query+Transformer+Models+from+Multi-Head+Checkpoints
9. DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model — DeepSeek-AI, 2024
https://scholar.google.com/scholar?q=DeepSeek-V2%3A+A+Strong%2C+Economical%2C+and+Efficient+Mixture-of-Experts+Language+Model
10. SwiftKV: Fast Prefill-Optimized Inference with Knowledge-Preserving Model Transformation — Aurick Qiao, Zhewei Yao, Samyam Rajbhandari, Yuxiong He (Snowflake AI Research), 2024
https://scholar.google.com/scholar?q=SwiftKV%3A+Fast+Prefill-Optimized+Inference+with+Knowledge-Preserving+Model+Transformation
11. Confident Adaptive Language Modeling (CALM) — Tal Schuster, Adam Fisch, Jai Gupta, Mostafa Dehghani, Dara Bahri, Vinh Q. Tran, Yi Tay, Donald Metzler, 2022
https://scholar.google.com/scholar?q=Confident+Adaptive+Language+Modeling+%28CALM%29
12. LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding — Mostafa Elhoushi, Akshat Shrivastava, Diana Liskovich, Basil Hosmer, Bram Wasti, Liangzhen Lai, Anas Mahmoud, Bilge Acun, Saurabh Agarwal, Ahmed Roman, Ahmed A. Aly, Beidi Chen, Carole-Jean Wu, 2024
https://scholar.google.com/scholar?q=LayerSkip%3A+Enabling+Early+Exit+Inference+and+Self-Speculative+Decoding
13. Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve — Amey Agrawal, Nitin Kedia, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav S. Gulavani, Alexey Tumanov, Ramachandran Ramjee, 2024
https://scholar.google.com/scholar?q=Taming+Throughput-Latency+Tradeoff+in+LLM+Inference+with+Sarathi-Serve
14. Retentive Network: A Successor to Transformer for Large Language Models — Yutao Sun, Li Dong, Shaohan Huang, Shuming Ma, Yuqing Xia, Jilong Xue, Jianyong Wang, Furu Wei, 2023
https://scholar.google.com/scholar?q=Retentive+Network%3A+A+Successor+to+Transformer+for+Large+Language+Models
15. Gated Linear Attention Transformers with Hardware-Efficient Training — Songlin Yang, Bailin Wang, Yikang Shen, Rameswar Panda, Yoon Kim, 2023
https://scholar.google.com/scholar?q=Gated+Linear+Attention+Transformers+with+Hardware-Efficient+Training
16. Mamba: Linear-Time Sequence Modeling with Selective State Spaces — Albert Gu, Tri Dao, 2023
https://scholar.google.com/scholar?q=Mamba%3A+Linear-Time+Sequence+Modeling+with+Selective+State+Spaces
17. Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention — Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, Francois Fleuret, 2020
https://scholar.google.com/scholar?q=Transformers+are+RNNs%3A+Fast+Autoregressive+Transformers+with+Linear+Attention
18. DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model (Multi-head Latent Attention) — DeepSeek-AI, 2024
https://scholar.google.com/scholar?q=DeepSeek-V2%3A+A+Strong%2C+Economical%2C+and+Efficient+Mixture-of-Experts+Language+Model+%28Multi-head+Latent+Attention%29
19. Jamba: A Hybrid Transformer-Mamba Language Model — Opher Lieber et al., 2024
https://scholar.google.com/scholar?q=Jamba%3A+A+Hybrid+Transformer-Mamba+Language+Model
20. Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models — Soham De et al., 2024
https://scholar.google.com/scholar?q=Griffin%3A+Mixing+Gated+Linear+Recurrences+with+Local+Attention+for+Efficient+Language+Models
21. Zoology: Measuring and Improving Recall in Efficient Language Models — Simran Arora, Sabri Eyuboglu, Aman Timalsina, Isys Johnson, Michael Poli, James Zou, Atri Rudra, Christopher Ré, 2023
https://scholar.google.com/scholar?q=Zoology%3A+Measuring+and+Improving+Recall+in+Efficient+Language+Models
22. Efficient Streaming Language Models with Attention Sinks (StreamingLLM) — Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, Mike Lewis, 2023
https://scholar.google.com/scholar?q=Efficient+Streaming+Language+Models+with+Attention+Sinks+%28StreamingLLM%29
23. H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models; and SnapKV — Zhenyu Zhang et al.; Yuhong Li et al., 2023/2024
https://scholar.google.com/scholar?q=H2O%3A+Heavy-Hitter+Oracle+for+Efficient+Generative+Inference+of+Large+Language+Models%3B+and+SnapKV
24. Efficiently Scaling Transformer Inference — Reiner Pope et al., 2022
https://scholar.google.com/scholar?q=Efficiently+Scaling+Transformer+Inference
25. RULER: What's the Real Context Size of Your Long-Context Language Models? — Cheng-Ping Hsieh et al., 2024
https://scholar.google.com/scholar?q=RULER%3A+What%27s+the+Real+Context+Size+of+Your+Long-Context+Language+Models%3F
Interactive Visualization: YOCO: Caching Once for Decoder-Decoder Long-Context Language Models

This episode examines Hyper-Connections (ICLR 2025, ByteDance Seed), which replaces the fixed residual connection in transformers with learned connection strengths. It traces the "seesaw" between Pre-Norm and Post-Norm: Pre-Norm gives stable gradients but suffers representation collapse in deep layers, while Post-Norm does the reverse. The method carries n parallel residual streams and learns depth-connections and width-connections, optionally predicted per token. It initializes as exactly Pre-Norm, so it strictly generalizes both variants. The paper's headline claim is 1.8x faster convergence and about six extra points on ARC-Challenge for OLMoE-1B-7B, for roughly 0.03% more parameters and 0.2% more FLOPs. The hosts place the idea against Highway Networks, DenseNet, and DenseFormer, and they flag that the collapse evidence rests largely on a single cosine-similarity figure. Listeners get a clear view of how a small learned wiring matrix can change a core piece of transformer design, along with a skeptical read of the reported gains.

Sources:
1. Hyper-Connections — Defa Zhu, Hongzhi Huang, Zihao Huang, Yutao Zeng, Yunyao Mao, Banggu Wu, Qiyang Min, Xun Zhou, 2024
http://arxiv.org/abs/2409.19606
2. Deep Residual Learning for Image Recognition — Kaiming He, Xiangyu Zhang, Shaoqing Ren, Jian Sun, 2016
https://scholar.google.com/scholar?q=Deep+Residual+Learning+for+Image+Recognition
3. On Layer Normalization in the Transformer Architecture — Ruibin Xiong, Yunchang Yang, Di He, Kai Zheng, Shuxin Zheng, Chen Xing, Huishuai Zhang, Yanyan Lan, Liwei Wang, Tie-Yan Liu, 2020
https://scholar.google.com/scholar?q=On+Layer+Normalization+in+the+Transformer+Architecture
4. Understanding the Difficulty of Training Transformers — Liyuan Liu, Xiaodong Liu, Jianfeng Gao, Weizhu Chen, Jiawei Han, 2020
https://scholar.google.com/scholar?q=Understanding+the+Difficulty+of+Training+Transformers
5. Densely Connected Convolutional Networks (DenseNet), with DenseFormer as the transformer analogue — Gao Huang, Zhuang Liu, Laurens van der Maaten, Kilian Q. Weinberger (DenseNet, 2017); Matteo Pagliardini, Amirkeivan Mohtashami, François Fleuret, Martin Jaggi (DenseFormer, 2024), 2017 / 2024
https://scholar.google.com/scholar?q=Densely+Connected+Convolutional+Networks+%28DenseNet%29%2C+with+DenseFormer+as+the+transformer+analogue
6. DenseFormer: Enhancing Information Flow in Transformers via Depth Weighted Averaging — Matteo Pagliardini, Amirkeivan Mohtashami, Francois Fleuret, Martin Jaggi, 2024
https://scholar.google.com/scholar?q=DenseFormer%3A+Enhancing+Information+Flow+in+Transformers+via+Depth+Weighted+Averaging
7. Highway Networks / Training Very Deep Networks — Rupesh Srivastava, Klaus Greff, Jürgen Schmidhuber, 2015
https://scholar.google.com/scholar?q=Highway+Networks+%2F+Training+Very+Deep+Networks
8. Densely Connected Convolutional Networks (DenseNet) — Gao Huang, Zhuang Liu, Laurens van der Maaten, Kilian Weinberger, 2017
https://scholar.google.com/scholar?q=Densely+Connected+Convolutional+Networks+%28DenseNet%29
9. DeepNet: Scaling Transformers to 1,000 Layers (DeepNorm) — Hongyu Wang et al., 2022
https://scholar.google.com/scholar?q=DeepNet%3A+Scaling+Transformers+to+1%2C000+Layers+%28DeepNorm%29
10. ReZero is All You Need: Fast Convergence at Large Depth — Thomas Bachlechner et al., 2021
https://scholar.google.com/scholar?q=ReZero+is+All+You+Need%3A+Fast+Convergence+at+Large+Depth
11. Peri-LN: Revisiting Layer Normalization in the Transformer Architecture / Sandwich-norm variants — Jeonghoon Kim et al., 2025
https://scholar.google.com/scholar?q=Peri-LN%3A+Revisiting+Layer+Normalization+in+the+Transformer+Architecture+%2F+Sandwich-norm+variants
12. The Curse of Depth in Large Language Models — Wenfang Sun et al., 2025
https://scholar.google.com/scholar?q=The+Curse+of+Depth+in+Large+Language+Models
13. Transformer Layers as Painters — Qi Sun, Marc Pickett, et al., 2024
https://scholar.google.com/scholar?q=Transformer+Layers+as+Painters
14. AltUp: Alternating Updates for Efficient Transformers — Cenk Baykal et al., 2023/2024
https://scholar.google.com/scholar?q=AltUp%3A+Alternating+Updates+for+Efficient+Transformers
15. ResiDual: Transformer with Dual Residual Connections — Shufang Xie et al., 2023
https://scholar.google.com/scholar?q=ResiDual%3A+Transformer+with+Dual+Residual+Connections
16. Value Residual Learning / ResFormer — Zhanchao Zhou et al., 2024
https://scholar.google.com/scholar?q=Value+Residual+Learning+%2F+ResFormer
17. A Mathematical Framework for Transformer Circuits (residual stream view) — Nelson Elhage et al., 2021
https://scholar.google.com/scholar?q=A+Mathematical+Framework+for+Transformer+Circuits+%28residual+stream+view%29
18. Reducing Activation Recomputation in Large Transformer Models — Vijay Korthikanti et al., 2022
https://scholar.google.com/scholar?q=Reducing+Activation+Recomputation+in+Large+Transformer+Models
19. mHC: Manifold-Constrained Hyper-Connections — DeepSeek-AI, 2025/2026
https://scholar.google.com/scholar?q=mHC%3A+Manifold-Constrained+Hyper-Connections
20. Frac-Connections / Virtual Width Networks (ByteDance Seed follow-ups) — ByteDance Seed team, 2025
https://scholar.google.com/scholar?q=Frac-Connections+%2F+Virtual+Width+Networks+%28ByteDance+Seed+follow-ups%29
Interactive Visualization: Hyper-Connections: Learning Residual Strengths for Faster Transformer Training

This episode explores "Subliminal Learning," a paper showing that a language model can pass behavioral traits to a student model through data that looks unrelated to the trait. In the headline experiment, a teacher prompted to love owls generates only number sequences. After filtering, a fresh copy of the base model is finetuned on those numbers, and its share of "owl" answers to a favorite-animal question rises from about 12% to over 60%. The hosts place this against earlier work: Hinton's "dark knowledge" in distillation, emergent misalignment from finetuning on insecure code, and adversarial examples as invisible predictive features. They also cover how the teacher-student pipeline works and why the effect seems to depend on a shared initialization. The episode matters for anyone who assumes that filtering distilled training data is enough to keep unwanted traits out, and it raises the possibility that some emergent misalignment is subliminal learning rather than a result of what the data says.

Sources:
1. Subliminal Learning: Language models transmit behavioral traits via hidden signals in data — Alex Cloud, Minh Le, James Chua, Jan Betley, Anna Sztyber-Betley, Jacob Hilton, Samuel Marks, Owain Evans, 2025
http://arxiv.org/abs/2507.14805
2. Distilling the Knowledge in a Neural Network — Geoffrey Hinton, Oriol Vinyals, Jeff Dean, 2015
https://scholar.google.com/scholar?q=Distilling+the+Knowledge+in+a+Neural+Network
3. Adversarial Examples Are Not Bugs, They Are Features — Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Logan Engstrom, Brandon Tran, Aleksander Madry, 2019
https://scholar.google.com/scholar?q=Adversarial+Examples+Are+Not+Bugs%2C+They+Are+Features
4. Preventing Language Models From Hiding Their Reasoning — Fabien Roger, Ryan Greenblatt, 2023
https://scholar.google.com/scholar?q=Preventing+Language+Models+From+Hiding+Their+Reasoning
5. AI Models Collapse When Trained on Recursively Generated Data — Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Nicolas Papernot, Ross Anderson, Yarin Gal, 2024
https://scholar.google.com/scholar?q=AI+Models+Collapse+When+Trained+on+Recursively+Generated+Data
6. Emergent Misalignment: Narrow Finetuning Can Produce Broadly Misaligned LLMs — Jan Betley, Daniel Tan, Niels Warncke, Anna Sztyber-Betley, Xuchan Bao, Martín Soto, Nathan Labenz, Owain Evans, 2025
https://scholar.google.com/scholar?q=Emergent+Misalignment%3A+Narrow+Finetuning+Can+Produce+Broadly+Misaligned+LLMs
7. Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training — Evan Hubinger, Carson Denison, Jesse Mu, Mike Lambert, Meg Tong, Monte MacDiarmid, et al., 2024
https://scholar.google.com/scholar?q=Sleeper+Agents%3A+Training+Deceptive+LLMs+that+Persist+Through+Safety+Training
8. Poisoning Web-Scale Training Datasets is Practical — Nicholas Carlini, Matthew Jagielski, Christopher A. Choquette-Choo, Daniel Paleka, Will Pearce, Hyrum Anderson, Andreas Terzis, Kurt Thomas, Florian Tramèr, 2023
https://scholar.google.com/scholar?q=Poisoning+Web-Scale+Training+Datasets+is+Practical
9. Persona Features Control Emergent Misalignment — Miles Wang, Tom Dupré la Tour, Olivia Watkins, Alex Makelov, et al. (OpenAI), 2025
https://scholar.google.com/scholar?q=Persona+Features+Control+Emergent+Misalignment
10. Born Again Neural Networks — Tommaso Furlanello, Zachary Lipton, Michael Tschannen, Laurent Itti, Anima Anandkumar, 2018
https://scholar.google.com/scholar?q=Born+Again+Neural+Networks
11. Distillation Robustifies Unlearning — Bruce W. Lee, Addie Foote, Alex Infanger, Leni Shor, Harish Kamath, Jacob Goldman-Wetzler, Bryce Woodworth, Alex Cloud, Alexander Matt Turner, 2025
https://scholar.google.com/scholar?q=Distillation+Robustifies+Unlearning
12. Model Organisms for Emergent Misalignment — Edward Turner, Anna Soligo, Mia Taylor, Senthooran Rajamanoharan, Neel Nanda, 2025
https://scholar.google.com/scholar?q=Model+Organisms+for+Emergent+Misalignment
13. Alignment Faking in Large Language Models — Ryan Greenblatt, Carson Denison, Benjamin Wright, Fabien Roger, et al., 2024
https://scholar.google.com/scholar?q=Alignment+Faking+in+Large+Language+Models
14. Unnatural Languages Are Not Bugs but Features for LLMs — Keyu Duan, Yiran Zhao, Zhili Feng, et al., 2025
https://scholar.google.com/scholar?q=Unnatural+Languages+Are+Not+Bugs+but+Features+for+LLMs
15. Linear Mode Connectivity and the Lottery Ticket Hypothesis — Jonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, Michael Carbin, 2020
https://scholar.google.com/scholar?q=Linear+Mode+Connectivity+and+the+Lottery+Ticket+Hypothesis
16. What is being transferred in transfer learning? — Behnam Neyshabur, Hanie Sedghi, Chiyuan Zhang, 2020
https://scholar.google.com/scholar?q=What+is+being+transferred+in+transfer+learning%3F
17. Sequence-Level Knowledge Distillation — Yoon Kim, Alexander M. Rush, 2016
https://scholar.google.com/scholar?q=Sequence-Level+Knowledge+Distillation
18. Poison Frogs! Targeted Clean-Label Poisoning Attacks on Neural Networks — Ali Shafahi, W. Ronny Huang, Mahyar Najibi, et al., 2018
https://scholar.google.com/scholar?q=Poison+Frogs%21+Targeted+Clean-Label+Poisoning+Attacks+on+Neural+Networks
19. Towards Understanding Subliminal Learning: When and How Hidden Biases Transfer (follow-up analyses of the effect) — Schrodi et al., 2025
https://scholar.google.com/scholar?q=Towards+Understanding+Subliminal+Learning%3A+When+and+How+Hidden+Biases+Transfer+%28follow-up+analyses+of+the+effect%29
Interactive Visualization: Subliminal Learning: Hidden Behavioral Traits Transmitted Through Model-Generated Data

This episode explores "Next-Latent Prediction Transformers Learn Compact World Models," a Microsoft Research paper proposing a small auxiliary loss that pushes an ordinary transformer to compress its history into a compact belief state, with no architecture changes. It covers why transformers lack the built-in compression that recurrent networks get from a fixed-size state, using the Manhattan taxi study where models reached 100 percent next-turn accuracy while their internal street maps were incoherent. It also covers the Clever Hans failure of myopic next-token training and the earlier fixes from the same group. The Belief State Transformer offers a guarantee at more than double the parameters, and joint multi-token prediction is cheaper but depends on an unknown k-observability horizon. NextLat borrows from reinforcement learning by training a small dynamics network to predict the next hidden state from the current state and token, which also allows self-speculative decoding with flexible draft lengths. The discussion then turns to the theorem behind the method and the conditions needed for the latents to converge to belief states.

Sources:
1. Next-Latent Prediction Transformers Learn Compact World Models — Jayden Teoh, Manan Tomar, Kwangjun Ahn, Edward S. Hu, Tim Pearce, Pratyusha Sharma, Akshay Krishnamurthy, Riashat Islam, Alex Lamb, John Langford, 2025
http://arxiv.org/abs/2511.05963
2. Bootstrap Your Own Latent: A New Approach to Self-Supervised Learning — Jean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec, Pierre H. Richemond, et al., 2020
https://scholar.google.com/scholar?q=Bootstrap+Your+Own+Latent%3A+A+New+Approach+to+Self-Supervised+Learning
3. Data-Efficient Reinforcement Learning with Self-Predictive Representations — Max Schwarzer, Ankesh Anand, Rishab Goel, R. Devon Hjelm, Aaron Courville, Philip Bachman, 2021
https://scholar.google.com/scholar?q=Data-Efficient+Reinforcement+Learning+with+Self-Predictive+Representations
4. Better & Faster Large Language Models via Multi-token Prediction — Fabian Gloeckle, Badr Youbi Idrissi, Baptiste Rozière, David Lopez-Paz, Gabriel Synnaeve, 2024
https://scholar.google.com/scholar?q=Better+%26+Faster+Large+Language+Models+via+Multi-token+Prediction
5. A Path Towards Autonomous Machine Intelligence (JEPA position paper) — Yann LeCun, 2022
https://scholar.google.com/scholar?q=A+Path+Towards+Autonomous+Machine+Intelligence+%28JEPA+position+paper%29
6. Planning and Acting in Partially Observable Stochastic Domains — Leslie Pack Kaelbling, Michael L. Littman, Anthony R. Cassandra, 1998
https://scholar.google.com/scholar?q=Planning+and+Acting+in+Partially+Observable+Stochastic+Domains
7. Predictive Representations of State — Michael L. Littman, Richard S. Sutton, Satinder Singh, 2001
https://scholar.google.com/scholar?q=Predictive+Representations+of+State
8. The Belief State Transformer — Edward S. Hu, Kwangjun Ahn, Qinghua Liu, Haoran Xu, Manan Tomar, Ada Langford, Dinesh Jayaraman, Alex Lamb, John Langford, 2024
https://scholar.google.com/scholar?q=The+Belief+State+Transformer
9. Transformers Represent Belief State Geometry in their Residual Stream — Adam S. Shai, Sarah E. Marzen, Lucas Teixeira, Alexander Gietelink Oldenziel, Paul M. Riechers, 2024
https://scholar.google.com/scholar?q=Transformers+Represent+Belief+State+Geometry+in+their+Residual+Stream
10. Fast Inference from Transformers via Speculative Decoding — Yaniv Leviathan, Matan Kalman, Yossi Matias, 2023
https://scholar.google.com/scholar?q=Fast+Inference+from+Transformers+via+Speculative+Decoding
11. Accelerating Large Language Model Decoding with Speculative Sampling — Charlie Chen, Sebastian Borgeaud, Geoffrey Irving, Jean-Baptiste Lespiau, Laurent Sifre, John Jumper, 2023
https://scholar.google.com/scholar?q=Accelerating+Large+Language+Model+Decoding+with+Speculative+Sampling
12. Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads — Tianle Cai, Yuhong Li, Zhengyang Geng, Hongwu Peng, Jason D. Lee, Deming Chen, Tri Dao, 2024
https://scholar.google.com/scholar?q=Medusa%3A+Simple+LLM+Inference+Acceleration+Framework+with+Multiple+Decoding+Heads
13. EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty — Yuhui Li, Fangyun Wei, Chao Zhang, Hongyang Zhang, 2024
https://scholar.google.com/scholar?q=EAGLE%3A+Speculative+Sampling+Requires+Rethinking+Feature+Uncertainty
14. Efficient Joint Prediction of Multiple Future Tokens (JTP) — Kwangjun Ahn, Alex Lamb, John Langford, 2025
https://scholar.google.com/scholar?q=Efficient+Joint+Prediction+of+Multiple+Future+Tokens+%28JTP%29
15. Bridging State and History Representations: Understanding Self-Predictive RL — Tianwei Ni, Benjamin Eysenbach, Erfan SeyedSalehi, Michel Ma, Clement Gehring, Aditya Mahajan, Pierre-Luc Bacon, 2024
https://scholar.google.com/scholar?q=Bridging+State+and+History+Representations%3A+Understanding+Self-Predictive+RL
16. The Pitfalls of Next-Token Prediction — Gregor Bachmann, Vaishnavh Nagarajan, 2024
https://scholar.google.com/scholar?q=The+Pitfalls+of+Next-Token+Prediction
17. Evaluating the World Model Implicit in a Generative Model — Keyon Vafa, Justin Y. Chen, Ashesh Rambachan, Jon Kleinberg, Sendhil Mullainathan, 2024
https://scholar.google.com/scholar?q=Evaluating+the+World+Model+Implicit+in+a+Generative+Model
18. EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty / EAGLE-2 — Yuhui Li, Fangyun Wei, Chao Zhang, Hongyang Zhang, 2024
https://scholar.google.com/scholar?q=EAGLE%3A+Speculative+Sampling+Requires+Rethinking+Feature+Uncertainty+%2F+EAGLE-2
19. Data-Efficient Reinforcement Learning with Self-Predictive Representations (SPR) — Max Schwarzer, Ankesh Anand, Rishab Goel, R Devon Hjelm, Aaron Courville, Philip Bachman, 2021
https://scholar.google.com/scholar?q=Data-Efficient+Reinforcement+Learning+with+Self-Predictive+Representations+%28SPR%29
20. The Parallelism Tradeoff: Limitations of Log-Precision Transformers — William Merrill, Ashish Sabharwal, 2023
https://scholar.google.com/scholar?q=The+Parallelism+Tradeoff%3A+Limitations+of+Log-Precision+Transformers
21. The Illusion of State in State-Space Models — William Merrill, Jackson Petty, Ashish Sabharwal, 2025
https://scholar.google.com/scholar?q=The+Illusion+of+State+in+State-Space+Models
22. LLM-JEPA: Large Language Models Meet Joint Embedding Predictive Architectures — Hai Huang, Yann LeCun, Randall Balestriero, 2025
https://scholar.google.com/scholar?q=LLM-JEPA%3A+Large+Language+Models+Meet+Joint+Embedding+Predictive+Architectures
23. Emergent Representations of Program Semantics / Transformers Represent Belief State Geometry in their Residual Stream — Adam S. Shai, Sarah E. Marzen, Lucas Teixeira, Alexander Gietelmark Oldenziel, Paul M. Riechers, 2024
https://scholar.google.com/scholar?q=Emergent+Representations+of+Program+Semantics+%2F+Transformers+Represent+Belief+State+Geometry+in+their+Residual+Stream
24. Multi-token prediction in DeepSeek-V3 (Technical Report) — DeepSeek-AI (Aixin Liu et al.), 2024
https://scholar.google.com/scholar?q=Multi-token+prediction+in+DeepSeek-V3+%28Technical+Report%29
Interactive Visualization: Next-Latent Prediction Lets Transformers Learn Compact World Models

This episode dissects the BASED architecture from Arora, Eyuboglu, Zhang et al. (Stanford/Buffalo), examining how it tries to resolve the tradeoff between recall accuracy and inference throughput in sequence models. The discussion traces the lineage of alternatives to standard softmax attention — state space models like Mamba, gated-convolution approaches like H3 and Hyena, linear attention, and sliding window attention — explaining why a growing KV-cache makes attention memory-bound at scale, while fixed-size-state models trade away precise recall by construction. It highlights the MQAR benchmark from the Zoology paper as a tool for exposing this recall-versus-state-size Pareto frontier, showing that every architecture, including full attention, sits somewhere on that curve rather than escaping it. The hosts then unpack BASED's hybrid design, which combines a linear attention component for cheap global context with a small sliding window for exact local comparisons, and start evaluating whether this combination actually delivers on its claimed 24x throughput gain over FlashAttention-2. Listeners interested in efficient LLM inference will get a clear framework for why architecture choices around memory and recall are fundamentally linked rather than independently solvable engineering problems.

Sources:
1. Simple linear attention language models balance the recall-throughput tradeoff — Simran Arora, Sabri Eyuboglu, Michael Zhang, Aman Timalsina, Silas Alberti, Dylan Zinsley, James Zou, Atri Rudra, Christopher Ré, 2024
http://arxiv.org/abs/2402.18668
2. Zoology: Measuring and Improving Recall in Efficient Language Models — Simran Arora, Sabri Eyuboglu, Aman Timalsina, Isys Johnson, Michael Poli, James Zou, Atri Rudra, Christopher Ré, 2023
https://scholar.google.com/scholar?q=Zoology%3A+Measuring+and+Improving+Recall+in+Efficient+Language+Models
3. The Hedgehog & the Porcupine: Expressive Linear Attentions with Softmax Mimicry — Michael Zhang, Kush Bhatia, Hermann Kumbong, Christopher Ré, 2024
https://scholar.google.com/scholar?q=The+Hedgehog+%26+the+Porcupine%3A+Expressive+Linear+Attentions+with+Softmax+Mimicry
4. Gated Linear Attention Transformers with Hardware-Efficient Training — Songlin Yang, Bailin Wang, Yikang Shen, Rameswar Panda, Yoon Kim, 2023
https://scholar.google.com/scholar?q=Gated+Linear+Attention+Transformers+with+Hardware-Efficient+Training
5. Mamba: Linear-Time Sequence Modeling with Selective State Spaces — Albert Gu, Tri Dao, 2023
https://scholar.google.com/scholar?q=Mamba%3A+Linear-Time+Sequence+Modeling+with+Selective+State+Spaces
6. Mistral 7B — Albert Q. Jiang et al., 2023
https://scholar.google.com/scholar?q=Mistral+7B

This episode examines APEX, a persistent-memory learned index from researchers at CUHK, MIT, Microsoft Research, and Simon Fraser University, presented at VLDB 2022. It traces the collision of two research threads — Intel's Optane persistent memory, which sits on the DDR bus but requires manual cache-line flushing and fencing to guarantee crash durability, and "learned indexes," which replace B-trees with lightweight regression models trained to predict a key's position. The discussion covers why naive approaches fail: running the learned index ALEX directly on persistent memory offers speed but zero crash safety, while wrapping it in standard transactional logging (PMDK) restores consistency at the cost of saturating scarce write bandwidth. It builds toward APEX's core innovation, "probe-and-stash," a technique designed to avoid the record-shifting that makes both alternatives fragile, enabling claimed 15x faster inserts and roughly 42-millisecond crash recovery. Listeners interested in database internals, memory hierarchies, or the practical gap between promising ML systems research and production-ready engineering will find the trade-offs — and the hosts' friendly disagreement over why learned indexes haven't seen wider industry adoption — a compelling entry point into the topic.

Sources:
1. APEX: A High-Performance Learned Index on Persistent Memory — Baotong Lu, Jialin Ding, Eric Lo, Umar Farooq Minhas, Tianzheng Wang, 2021
http://arxiv.org/abs/2105.00683
2. The Case for Learned Index Structures — Tim Kraska, Alex Beutel, Ed H. Chi, Jeffrey Dean, Neoklis Polyzotis, 2018
https://scholar.google.com/scholar?q=The+Case+for+Learned+Index+Structures
3. ALEX: An Updatable Adaptive Learned Index — Jialin Ding, Umar Farooq Minhas, Jia Yu, Chi Zhang, Yinan Li, Chi Wang, Badrish Chandramouli, Johannes Gehrke, Donald Kossmann, David Lomet, 2020
https://scholar.google.com/scholar?q=ALEX%3A+An+Updatable+Adaptive+Learned+Index
4. Let's Talk About Storage & Recovery Methods for Non-Volatile Memory Database Systems — Joy Arulraj, Andrew Pavlo, Subramanya R. Dulloor, 2015
https://scholar.google.com/scholar?q=Let%27s+Talk+About+Storage+%26+Recovery+Methods+for+Non-Volatile+Memory+Database+Systems
5. An Empirical Guide to the Behavior and Use of Scalable Persistent Memory — Jian Yang, Juno Kim, Morteza Hoseinzadeh, Joseph Izraelevitz, Steve Swanson, 2020
https://scholar.google.com/scholar?q=An+Empirical+Guide+to+the+Behavior+and+Use+of+Scalable+Persistent+Memory
6. How Does Updatable Learned Index Perform on Non-Volatile Main Memory? — Leying Chen, Shimin Chen, 2021
https://scholar.google.com/scholar?q=How+Does+Updatable+Learned+Index+Perform+on+Non-Volatile+Main+Memory%3F
7. Dash: Scalable Hashing on Persistent Memory — Baotong Lu, Xiangpeng Hao, Tianzheng Wang, Eric Lo, 2020
https://scholar.google.com/scholar?q=Dash%3A+Scalable+Hashing+on+Persistent+Memory
8. FPTree: A Hybrid SCM-DRAM Persistent and Concurrent B-Tree for Storage Class Memory — Ismail Oukid, Johan Lasperas, Anisoara Nica, Thomas Willhalm, Wolfgang Lehner, 2016
https://scholar.google.com/scholar?q=FPTree%3A+A+Hybrid+SCM-DRAM+Persistent+and+Concurrent+B-Tree+for+Storage+Class+Memory
9. RECIPE: Converting Concurrent DRAM Indexes to Persistent-Memory Indexes — Se Kwon Lee, Jayashree Mohan, Sanidhya Kashyap, Taesoo Kim, Vijay Chidambaram, 2019
https://scholar.google.com/scholar?q=RECIPE%3A+Converting+Concurrent+DRAM+Indexes+to+Persistent-Memory+Indexes
Interactive Visualization: APEX: A High-Performance Learned Index on Persistent Memory

This episode examines SALI, a 2024 SIGMOD systems paper diagnosing why learned database indexes like ALEX and LIPP see throughput drop as thread counts rise, rather than scale with additional concurrency. The discussion traces the lineage from Google's 2018 Recursive Model Index through the buffer-based and model-based approaches that emerged to handle inserts, focusing on the shift-versus-chain tradeoff between ALEX's coarse-grained locking and LIPP's fine-grained, chain-based node design. The key finding is that fine-grained locking solves data contention but not statistics contention — shared counters tracking when a node needs reorganization become a cacheline-thrashing bottleneck under heavy concurrent writes. Listeners interested in database internals or concurrent data structures will find the breakdown of exactly where and why two well-regarded designs fail at scale particularly compelling, especially the setup for SALI's proposed fix using self-adapting nodes instead of shared bottleneck counters.

Sources:
1. SALI: A Scalable Adaptive Learned Index Framework based on Probability Models — Jiake Ge, Huanchen Zhang, Boyu Shi, Yuanhui Luo, Yunda Guo, Yunpeng Chai, Yuxing Chen, Anqun Pan, 2023
http://arxiv.org/abs/2308.15012
2. The Case for Learned Index Structures — Tim Kraska, Alex Beutel, Ed H. Chi, Jeffrey Dean, Neoklis Polyzotis, 2018
https://scholar.google.com/scholar?q=The+Case+for+Learned+Index+Structures
3. ALEX: An Updatable Adaptive Learned Index — Jialin Ding, Umar Farooq Minhas, Jia Yu, Chi Wang, Jaeyoung Do, Yinan Li, Hantian Zhang, Badrish Chandramouli, Johannes Gehrke, Donald Kossmann, David Lomet, Tim Kraska, 2020
https://scholar.google.com/scholar?q=ALEX%3A+An+Updatable+Adaptive+Learned+Index
4. Updatable Learned Index with Precise Positions (LIPP) — Jiacheng Wu, Yong Zhang, Shimin Chen, Jin Wang, Yu Chen, Chunxiao Xing, 2021
https://scholar.google.com/scholar?q=Updatable+Learned+Index+with+Precise+Positions+%28LIPP%29
5. XIndex: A Scalable Learned Index for Multicore Data Storage — Chuzhe Tang, Youyun Wang, Zhiyuan Dong, Gansen Hu, Zhaoguo Wang, Minjie Wang, Haibo Chen, 2020
https://scholar.google.com/scholar?q=XIndex%3A+A+Scalable+Learned+Index+for+Multicore+Data+Storage
6. Are Updatable Learned Indexes Ready? — Chaichon Wongkham, Baotong Lu, Chris Liu, Zhicong Zhong, Eric Lo, Tianzheng Wang, 2022
https://scholar.google.com/scholar?q=Are+Updatable+Learned+Indexes+Ready%3F
7. Adaptive Hybrid Indexes — Christoph Anneser, Andreas Kipf, Huanchen Zhang, Thomas Neumann, Alfons Kemper, 2022
https://scholar.google.com/scholar?q=Adaptive+Hybrid+Indexes
8. DILI: A Distribution-Driven Learned Index — Pengfei Li, Hua Lu, Rong Zhu, Bolin Ding, Long Yang, Gang Pan, 2023
https://scholar.google.com/scholar?q=DILI%3A+A+Distribution-Driven+Learned+Index
9. Updatable Learned Indexes Meet Disk-Resident DBMS: From Evaluations to Design Choices — Hai Lan, Zhifeng Bao, J. Shane Culpepper, Renata Borovica-Gajic, 2023
https://scholar.google.com/scholar?q=Updatable+Learned+Indexes+Meet+Disk-Resident+DBMS%3A+From+Evaluations+to+Design+Choices
Interactive Visualization: SALI: Fixing Concurrency Bottlenecks in Learned Indexes

This episode unpacks ALEX, a 2020 paper from MIT, Microsoft Research, Arizona State, and Georgia Tech that reworks the "learned index" idea for real-world databases. It traces the lineage from Kraska et al.'s 2018 Learned Index, which used a hierarchy of regression models to predict a key's position in a sorted array but only worked on static, read-only data since any insert would break the model's predictions. The discussion explains how ALEX solves this with a Gapped Array that leaves deliberate empty slots for near-free inserts, plus a tree structure whose nodes can grow, shrink, split, or retrain their local models on the fly, letting it support inserts, updates, and deletes alongside lookups. The headline numbers — up to 4.1x the throughput of a B+Tree with an index up to 2000 times smaller — anchor a broader explanation of why B+Trees have dominated databases since the 1970s and what it takes for a learned alternative to finally handle OLTP-style workloads. Listeners interested in database internals or machine learning applied to systems problems will get a clear, building-block explanation of both the classic B+Tree and the model-based alternative trying to replace it.

Sources:
1. ALEX: An Updatable Adaptive Learned Index — Jialin Ding, Umar Farooq Minhas, Jia Yu, Chi Wang, Jaeyoung Do, Yinan Li, Hantian Zhang, Badrish Chandramouli, Johannes Gehrke, Donald Kossmann, David Lomet, Tim Kraska, 2019
http://arxiv.org/abs/1905.08898
2. The Case for Learned Index Structures — Tim Kraska, Alex Beutel, Ed H. Chi, Jeffrey Dean, Neoklis Polyzotis, 2018
https://scholar.google.com/scholar?q=The+Case+for+Learned+Index+Structures
3. FITing-Tree: A Data-aware Index Structure — Alex Galakatos, Michael Markovitch, Carsten Binnig, Rodrigo Fonseca, Tim Kraska, 2019
https://scholar.google.com/scholar?q=FITing-Tree%3A+A+Data-aware+Index+Structure
4. The PGM-index: a fully-dynamic compressed learned index with provable worst-case bounds — Paolo Ferragina, Giorgio Vinciguerra, 2020
https://scholar.google.com/scholar?q=The+PGM-index%3A+a+fully-dynamic+compressed+learned+index+with+provable+worst-case+bounds
5. Benchmarking Learned Indexes — Ryan Marcus, Andreas Kipf, Alexander van Renen, Mihail Stoian, Sanchit Misra, Alfons Kemper, Thomas Neumann, Tim Kraska, 2020
https://scholar.google.com/scholar?q=Benchmarking+Learned+Indexes
6. A Sparse Table Implementation of Priority Queues — Alon Itai, Alan G. Konheim, Michael Rodeh, 1981
https://scholar.google.com/scholar?q=A+Sparse+Table+Implementation+of+Priority+Queues
7. Cache-Oblivious Streaming B-trees — Michael A. Bender, Martin Farach-Colton, Jeremy T. Fineman, Yonatan R. Fogel, Bradley C. Kuszmaul, Jelani Nelson, 2007
https://scholar.google.com/scholar?q=Cache-Oblivious+Streaming+B-trees
8. Hekaton: SQL Server's Memory-Optimized OLTP Engine — Cristian Diaconu, Craig Freedman, Erik Ismert, Per-Åke Larson, Pravin Mittal, Ryan Stonecipher, Nitin Verma, Mike Zwilling, 2013
https://scholar.google.com/scholar?q=Hekaton%3A+SQL+Server%27s+Memory-Optimized+OLTP+Engine
9. Cache Craftiness for Fast Multicore Key-Value Storage (Masstree) — Yandong Mao, Eddie Kohler, Robert Morris, 2012
https://scholar.google.com/scholar?q=Cache+Craftiness+for+Fast+Multicore+Key-Value+Storage+%28Masstree%29
10. The Adaptive Radix Tree: ARTful Indexing for Main-Memory Databases — Viktor Leis, Alfons Kemper, Thomas Neumann, 2013
https://scholar.google.com/scholar?q=The+Adaptive+Radix+Tree%3A+ARTful+Indexing+for+Main-Memory+Databases
11. Speedy Transactions in Multicore In-Memory Databases (Silo) — Stephen Tu, Wenting Zheng, Eddie Kohler, Barbara Liskov, Samuel Madden, 2013
https://scholar.google.com/scholar?q=Speedy+Transactions+in+Multicore+In-Memory+Databases+%28Silo%29
12. An Adaptive Packed-Memory Array — Michael A. Bender, Haodong Hu, 2007
https://scholar.google.com/scholar?q=An+Adaptive+Packed-Memory+Array
13. SageDB: A Learned Database System — Tim Kraska, Mohammad Alizadeh, Alex Beutel, et al., 2019
https://scholar.google.com/scholar?q=SageDB%3A+A+Learned+Database+System
Interactive Visualization: ALEX: Building an Updatable Learned Index for Databases

This episode breaks down DeepSeek-V4.1-Flash, a 552-billion-parameter Mixture-of-Experts model whose KV cache footprint has shrunk roughly 437-fold since DeepSeek's original V1, including a 4x drop in this single generational jump from V4-Flash. The discussion explains why long-context, tool-using agents make the KV cache the real serving bottleneck, and how the model's architecture attacks it from multiple angles: a Causal Encoder-Decoder design that nearly halves prefill compute by having upper layers project keys and values from a midpoint boundary rather than computing their own, Sliding Window Attention for bounded recent-context memory, and CSA2 (Compressed Sparse Attention 2), which layers entry-size compression, sequence compression, and cross-layer sharing on top of FP4 quantization to reach roughly 890 bytes of cache per token. It also flags the gap between DeepSeek's architectural and precision claims and actual measured deployment gains, previewing a closer look at what these compression tricks mean once put into real serving conditions. Listeners interested in efficient LLM inference, MoE architectures, or the mechanics of attention and caching will find concrete, mechanism-level explanations rather than surface-level hype.

Sources:
1. DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression — DeepSeek-AI, :, Anyi Xu, B. Li, Bangcai Lin, Bing Xue, BingCheng Xian, Bingzheng Xu, Bochao Wu, Bowei Zhang, Boyi Deng, C. C. Yu, Chao Jin, Chaofan Lin, Chen Dong, Chenbing Wang, Chenfan Feng, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chengyuan Zhang, Chenhao Xu, Chenqi Zhao, Chenze Shao, Chuhao Wang, Chuqi Zhang, Damai Dai, Dejian Yang, Deli Chen, Di Huang, Di Wu, Donghao Li, Erhang Li, Eric Fu, F. Zhou, Fangwei Zhou, Fangyun Lin, Fangzhou Yuan, Feiyu Xia, Fucong Dai, Guangbo Hao, Guanglin Li, Guanting Chen, Guoai Cao, Guofan Fan, Guolai Meng, Guowei Li, Haichuan Zhang, Haiyang Ma, Haiyang Shen, Han Li, Han Yu, Han Zhang, Hangyuan Deng, Hanwei Xu, Hanxiang Xu, Hanxun Zhong, Hao Guo, Hao Jiang, Hao Li, Hao Qin, Haodong Wen, Haofen Liang, Haofeng Huang, Haohua Liu, Haoling Zhang, Haoming Luo, Haoran Yang, Haotian Xu, Haotian Yuan, Haoting Huang, Haowen Luo, Haoyang Cai, Haoyu Chen, Haozhe Ji, Hengran Zhang, Hengrui Wang, Hengxu Wu, Honghui Ding, Hongxuan Tang, Huadong Wang, Huanqi Cao, Huazuo Gao, Hui Qu, Hui Zeng, J. Yang, J. H. Jin, J. H. Zhang, J. X. Zou, Jia Yu, Jiahui Zhou, Jiajun Chen, Jialiang Huang, Jialin Zhao, Jiamin Tang, Jian Zhou, Jianan Tong, Jianwen Li, Jiaqi Zhu, Jiarui Wang, Jiasheng Ye, Jiashi Li, Jiaxin Xu, Jiaying Ding, Jibai Lu, Jiewen Hu, Jin Yan, Jincheng Zhai, Jingchang Chen, Jingcheng Hu, Jingli Zhou, Jingsheng Xu, Jingting Xiang, Jingyan Yun, Jingyang Yuan, Jingyuan Cheng, Jinhua Zhu, Jinpeng Wang, Jinyi Chen, Jinyi Hu, Jiping Yu, Jueliang Guo, Junbo Pei, Junbo Sun, Junguang Jiang, Junjie Qiu, Junkang Zhou, Junqi Liu, Junren Li, Junxian Li, Junxiao Song, Junyi Guo, Kai Dong, Kaifeng Chen, Kaige Gao, Kang Guan, Kangdong Yuan, Ke Hong, Ke Xu, Kefan Zhao, Kexin Ji, Kexin Zhang, Kexing Zhou, Kuai Yu, Lan Zhang, Lean Wang, Lecong Zhang, Lei Wang, Letian Gao, Liang Zhao, Liansheng Xu, Lihua Guo, Lingxiao Luo, Lingyue Fu, Litao Deng, Litong Wang, Liyue Zhang, Longhao Chen, Lu Chen, Luotian Huang, Luyao Ma, Luyao Wang, M. S. Di, Max Mei, Menghao Ye, Miao Cui, Mingchuan Zhang, Minghua Zhang, Minghui Tang, Mingjing Zhang, Mingqi Wei, Mingshu Chen, Mingxing Liu, Mingxu Zhou, Mingyu Xu, Mingyu Yang, Mingze Wang, Muyang Chen, Ni Shentu, Ning Wang, Niufang Ning, Panpan Huang, Peixin Cong, Peiyi Wang, Peiyuan Xin, Pengfei Ren, Pengfei Yan, Pengle Zhang, Qi Kang, Qi Tang, Qiancheng Wang, Qiang Li, Qihao Zhu, Qingyang Li, Qinyu Chen, Qiushi Du, Qizhou Guo, Rongxian Xu, Rui Ding, Rui Hu, Rui Tian, Rui Yu, Ruidong Zhu, Ruifan Xu, Ruihan Yang, Ruihang Xia, Ruijie Lu, Ruilin Geng, Ruipeng Hong, Ruiqi Ge, Ruisong Zhang, Ruize Sun, Ruizhe Pan, Runji Wang, Runqian Chen, Runxin Xu, Ruohong Tian, Ruomeng Shen, Ruoyu Zhang, Ryan X., S. H. Liu, Shanghao Lu, Shangyan Zhou, Shanhuang Chen, Shaofei Cai, Shaoheng Nie, Shaoyuan Chen, Shengding Hu, Shengkai Lin, Shengwen Ran, Shengyu Liu, Shengyuan Jia, Shi Bai, Shi Feng, Shicheng Xu, Shichun Liu, Shiqiang Hu, Shirong Ma, Shiyu Wang, Shiyuan Feng, Shufan Gong, Shuhan Lin, Shuiping Yu, Shunfeng Zhou, Shuo Yang, Shuomeng Wang, Shuting Guo, Shuting Pan, Shuying Yu, Sinuo Cao, Siyi Lin, Sizhe Chen, Songyang Chen, Songyang Zhou, Tao Ni, Tao Yun, Tian Jin, Tian Pei, Tian Ye, Tianle Lin, Tianran Ji, Tianyi Cui, Tianyuan Yue, Tingting Yu, Tongrui Xiong, Wangding Zeng, Wei Liu, Wei Zhang, Weibin Xu, Weihao Zeng, Weilin Zhao, Wen Liu, Wenfeng Liang, Wenjie Pang, Wenjing Luo, Wenjing Yao, Wenjun Gao, Wenkai Shao, Wenkai Yang, Wenli Zhang, Wenlu Wang, Wenlve Huang, Wenqian Yan, Wentao Zhang, Xi Gao, Xiang He, Xiang Li, Xiangli Li, Xiangwen Wang, Xiangying Zhang, Xiankui Wei, Xiao Bi, Xiaodong Liu, Xiaohan Wang, Xiaojian Qu, Xiaokang Chen, Xiaokang Zhang, Xiaotao Nie, Xiaoyao Zou, Xiaoyuan Li, Xicheng Guo, Xieting Chu, Xin Cheng, Xin Liu, Xin Xie, Xinbo Xu, Xingchao Liu, Xingchen Liu, Xingkai Yu, Xingyou Li, Xintong Yao, Xinyang Chen, Xinyong Jiang, Xinyu Yang, Xinyu Yang, Xu Chen, Xuanyu Wang, Xubei Zhong, Xuecheng Su, Xuejie Liu, Xuheng Lin, Xujie Fan, Xuncheng Zhao, Xuwei Fu, Y. C. Yan, Y. H. Jiang, Y. T. Wu, Y. W. M., Y. Z. Wang, Yafei Gao, Yang Yang, Yang Zhang, Yanru Ma, Yanwen Huang, Yao Li, Yao Li, Yao Meng, Yao Zhao, Yaofeng Sun, Yaohui Wang, Yaoyang Ye, Yehang Yin, Yexinrui Wu, Yi Qian, Yi Tao, Yi Yu, Yichao Zhang, Yichen Jiang, Yicheng Wang, Yifan Ding, Yifan Shi, Yifeng Peng, Yifeng Zhai, Yijia Wu, Yiliang Xiong, Yilun Wang, Ying He, Ying Zhou, Yingjia Luo, Yinmin Zhong, Yiping Wang, Yisong Wang, Yixiang Zhang, Yixiao Chen, Yixuan Tan, Yixuan Wei, Yiyang Ma, Yiyao Yang, Yiyuan Liu, Yizai Cai, Yizhen Wei, Yizhi Wang, Yonglun Yang, Yongqi Zhuo, Yongqiang Guo, Yongtong Wu, Yu Wu, Yu Zhang, Yuan Bian, Yuan Cheng, Yuan Ou, Yuan Sun, Yuanfan Xu, Yuanhang Sun, Yuanhao Li, Yuchen Liu, Yuchen Yao, Yudong Han, Yuduan Wang, Yuhan Wu, Yuhao Meng, Yuheng Zou, YuKun Li, Yunchuan Wang, Yunfan Xiao, Yunfan Xiong, Yupeng Chen, Yuqian Cao, Yuqian Wang, Yuqing Chen, Yushun Zhang, Yutong Lin, Yuwei Xiao, Yuxian Gu, Yuxiang Chen, Yuxiang Huang, Yuxiang Luo, Yuxiang You, Yuxin Chen, Yuxin Xiang, Yuxuan Liu, Yuxuan Zhou, Yuyang Zhou, Yuzhe Guo, Yuzhen Huang, Yuzhuo Bai, Z. Y. Z., Zanlin Ni, Zehao Wang, Zehua Zhao, Zehui Ren, Zejun Zhao, Zhangli Sha, Zhanying Wang, Zhaochen Zhang, Zhaoshuai Du, Zhe Fu, Zhean Xu, Zhenda Xie, Zheng Liu, Zhengyan Zhang, Zhenhua Dong, Zhewen Hao, Zhibang Wang, Zhibin Gou, Zhicheng Ma, Zhihao Li, Zhihong Shao, Zhihuan Huang, Zhijie Li, Zhirui Lu, Zhixian Huang, Zhixuan Chen, Zhixuan Chen, Zhixuan Pan, Zhiyu Wu, Zhizhou Ren, Zhu He, Zhuoshu Li, Zhuping Zhang, Zian Xu, Zihao Wang, Zihui Gu, Zijia Zhu, Zili Zhang, Zilin Li, Zilong Hou, Zilong Lyu, Ziqiao Wang, Ziwei Xie, Ziya Zhang, Ziyi Gao, Zizheng Pan, Zonglin Li, Zongqing Yao, Zui Chen, Zuofan Wu, Chenchen Ling, Chengyu Hou, Chong Chen, D. Li, Di Qi, Dongjie Ji, Fang Wei, Fanyi Xia, Fei Xie, Feiyi Tan, Hailong Guo, Haiyan Zhai, Hui Zhou, Huihui Tan, Huijie Li, Jia Luo, Jia Song, Jialu Cai, Jian Liang, Jiangting Zhou, Jiaqi Gao, Jiayi Shao, Jie Chen, Jieyu Yang, Jin Chen, Jingde Zhang, Jingzi Zhou, Jinqian Wang, Jinyang Liu, JinZhao Sun, Junhua Ling, Junmin Zheng, Kaicheng Yang, Ke Xu, Le Su, Leyi Xia, Liangfeng Ding, Lin Zhuo, Linwang Ma, Linyan Zhu, Liyu Cai, Luqi Yao, M. K. Zhang, Meng Li, Miao Lin, Miaojun Wang, Min Zhang, Mingming Li, Mingming Wang, Mingze Yin, Minmin Han, Nan Cao, Ning Wang, Ningxin Ma, Panpan Wang, Peihan Lin, Peng Sun, Peng Zhang, Qian Ying, Qiang Xiang, Qiao Wang, Qingmiao Mao, Qiwei Jiang, Rongli Jin, Ruyi Chen, Sha Tao, Shangmian Sun, Shaoqing Wu, Shichao Zou, Si Lei, Tianyang Zhang, Tianyu Sun, Tingting Yin, W. L. Xiao, Wei An, Wei Li, Wei Wang, Weiwei Lin, Wenqing Hou, X. Lin, Xiangfei Meng, Xianzhu Huang, Xiao Peng, Xiaoqian Li, Xiaoting Zhang, Xiaowen Sun, Xiaoxiang Wang, Xiaoyu Ye, Xinrou Zhang, Xinyu Zhang, Xue Cao, Xueyin Chen, Yanan Zhou, Yanhong Xu, Yao Xia, Yao Xu, Yi Shao, Yihong Zhang, Yiling Ma, Ying Tang, Yining Lou, Yiru Chen, Yishi Piao, Yixuan Chen, Yong Xiong, Yuchen Xuan, Yuehan Yang, Yuer Xu, Yukun Zha, Yunxian Ma, Yuping Lin, Yuting Yan, Yutong Xie, Yuwen Sheng, Yuxuan Zhu, Zekai Zhang, Zhe Ju, Zhenzhen Lin, Zheren Gao, Zheyang Sun, Zhigang Yan, Zhongyu Wu, Zi Wang, Zihua Qu, Ziling Yan, Ziyi Wan, 2026
http://arxiv.org/abs/2609.19969
2. Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention — Jingyang Yuan, Huazuo Gao, Damai Dai, et al. (DeepSeek-AI), 2025
https://scholar.google.com/scholar?q=Native+Sparse+Attention%3A+Hardware-Aligned+and+Natively+Trainable+Sparse+Attention
3. MoBA: Mixture of Block Attention for Long-Context LLMs — Enzhe Lu et al. (Moonshot AI), 2025
https://scholar.google.com/scholar?q=MoBA%3A+Mixture+of+Block+Attention+for+Long-Context+LLMs
4. Longformer: The Long-Document Transformer — Iz Beltagy, Matthew E. Peters, Arman Cohan, 2020
https://scholar.google.com/scholar?q=Longformer%3A+The+Long-Document+Transformer
5. H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models — Zhenyu Zhang, Ying Sheng, Tianyi Zhou, et al., 2023
https://scholar.google.com/scholar?q=H2O%3A+Heavy-Hitter+Oracle+for+Efficient+Generative+Inference+of+Large+Language+Models
6. Attention Is All You Need — Ashish Vaswani, Noam Shazeer, Niki Parmar, et al., 2017
https://scholar.google.com/scholar?q=Attention+Is+All+You+Need
7. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer (T5) — Colin Raffel, Noam Shazeer, Adam Roberts, et al. (Google), 2020
https://scholar.google.com/scholar?q=Exploring+the+Limits+of+Transfer+Learning+with+a+Unified+Text-to-Text+Transformer+%28T5%29
8. You Only Cache Once: Decoder-Decoder Architectures for Language Models (YOCO) — Yutao Sun, Li Dong, Yi Zhu, et al. (Microsoft Research), 2024
https://scholar.google.com/scholar?q=You+Only+Cache+Once%3A+Decoder-Decoder+Architectures+for+Language+Models+%28YOCO%29
9. DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving — Yinmin Zhong, Shengyu Liu, Junda Chen, et al., 2024
https://scholar.google.com/scholar?q=DistServe%3A+Disaggregating+Prefill+and+Decoding+for+Goodput-optimized+Large+Language+Model+Serving
10. You Only Cache Once: Decoder-Decoder Architectures for Language Models — Y. Sun, L. Dong, Y. Zhu, S. Huang, W. Wang, S. Ma, Q. Zhang, J. Wang, F. Wei, 2024
https://scholar.google.com/scholar?q=You+Only+Cache+Once%3A+Decoder-Decoder+Architectures+for+Language+Models
11. IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse — Y. Bai, Q. Dong, T. Jiang, X. Lv, Z. Du, A. Zeng, J. Tang, J. Li, 2026
https://scholar.google.com/scholar?q=IndexCache%3A+Accelerating+Sparse+Attention+via+Cross-Layer+Index+Reuse
12. You Only Index Once: Cross-layer sparse attention with shared routing — Y. Sun, Y. Zhang, L. Dong, J. Wang, F. Wei, 2026
https://scholar.google.com/scholar?q=You+Only+Index+Once%3A+Cross-layer+sparse+attention+with+shared+routing
13. HySparse: A Hybrid Sparse Attention Architecture with Oracle Token Selection and KV Cache Sharing — Y. Gao, J. Wei, Q. Zhang, Y. Cheng, S. Chen, et al., 2026
https://scholar.google.com/scholar?q=HySparse%3A+A+Hybrid+Sparse+Attention+Architecture+with+Oracle+Token+Selection+and+KV+Cache+Sharing
14. Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models (Engram) — X. Cheng, W. Zeng, D. Dai, Q. Chen, B. Wang, et al., 2026
https://scholar.google.com/scholar?q=Conditional+Memory+via+Scalable+Lookup%3A+A+New+Axis+of+Sparsity+for+Large+Language+Models+%28Engram%29
15. Muon is Scalable for LLM Training — J. Liu, J. Su, X. Yao, Z. Jiang, et al., 2025
https://scholar.google.com/scholar?q=Muon+is+Scalable+for+LLM+Training
Interactive Visualization: DeepSeek-V4.1-Flash: Shrinking KV Cache at 552B Scale

This episode examines RewardingDoubt, a reinforcement-learning method for training large language models to express calibrated confidence rather than defaulting to near-maximal certainty regardless of correctness. The discussion covers why miscalibrated confidence is dangerous in high-stakes domains like medicine, legal consultation, and customer service, where deferring to a human reviewer only works if the model's stated confidence is trustworthy. It walks through the technical core of the approach: a proper scoring rule (specifically the logarithmic scoring rule) that mathematically rewards models for reporting their true beliefs rather than gaming their confidence scores, framed intuitively as a betting game where lying about certainty costs real "money." The conversation contrasts this RL-based approach with prior black-box (output-only) and white-box (internals-probing) calibration methods, including Kadavath et al.'s self-evaluation technique and Lin, Hilton, and Evans' supervised fine-tuning approach, arguing RL sidesteps the quality ceiling imposed by fixed ground-truth labels. It also details the paper's MDP formulation, where the model's answer is frozen before a separate PPO-trained pass generates the confidence sequence, and explains the two evaluation metrics—Expected Calibration Error and AUROC—used to judge whether the method actually works.

Sources:
1. Rewarding Doubt: A Reinforcement Learning Approach to Calibrated Confidence Expression of Large Language Models — David Bani-Harouni, Chantal Pellegrini, Paul Stangel, Ege Özsoy, Kamilia Zaripova, Nassir Navab, Matthias Keicher, 2025
http://arxiv.org/abs/2503.02623v6
2. Language Models (Mostly) Know What They Know — Kadavath, Conerly, Askell, Henighan, Drain, Perez, Schiefer, Hatfield-Dodds, DasSarma, Tran-Johnson, et al. (Anthropic), 2022
https://scholar.google.com/scholar?q=Language+Models+%28Mostly%29+Know+What+They+Know
3. Teaching Models to Express Their Uncertainty in Words — Stephanie Lin, Jacob Hilton, Owain Evans, 2022
https://scholar.google.com/scholar?q=Teaching+Models+to+Express+Their+Uncertainty+in+Words
4. LACIE: Listener-Aware Finetuning for Confidence Calibration in Large Language Models — Elias Stengel-Eskin, Peter Hase, Mohit Bansal, 2024
https://scholar.google.com/scholar?q=LACIE%3A+Listener-Aware+Finetuning+for+Confidence+Calibration+in+Large+Language+Models
5. Taming Overconfidence in LLMs: Reward Calibration in RLHF — Jixuan Leng, Chengsong Huang, Banghua Zhu, Jiaxin Huang, 2024
https://scholar.google.com/scholar?q=Taming+Overconfidence+in+LLMs%3A+Reward+Calibration+in+RLHF
6. Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs — Miao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li, Jie Fu, Junxian He, Bryan Hooi, 2024
https://scholar.google.com/scholar?q=Can+LLMs+Express+Their+Uncertainty%3F+An+Empirical+Evaluation+of+Confidence+Elicitation+in+LLMs
7. Calibration-Tuning: Teaching Large Language Models to Know What They Don't Know — Sanyam Kapoor, Nate Gruver, Manley Roberts, Arka Pal, Samuel Dooley, Micah Goldblum, Andrew Wilson, 2024
https://scholar.google.com/scholar?q=Calibration-Tuning%3A+Teaching+Large+Language+Models+to+Know+What+They+Don%27t+Know
Interactive Visualization: Teaching Language Models to Bet Honestly on Their Own Answers

This episode explores Simula, a system from EPFL and Google DeepMind researchers that generates entire specialized training datasets from scratch, with no seed examples required. It breaks down the reasoning-driven, multi-role agentic pipeline — where separate model roles propose, critique, sample, generate, and filter data — contrasting it with older approaches like Self-Instruct's hand-written templates and Promptbreeder's opaque evolutionary search. The discussion covers the QDC framework (quality, diversity, complexity) used to define "good" synthetic data, and how Simula builds an explicit taxonomy tree that serves as both the sampling mechanism and an auditable record of why each data point exists, echoing the transparency goals of Datasheets for Datasets. The hosts also flag a key tension worth scrutinizing: the system relies on the model to grade its own taxonomy quality, critique its own outputs, and judge complexity — assumptions that deserve real pushback rather than blind trust. Listeners interested in synthetic data generation, dataset auditability, or the mechanics of multi-stage agentic pipelines will find the walkthrough of Simula's three-stage architecture a concrete look at how reasoning-first systems aim to replace scarce human annotation.

Sources:
1. Reasoning-Driven Synthetic Data Generation and Evaluation — Tim R. Davidson, Benoit Seguin, Enrico Bacis, Cesar Ilharco, Hamza Harkous, 2026
http://arxiv.org/abs/2603.29791v1
2. Surveying the effects of quality, diversity, and complexity in synthetic data from large language models — Havrilla, Dai, O'Mahony, Oostermeijer, Zisler, Albalak, Milo, Raparthy, Gandhi, Abbasi, et al., 2024
https://scholar.google.com/scholar?q=Surveying+the+effects+of+quality%2C+diversity%2C+and+complexity+in+synthetic+data+from+large+language+models
3. LLM evaluators recognize and favor their own generations — Panickssery, Bowman, Feng, 2024
https://scholar.google.com/scholar?q=LLM+evaluators+recognize+and+favor+their+own+generations
4. Position: will we run out of data? Limits of LLM scaling based on human-generated data — Villalobos, Ho, Sevilla, Besiroglu, Heim, Hobbhahn, 2024
https://scholar.google.com/scholar?q=Position%3A+will+we+run+out+of+data%3F+Limits+of+LLM+scaling+based+on+human-generated+data
5. Scaling data-constrained language models — Muennighoff, Rush, Barak, Le Scao, Tazi, Piktus, Pyysalo, Wolf, Raffel, 2023
https://scholar.google.com/scholar?q=Scaling+data-constrained+language+models
6. Lora without regret — Schulman, Thinking Machines Lab, 2025
https://scholar.google.com/scholar?q=Lora+without+regret
7. Self-recognition in language models — Davidson, Surkov, Veselovsky, Russo, West, Gulcehre, 2024
https://scholar.google.com/scholar?q=Self-recognition+in+language+models
Interactive Visualization: Zero-Seed Data Generation: Simula's Reasoning-Driven Pipeline

This episode examines "Skip a Layer or Loop It? Learning Program-of-Layers in LLMs" (ICML 2026), which asks whether a frozen pretrained LLM can answer better if each input gets its own sequence of skipped and repeated layers, with no retraining of the base model. The hosts place it against prior work: layer-dropping and early-exit methods built to save compute, recurrent-depth models trained to loop, and studies showing that middle layers in frozen models tolerate skipping, repeating and reordering. They then explain how programs are searched with Monte Carlo Tree Search. That search uses skip or repeat actions on blocks of up to four layers, a binary correct-answer reward, and a penalty on program length. A learned predictor, POLAR, is meant to replace that search at inference. The hosts stress that a "valid program exists" result is judged against the ground-truth label, so it is an oracle coverage number and not an accuracy you can get at inference. They test that reading on four models and DART-Math difficulty levels, asking whether the gains are real or just search luck.

Sources:
1. Skip a Layer or Loop It? Learning Program-of-Layers in LLMs — Ziyue Li, Yang Li, Tianyi Zhou, 2026
http://arxiv.org/abs/2606.06574
2. Transformer Layers as Painters — Qi Sun, Marc Pickett, Aakash Kumar Nain, Llion Jones, 2024
https://scholar.google.com/scholar?q=Transformer+Layers+as+Painters
3. The Remarkable Robustness of LLMs: Stages of Inference? — Vedang Lad, Wes Gurnee, Max Tegmark, 2024
https://scholar.google.com/scholar?q=The+Remarkable+Robustness+of+LLMs%3A+Stages+of+Inference%3F
4. Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach — Jonas Geiping, Sean McLeish, Neel Jain, John Kirchenbauer, et al., 2025
https://scholar.google.com/scholar?q=Scaling+up+Test-Time+Compute+with+Latent+Reasoning%3A+A+Recurrent+Depth+Approach
5. Teaching Pretrained Language Models to Think Deeper with Retrofitted Recurrence — Sean McLeish et al., 2025
https://scholar.google.com/scholar?q=Teaching+Pretrained+Language+Models+to+Think+Deeper+with+Retrofitted+Recurrence
6. Reasoning with Latent Thoughts: On the Power of Looped Transformers — Nikunj Saunshi, Nishanth Dikkala, Zhiyuan Li, Sanjiv Kumar, Sashank Reddi, 2025
https://scholar.google.com/scholar?q=Reasoning+with+Latent+Thoughts%3A+On+the+Power+of+Looped+Transformers
7. Scaling Latent Reasoning via Looped Language Models (Ouro) — Rui-Jie Zhu et al., 2025
https://scholar.google.com/scholar?q=Scaling+Latent+Reasoning+via+Looped+Language+Models+%28Ouro%29
8. Do Language Models Use Their Depth Efficiently? — Róbert Csordás, Christopher D. Manning, Christopher Potts, 2025
https://scholar.google.com/scholar?q=Do+Language+Models+Use+Their+Depth+Efficiently%3F
9. The Unreasonable Ineffectiveness of the Deeper Layers — Andrey Gromov, Kushal Tirumala, Hassan Shapourian, Paolo Glorioso, Daniel A. Roberts, 2024
https://scholar.google.com/scholar?q=The+Unreasonable+Ineffectiveness+of+the+Deeper+Layers
10. Large Language Monkeys: Scaling Inference Compute with Repeated Sampling — Bradley Brown, Jordan Juravsky, Ryan Ehrlich, et al., 2024
https://scholar.google.com/scholar?q=Large+Language+Monkeys%3A+Scaling+Inference+Compute+with+Repeated+Sampling
11. Not All Layers of LLMs Are Necessary During Inference — Siqi Fan, Xin Jiang, Xuying Meng, et al., 2024
https://scholar.google.com/scholar?q=Not+All+Layers+of+LLMs+Are+Necessary+During+Inference
12. Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding — Jun Zhang, Jue Wang, Huan Li, et al., 2023
https://scholar.google.com/scholar?q=Draft+%26+Verify%3A+Lossless+Large+Language+Model+Acceleration+via+Self-Speculative+Decoding
13. Adaptive Computation Time for Recurrent Neural Networks — Alex Graves, 2016
https://scholar.google.com/scholar?q=Adaptive+Computation+Time+for+Recurrent+Neural+Networks
14. Training Large Language Models to Reason in a Continuous Latent Space (Coconut) — Shibo Hao, Sainbayar Sukhbaatar, DiJia Su, et al., 2024
https://scholar.google.com/scholar?q=Training+Large+Language+Models+to+Reason+in+a+Continuous+Latent+Space+%28Coconut%29
Interactive Visualization: Skip or Loop Layers: Learning Programs-of-Layers in Frozen LLMs

This episode examines "Skip a Layer or Loop it? Test-Time Depth Adaptation of Pretrained LLMs," which proposes Chain-of-Layers (CoLa). CoLa takes a frozen pretrained model and, for each input, skips some layers, repeats others, and reorders them, with no finetuning. The hosts explain why skipping works, citing residual connections, ResNets behaving like ensembles of shallow paths, and static pruning results like ShortGPT and "The Unreasonable Ineffectiveness of the Deeper Layers". They also cover dynamic early-exit methods and looped-depth models such as Universal Transformers. They disagree about whether pruning results mean deeper layers are dead weight, since those results come mostly from multiple-choice tasks and multi-step reasoning degrades faster. The paper's claim is that many correctly answered samples still work with a shorter layer chain, and many wrong ones can be fixed by some other chain. The episode walks through the Monte Carlo Tree Search that finds these paths, including its skip and repeat edits, its scoring rule with an exploration bonus and a length penalty, and the resulting Pareto set of short, accurate paths. It also notes that the paper omits closely related prior work on frozen-model layer manipulation, which affects how novel the result is.

Sources:
1. Skip a Layer or Loop it? Test-Time Depth Adaptation of Pretrained LLMs — Ziyue Li, Yang Li, Tianyi Zhou, 2025
http://arxiv.org/abs/2507.07996
2. Universal Transformers — Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, Łukasz Kaiser, 2018 (ICLR 2019)
https://scholar.google.com/scholar?q=Universal+Transformers
3. Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach — Jonas Geiping, Sean McLeish, Neel Jain, John Kirchenbauer, Siddharth Singh, Brian R. Bartoldson, Bhavya Kailkhura, Abhinav Bhatele, Tom Goldstein, 2025
https://scholar.google.com/scholar?q=Scaling+up+Test-Time+Compute+with+Latent+Reasoning%3A+A+Recurrent+Depth+Approach
4. Mixture-of-Depths: Dynamically Allocating Compute in Transformer-Based Language Models — David Raposo, Sam Ritter, Blake Richards, Timothy Lillicrap, Peter Conway Humphreys, Adam Santoro, 2024
https://scholar.google.com/scholar?q=Mixture-of-Depths%3A+Dynamically+Allocating+Compute+in+Transformer-Based+Language+Models
5. Transformer Layers as Painters — Qi Sun, Marc Pickett, Aakash Kumar Nain, Llion Jones, 2024
https://scholar.google.com/scholar?q=Transformer+Layers+as+Painters
6. The Unreasonable Ineffectiveness of the Deeper Layers — Andrey Gromov, Kushal Tirumala, Hassan Shapourian, Paolo Glorioso, Daniel A. Roberts, 2024
https://scholar.google.com/scholar?q=The+Unreasonable+Ineffectiveness+of+the+Deeper+Layers
7. ShortGPT: Layers in Large Language Models are More Redundant Than You Expect — Xin Men, Mingyu Xu, Qingyu Zhang, Bingning Wang, Hongyu Lin, Yaojie Lu, Xianpei Han, Weipeng Chen, 2024
https://scholar.google.com/scholar?q=ShortGPT%3A+Layers+in+Large+Language+Models+are+More+Redundant+Than+You+Expect
8. LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding — Mostafa Elhoushi, Akshat Shrivastava, Diana Liskovich, Basil Hosmer, Bram Wasti, Liangzhen Lai, Anas Mahmoud, Bilge Acun, Saurabh Agarwal, Ahmed Roman, Ahmed A. Aly, Beidi Chen, Carole-Jean Wu, 2024 (ACL 2024)
https://scholar.google.com/scholar?q=LayerSkip%3A+Enabling+Early+Exit+Inference+and+Self-Speculative+Decoding
9. Confident Adaptive Language Modeling — Tal Schuster, Adam Fisch, Jai Gupta, Mostafa Dehghani, Dara Bahri, Vinh Q. Tran, Yi Tay, Donald Metzler, 2022 (NeurIPS 2022)
https://scholar.google.com/scholar?q=Confident+Adaptive+Language+Modeling
10. Dr.LLM: Dynamic Layer Routing in LLMs — Ahmed Heakl et al., 2025
https://scholar.google.com/scholar?q=Dr.LLM%3A+Dynamic+Layer+Routing+in+LLMs
11. Confident Adaptive Language Modeling (CALM) — Tal Schuster et al., 2022
https://scholar.google.com/scholar?q=Confident+Adaptive+Language+Modeling+%28CALM%29
12. Router-Tuning / FlexiDepth: dynamic layer skipping in pretrained LLMs — Shwai He et al. (Router-Tuning); Xuan Luo et al. (FlexiDepth), 2024-2025
https://scholar.google.com/scholar?q=Router-Tuning+%2F+FlexiDepth%3A+dynamic+layer+skipping+in+pretrained+LLMs
13. Mixture-of-Recursions: Learning Dynamic Recursive Depths for Adaptive Token-Level Computation — Sangmin Bae et al., 2025
https://scholar.google.com/scholar?q=Mixture-of-Recursions%3A+Learning+Dynamic+Recursive+Depths+for+Adaptive+Token-Level+Computation
14. Do Language Models Use Their Depth Efficiently? — Róbert Csordás, Christopher D. Manning, Christopher Potts, 2025
https://scholar.google.com/scholar?q=Do+Language+Models+Use+Their+Depth+Efficiently%3F
15. Large Language Monkeys: Scaling Inference Compute with Repeated Sampling — Bradley Brown et al., 2024
https://scholar.google.com/scholar?q=Large+Language+Monkeys%3A+Scaling+Inference+Compute+with+Repeated+Sampling
16. Residual Networks Behave Like Ensembles of Relatively Shallow Networks — Andreas Veit, Michael Wilber, Serge Belongie, 2016
https://scholar.google.com/scholar?q=Residual+Networks+Behave+Like+Ensembles+of+Relatively+Shallow+Networks
17. The Remarkable Robustness of LLMs: Stages of Inference? — Vedang Lad, Wes Gurnee, Max Tegmark, 2024
https://scholar.google.com/scholar?q=The+Remarkable+Robustness+of+LLMs%3A+Stages+of+Inference%3F
18. Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding — Jun Zhang et al., 2023
https://scholar.google.com/scholar?q=Draft+%26+Verify%3A+Lossless+Large+Language+Model+Acceleration+via+Self-Speculative+Decoding
19. Interpreting GPT: the logit lens / Eliciting Latent Predictions from Transformers with the Tuned Lens — nostalgebraist (2020); Nora Belrose et al. (2023), 2020/2023
https://scholar.google.com/scholar?q=Interpreting+GPT%3A+the+logit+lens+%2F+Eliciting+Latent+Predictions+from+Transformers+with+the+Tuned+Lens
Interactive Visualization: Chain-of-Layers: Skipping and Looping Frozen LLM Layers Per Input

This episode explores Structural Language Models of Code, a 2020 paper from Technion, Tel Aviv University and Facebook AI Research that tackles "any-code completion": predicting a missing piece of a program with no restriction on vocabulary or structure. It traces the history from 1969-era automatic programming through domain-specific systems like FlashFill and DeepCoder, and prior general-language work that limited APIs, types or syntax. Then it contrasts a subtoken sequence-to-sequence baseline with the paper's approach of modeling code as an abstract syntax tree. That approach applies the language-model chain rule over a depth-first tree traversal, using partial AST paths that end at the node being generated, which extends the authors' earlier code2vec and code2seq work from reading code to writing it. The discussion covers the headline exact-match gains, Java accuracy@1 of 18.04 versus 16.93 and C# 37.61 versus 26.42. It also raises the question of whether a one-point win is meaningful, given that exact match undercounts logically equivalent code. Listeners interested in how syntax-aware models can generate code with unseen identifiers will find the mechanics and the skeptical read of the metrics useful.

Sources:
1. Structural Language Models of Code: Any-Code Completion with Trees
https://arxiv.org/pdf/1910.00577
2. Structured Generative Models of Natural Source Code — Chris J. Maddison, Daniel Tarlow, 2014
https://scholar.google.com/scholar?q=Structured+Generative+Models+of+Natural+Source+Code
3. On the Naturalness of Software — Abram Hindle, Earl T. Barr, Zhendong Su, Mark Gabel, Premkumar Devanbu, 2012
https://scholar.google.com/scholar?q=On+the+Naturalness+of+Software
4. A Syntactic Neural Model for General-Purpose Code Generation — Pengcheng Yin, Graham Neubig, 2017
https://scholar.google.com/scholar?q=A+Syntactic+Neural+Model+for+General-Purpose+Code+Generation
5. Generative Code Modeling with Graphs — Marc Brockschmidt, Miltiadis Allamanis, Alexander L. Gaunt, Oleksandr Polozov, 2019
https://scholar.google.com/scholar?q=Generative+Code+Modeling+with+Graphs
6. Code Completion with Statistical Language Models — Veselin Raychev, Martin Vechev, Eran Yahav, 2014
https://scholar.google.com/scholar?q=Code+Completion+with+Statistical+Language+Models
7. IntelliCode Compose: Code Generation Using Transformer — Alexey Svyatkovskiy, Shao Kun Deng, Shengyu Fu, Neel Sundaresan, 2020
https://scholar.google.com/scholar?q=IntelliCode+Compose%3A+Code+Generation+Using+Transformer
8. Efficient Training of Language Models to Fill in the Middle — Mohammad Bavarian, Heewoo Jun, Nikolas Tezak, John Schulman, Christine McLeavey, Jerry Tworek, Mark Chen, 2022
https://scholar.google.com/scholar?q=Efficient+Training+of+Language+Models+to+Fill+in+the+Middle
9. A General Path-Based Representation for Predicting Program Properties — Uri Alon, Meital Zilberstein, Omer Levy, Eran Yahav, 2018
https://scholar.google.com/scholar?q=A+General+Path-Based+Representation+for+Predicting+Program+Properties
10. code2vec: Learning Distributed Representations of Code — Uri Alon, Meital Zilberstein, Omer Levy, Eran Yahav, 2019
https://scholar.google.com/scholar?q=code2vec%3A+Learning+Distributed+Representations+of+Code
11. code2seq: Generating Sequences from Structured Representations of Code — Uri Alon, Shaked Brody, Omer Levy, Eran Yahav, 2019
https://scholar.google.com/scholar?q=code2seq%3A+Generating+Sequences+from+Structured+Representations+of+Code
12. Learning to Represent Programs with Graphs — Miltiadis Allamanis, Marc Brockschmidt, Mahmoud Khademi, 2018
https://scholar.google.com/scholar?q=Learning+to+Represent+Programs+with+Graphs
13. Abstract Syntax Networks for Code Generation and Semantic Parsing — Maxim Rabinovich, Mitchell Stern, Dan Klein, 2017
https://scholar.google.com/scholar?q=Abstract+Syntax+Networks+for+Code+Generation+and+Semantic+Parsing
14. PHOG: Probabilistic Model for Code — Pavol Bielik, Veselin Raychev, Martin Vechev, 2016
https://scholar.google.com/scholar?q=PHOG%3A+Probabilistic+Model+for+Code
15. Incorporating Copying Mechanism in Sequence-to-Sequence Learning — Jiatao Gu, Zhengdong Lu, Hang Li, Victor O.K. Li, 2016
https://scholar.google.com/scholar?q=Incorporating+Copying+Mechanism+in+Sequence-to-Sequence+Learning
16. On the Bottleneck of Graph Neural Networks and its Practical Implications — Uri Alon, Eran Yahav, 2020
https://scholar.google.com/scholar?q=On+the+Bottleneck+of+Graph+Neural+Networks+and+its+Practical+Implications
17. InCoder: A Generative Model for Code Infilling and Synthesis — Daniel Fried et al., 2022
https://scholar.google.com/scholar?q=InCoder%3A+A+Generative+Model+for+Code+Infilling+and+Synthesis
18. Evaluating Large Language Models Trained on Code (Codex / HumanEval) — Mark Chen et al., 2021
https://scholar.google.com/scholar?q=Evaluating+Large+Language+Models+Trained+on+Code+%28Codex+%2F+HumanEval%29
19. RepoCoder: Repository-Level Code Completion Through Iterative Retrieval and Generation — Fengji Zhang et al., 2023
https://scholar.google.com/scholar?q=RepoCoder%3A+Repository-Level+Code+Completion+Through+Iterative+Retrieval+and+Generation
20. UniXcoder: Unified Cross-Modal Pre-training for Code Representation — Daya Guo et al., 2022
https://scholar.google.com/scholar?q=UniXcoder%3A+Unified+Cross-Modal+Pre-training+for+Code+Representation
Interactive Visualization: Structural Language Models of Code: Any-Code Completion with Trees

This episode examines "The Unreasonable Ineffectiveness of the Deeper Layers," a 2025 study showing that up to half the layers of a 70-billion-parameter model can be removed with almost no drop in standard QA benchmark performance. The discussion covers the residual stream architecture that makes such pruning possible, why later transformer layers often contribute diminishing changes to accumulated representations, and how researchers identify which layer blocks are safe to cut by comparing input and output similarity. It also unpacks prior work on knowledge localization, including causal tracing of factual associations and feed-forward layers acting as key-value memories, to explain why deep layers might be redundant rather than essential. Finally, it details how QLoRA enables lightweight healing of the pruned model's mismatched seams using minimal compute, making the entire process feasible on a single GPU rather than a training cluster. Listeners interested in model efficiency, interpretability, or the surprising redundancy inside trusted large language models will find the core result — and the mechanics behind it — genuinely counterintuitive.

Sources:
1. The Unreasonable Ineffectiveness of the Deeper Layers — Andrey Gromov, Kushal Tirumala, Hassan Shapourian, Paolo Glorioso, Daniel A. Roberts, 2024
http://arxiv.org/abs/2403.17887
2. ShortGPT: Layers in Large Language Models are More Redundant Than You Expect — Xin Men, Mingyu Xu, Qingyu Zhang, Bingning Wang, Hongyu Lin, Yaojie Lu, Xianpei Han, Weipeng Chen, 2024
https://scholar.google.com/scholar?q=ShortGPT%3A+Layers+in+Large+Language+Models+are+More+Redundant+Than+You+Expect
3. Locating and Editing Factual Associations in GPT (ROME) — Kevin Meng, David Bau, Alex Andonian, Yonatan Belinkov, 2022
https://scholar.google.com/scholar?q=Locating+and+Editing+Factual+Associations+in+GPT+%28ROME%29
4. Transformer Feed-Forward Layers Are Key-Value Memories — Mor Geva, Roei Schuster, Jonathan Berant, Omer Levy, 2021
https://scholar.google.com/scholar?q=Transformer+Feed-Forward+Layers+Are+Key-Value+Memories
5. Compact Language Models via Pruning and Knowledge Distillation (Minitron) — Saurav Muralidharan, Sharath Turuvekere Sreenivas, Raviraj Joshi, Marcin Chochowski, Mostofa Patwary, Mohammad Shoeybi, Bryan Catanzaro, Jan Kautz, Pavlo Molchanov (NVIDIA), 2024
https://scholar.google.com/scholar?q=Compact+Language+Models+via+Pruning+and+Knowledge+Distillation+%28Minitron%29
6. The Truth is in There: Improving Reasoning in Language Models with Layer-Selective Rank Reduction (LASER) — Pratyusha Sharma, Jordan T. Ash, Dipendra Misra, 2023
https://scholar.google.com/scholar?q=The+Truth+is+in+There%3A+Improving+Reasoning+in+Language+Models+with+Layer-Selective+Rank+Reduction+%28LASER%29
7. LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding — Mostafa Elhoushi, Akshat Shrivastava, Diana Liskovich, Basil Hosmer, et al., 2024
https://scholar.google.com/scholar?q=LayerSkip%3A+Enabling+Early+Exit+Inference+and+Self-Speculative+Decoding
8. Are Emergent Abilities of Large Language Models a Mirage? — Rylan Schaeffer, Brando Miranda, Sanmi Koyejo, 2023
https://scholar.google.com/scholar?q=Are+Emergent+Abilities+of+Large+Language+Models+a+Mirage%3F
9. SliceGPT: Compress Large Language Models by Deleting Rows and Columns — Saleh Ashkboos, Maximilian L. Croci, Marcelo Gennari do Nascimento, Torsten Hoefler, James Hensman, 2024
https://scholar.google.com/scholar?q=SliceGPT%3A+Compress+Large+Language+Models+by+Deleting+Rows+and+Columns
10. Tensor Programs VI: Feature Learning in Infinite-Depth Neural Networks — Greg Yang, Dingli Yu, Chen Zhu, Soufiane Hayou, 2023
https://scholar.google.com/scholar?q=Tensor+Programs+VI%3A+Feature+Learning+in+Infinite-Depth+Neural+Networks
Interactive Visualization: The Unreasonable Ineffectiveness of Deep Transformer Layers

This episode examines ShortGPT, a layer-pruning method that ranks the layers of a pre-norm LLM by Block Influence (one minus the average cosine similarity between a layer's input and output hidden states) and deletes the lowest-scoring ones with no gradients or retraining. It sets the method against unstructured pruning and the structured baselines LLM-Pruner, SliceGPT, and LaCo. It also covers the pre-norm residual-stream argument for why deep layers may be redundant, along with the limits of that argument. The central tension is the headline result: removing nine layers from a 32-layer model barely dents MMLU (45.4 to 44.0) but collapses XSum summarization (19.40 to 0.67). That gap prompts the question of whether "roughly 90% of performance retained" reflects real compression or an artifact of multiple-choice metrics that don't test multi-step generation. Listeners get a skeptical look at why a small-angle update isn't necessarily an unimportant one, illustrated by the last layer's FFN, whose removal sends perplexity from 7.60 to 12.35.

Sources:
1. ShortGPT: Layers in Large Language Models are More Redundant Than You Expect — Xin Men, Mingyu Xu, Qingyu Zhang, Bingning Wang, Hongyu Lin, Yaojie Lu, Xianpei Han, Weipeng Chen, 2024
http://arxiv.org/abs/2403.03853
2. ShortGPT: Layers in Large Language Models are More Redundant Than You Expect — Xin Men, Mingyu Xu, Qingyu Zhang, Bingning Wang, Hongyu Lin, Yaojie Lu, Xianpei Han, Weipeng Chen, 2024
https://scholar.google.com/scholar?q=ShortGPT%3A+Layers+in+Large+Language+Models+are+More+Redundant+Than+You+Expect
3. The Unreasonable Ineffectiveness of the Deeper Layers — Andrey Gromov, Kushal Tirumala, Hassan Shapourian, Paolo Glorioso, Daniel A. Roberts, 2024 (arXiv; ICLR 2025)
https://scholar.google.com/scholar?q=The+Unreasonable+Ineffectiveness+of+the+Deeper+Layers
4. Shortened LLaMA: Depth Pruning for Large Language Models with Comparison of Retraining Methods — Bo-Kyeong Kim, Geonmin Kim, Tae-Ho Kim, Thibault Castells, Shinkook Choi, Junho Shin, Hyoung-Kyu Song, 2024
https://scholar.google.com/scholar?q=Shortened+LLaMA%3A+Depth+Pruning+for+Large+Language+Models+with+Comparison+of+Retraining+Methods
5. Compact Language Models via Pruning and Knowledge Distillation (Minitron) — Saurav Muralidharan, Sharath Turuvekere Sreenivas, Raviraj Joshi, Marcin Chochowski, Mostofa Patwary, Mohammad Shoeybi, Bryan Catanzaro, Jan Kautz, Pavlo Molchanov, 2024
https://scholar.google.com/scholar?q=Compact+Language+Models+via+Pruning+and+Knowledge+Distillation+%28Minitron%29
6. SLEB: Streamlining LLMs through Redundancy Verification and Elimination of Transformer Blocks — Jiwon Song, Kyungseok Oh, Taesu Kim, Hyungjun Kim, Yulhwa Kim, Jae-Joon Kim, 2024
https://scholar.google.com/scholar?q=SLEB%3A+Streamlining+LLMs+through+Redundancy+Verification+and+Elimination+of+Transformer+Blocks
7. Reassessing Layer Pruning in LLMs: New Insights and Methods — Yao Lu, Hao Cheng, Yujie Fang, Zeyu Wang, Jiaheng Wei, Dongwei Xu, Qi Xuan, Xiaoniu Yang, Zhaowei Zhu, 2024
https://scholar.google.com/scholar?q=Reassessing+Layer+Pruning+in+LLMs%3A+New+Insights+and+Methods
8. Do Language Models Use Their Depth Efficiently? — Róbert Csordás, Christopher D. Manning, Christopher Potts, 2025
https://scholar.google.com/scholar?q=Do+Language+Models+Use+Their+Depth+Efficiently%3F
9. The Curse of Depth in Large Language Models — Wenfang Sun, Xinyuan Song, Pengxiang Li, Lu Yin, Yefeng Zheng, Shiwei Liu, 2025
https://scholar.google.com/scholar?q=The+Curse+of+Depth+in+Large+Language+Models
10. Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN — Pengxiang Li, Lu Yin, Shiwei Liu, 2024
https://scholar.google.com/scholar?q=Mix-LN%3A+Unleashing+the+Power+of+Deeper+Layers+by+Combining+Pre-LN+and+Post-LN
11. Transformer Layers as Painters — Qi Sun, Marc Pickett, Aakash Kumar Nain, Llion Jones, 2024
https://scholar.google.com/scholar?q=Transformer+Layers+as+Painters
12. Massive Activations in Large Language Models — Mingjie Sun, Xinlei Chen, J. Zico Kolter, Zhuang Liu, 2024
https://scholar.google.com/scholar?q=Massive+Activations+in+Large+Language+Models
13. A Deeper Look at Depth Pruning of LLMs — Shoaib Ahmed Siddiqui, Xin Dong, Greg Heinrich, Thomas Breuel, Jan Kautz, David Krueger, Pavlo Molchanov, 2024
https://scholar.google.com/scholar?q=A+Deeper+Look+at+Depth+Pruning+of+LLMs
14. What Matters in Transformers? Not All Attention is Needed — Shwai He, Guoheng Sun, Zheyu Shen, Ang Li, 2024
https://scholar.google.com/scholar?q=What+Matters+in+Transformers%3F+Not+All+Attention+is+Needed
15. Residual Networks Behave Like Ensembles of Relatively Shallow Networks — Andreas Veit, Michael Wilber, Serge Belongie, 2016
https://scholar.google.com/scholar?q=Residual+Networks+Behave+Like+Ensembles+of+Relatively+Shallow+Networks
16. LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding — Mostafa Elhoushi et al., 2024
https://scholar.google.com/scholar?q=LayerSkip%3A+Enabling+Early+Exit+Inference+and+Self-Speculative+Decoding
17. Compressing LLMs: The Truth is Rarely Pure and Never Simple — Ajay Jaiswal, Zhenyu Zhang, Zhangyang Wang, Yinfei Yang, Shiwei Liu, Yi Zhang, et al., 2023
https://scholar.google.com/scholar?q=Compressing+LLMs%3A+The+Truth+is+Rarely+Pure+and+Never+Simple
Interactive Visualization: ShortGPT: Deleting Redundant LLM Layers, Free Lunch or Artifact?

This episode examines "Transformer Layers as Painters," which asks whether a frozen, ordinarily trained Llama 2 can tolerate having its layers skipped, swapped, reordered, or run in parallel without any retraining. The discussion uses a painter-on-an-assembly-line analogy: the residual stream serves as a shared canvas, so layers read and write the same space. It places the paper alongside related work on residual networks, logit and tuned lens, depth pruning, and the "stages of inference" study. Testing on Llama2-7B, 13B, and 70B plus BERT-Large across five benchmarks, the first results show a sharp split. Removing or swapping the first and last layers collapses performance, while the uniform middle layers barely register a change, and accuracy falls off gradually instead of breaking. Listeners interested in layer pruning, conditional computation, and the latency cost of depth get a look at what a pretrained model can absorb without being trained for it.

Sources:
1. Transformer Layers as Painters — Qi Sun, Marc Pickett, Aakash Kumar Nain, Llion Jones, 2024
http://arxiv.org/abs/2407.09298
2. Eliciting Latent Predictions from Transformers with the Tuned Lens — Nora Belrose, Zach Furman, Logan Smith, Danny Halawi, Igor Ostrovsky, Lev McKinney, Stella Biderman, Jacob Steinhardt (building on nostalgebraist's 2020 'interpreting GPT: the logit lens'), 2023
https://scholar.google.com/scholar?q=Eliciting+Latent+Predictions+from+Transformers+with+the+Tuned+Lens
3. The Remarkable Robustness of LLMs: Stages of Inference? — Vedang Lad, Wes Gurnee, Max Tegmark, 2024
https://scholar.google.com/scholar?q=The+Remarkable+Robustness+of+LLMs%3A+Stages+of+Inference%3F
4. The Unreasonable Ineffectiveness of the Deeper Layers — Andrey Gromov, Kushal Tirumala, Hassan Shapourian, Paolo Glorioso, Daniel A. Roberts, 2024
https://scholar.google.com/scholar?q=The+Unreasonable+Ineffectiveness+of+the+Deeper+Layers
5. Residual Networks Behave Like Ensembles of Relatively Shallow Networks — Andreas Veit, Michael Wilber, Serge Belongie, 2016
https://scholar.google.com/scholar?q=Residual+Networks+Behave+Like+Ensembles+of+Relatively+Shallow+Networks
6. Reducing Transformer Depth on Demand with Structured Dropout (LayerDrop) — Angela Fan, Edouard Grave, Armand Joulin, 2019 (ICLR 2020)
https://scholar.google.com/scholar?q=Reducing+Transformer+Depth+on+Demand+with+Structured+Dropout+%28LayerDrop%29
7. LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding — Mostafa Elhoushi, Akshat Shrivastava, Diana Liskovich, Basil Hosmer, Bram Wasti, et al. (Meta), 2024
https://scholar.google.com/scholar?q=LayerSkip%3A+Enabling+Early+Exit+Inference+and+Self-Speculative+Decoding
8. Mixture-of-Depths: Dynamically Allocating Compute in Transformer-Based Language Models — David Raposo, Sam Ritter, Blake Richards, Timothy Lillicrap, Peter Conway Humphreys, Adam Santoro (Google DeepMind), 2024
https://scholar.google.com/scholar?q=Mixture-of-Depths%3A+Dynamically+Allocating+Compute+in+Transformer-Based+Language+Models
9. Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach — Jonas Geiping, Sean McLeish, Neel Jain, John Kirchenbauer, Siddharth Singh, Brian R. Bartoldson, Bhavya Kailkhura, Abhinav Bhatele, Tom Goldstein, 2025
https://scholar.google.com/scholar?q=Scaling+up+Test-Time+Compute+with+Latent+Reasoning%3A+A+Recurrent+Depth+Approach
10. Interpreting GPT: the logit lens (blog post) and Eliciting Latent Predictions from Transformers with the Tuned Lens — nostalgebraist; Nora Belrose et al., 2020 / 2023
https://scholar.google.com/scholar?q=Interpreting+GPT%3A+the+logit+lens+%28blog+post%29+and+Eliciting+Latent+Predictions+from+Transformers+with+the+Tuned+Lens
11. Do Language Models Use Their Depth Efficiently? — Róbert Csordás, Christopher D. Manning, Christopher Potts, 2025
https://scholar.google.com/scholar?q=Do+Language+Models+Use+Their+Depth+Efficiently%3F
12. The Curse of Depth in Large Language Models — Wenfang Sun, Xinyuan Song, Pengxiang Li, Lu Yin, Yefeng Zheng, Shiwei Liu, 2025
https://scholar.google.com/scholar?q=The+Curse+of+Depth+in+Large+Language+Models
13. Massive Activations in Large Language Models — Mingjie Sun, Xinlei Chen, J. Zico Kolter, Zhuang Liu, 2024
https://scholar.google.com/scholar?q=Massive+Activations+in+Large+Language+Models
14. Highway and Residual Networks Learn Unrolled Iterative Estimation — Klaus Greff, Rupesh K. Srivastava, Jürgen Schmidhuber, 2016
https://scholar.google.com/scholar?q=Highway+and+Residual+Networks+Learn+Unrolled+Iterative+Estimation
15. Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding — Jun Zhang et al., 2023
https://scholar.google.com/scholar?q=Draft+%26+Verify%3A+Lossless+Large+Language+Model+Acceleration+via+Self-Speculative+Decoding
16. Universal Transformers / Deep Equilibrium Models — Mostafa Dehghani et al.; Shaojie Bai, J. Zico Kolter, Vladlen Koltun, 2019
https://scholar.google.com/scholar?q=Universal+Transformers+%2F+Deep+Equilibrium+Models
17. Your Transformer is Secretly Linear — Anton Razzhigaev et al., 2024
https://scholar.google.com/scholar?q=Your+Transformer+is+Secretly+Linear
18. Layer by Layer: Uncovering Hidden Representations in Language Models — Oscar Skean et al., 2025
https://scholar.google.com/scholar?q=Layer+by+Layer%3A+Uncovering+Hidden+Representations+in+Language+Models
19. Transformer Feed-Forward Layers Are Key-Value Memories — Mor Geva, Roei Schuster, Jonathan Berant, Omer Levy, 2021
https://scholar.google.com/scholar?q=Transformer+Feed-Forward+Layers+Are+Key-Value+Memories
20. BERT Rediscovers the Classical NLP Pipeline — Ian Tenney, Dipanjan Das, Ellie Pavlick, 2019
https://scholar.google.com/scholar?q=BERT+Rediscovers+the+Classical+NLP+Pipeline

This episode examines "Redwood," a company technical report claiming an AI system took a two-architect spec to verified RTL for a spatial dataflow inference accelerator in under two weeks, with 95% coverage per block and a Qwen3-0.6B bring-up in week three. The hosts argue over whether the 14% first-silicon success statistic actually motivates the paper's single-spec approach, or whether handoffs between architecture, RTL, verification and kernels are the real schedule bottleneck. They explain why batch-size-one physical-AI inference is memory-bound rather than compute-bound, and walk through the tile design: RISC-V control core, matrix engine, vector engine, 512 KB scratchpad and a credit-based network-on-chip. Throughout, they separate the FPGA results measured on a Versal VPK180 from the Samsung 8 nm figures, which are projections. Those projections include the headline 1.75x throughput, 1.9x lower power and 3.4x performance-per-watt versus a Jetson Orin Nano, and the authors' claim of early recursive self-improvement. It's useful for listeners who want to judge AI-driven chip design claims by what was actually demonstrated.

Sources:
1. Redwood: A Frontier AI Accelerator Designed, Verified, and Deployed from Scratch in 2 Weeks by AI — Architect Labs, 2026
http://arxiv.org/abs/2608.26418
2. ChipNeMo: Domain-Adapted LLMs for Chip Design — Mingjie Liu et al. (NVIDIA), 2023
https://scholar.google.com/scholar?q=ChipNeMo%3A+Domain-Adapted+LLMs+for+Chip+Design
3. Pushing the Limits of Machine Design: Automated CPU Design with AI — Shuyao Cheng et al., 2023
https://scholar.google.com/scholar?q=Pushing+the+Limits+of+Machine+Design%3A+Automated+CPU+Design+with+AI
4. Chip-Chat: Challenges and Opportunities in Conversational Hardware Design — Jason Blocklove, Siddharth Garg, Ramesh Karri, Hammond Pearce, 2023
https://scholar.google.com/scholar?q=Chip-Chat%3A+Challenges+and+Opportunities+in+Conversational+Hardware+Design
5. VerilogEval: Evaluating Large Language Models for Verilog Code Generation — Mingjie Liu, Nathaniel Pinckney, Brucek Khailany, Haoxing Ren, 2023
https://scholar.google.com/scholar?q=VerilogEval%3A+Evaluating+Large+Language+Models+for+Verilog+Code+Generation
6. CVDP: Verilog Design Problems benchmark for RTL design and verification agents — Nathaniel Pinckney et al. (NVIDIA), 2025
https://scholar.google.com/scholar?q=CVDP%3A+Verilog+Design+Problems+benchmark+for+RTL+design+and+verification+agents
7. FVEval: Understanding Language Model Capabilities in Formal Verification of Digital Hardware — Minwoo Kang et al., 2024
https://scholar.google.com/scholar?q=FVEval%3A+Understanding+Language+Model+Capabilities+in+Formal+Verification+of+Digital+Hardware
8. Roofline: An Insightful Visual Performance Model for Multicore Architectures — Samuel Williams, Andrew Waterman, David Patterson, 2009
https://scholar.google.com/scholar?q=Roofline%3A+An+Insightful+Visual+Performance+Model+for+Multicore+Architectures
9. FlightLLM: Efficient Large Language Model Inference with a Complete Mapping Flow on FPGAs — Shulin Zeng et al., 2024
https://scholar.google.com/scholar?q=FlightLLM%3A+Efficient+Large+Language+Model+Inference+with+a+Complete+Mapping+Flow+on+FPGAs
10. MLPerf Power: Benchmarking the Energy Efficiency of Machine Learning Systems from Microwatts to Megawatts for Sustainable AI — Arya Tschand et al., 2024
https://scholar.google.com/scholar?q=MLPerf+Power%3A+Benchmarking+the+Energy+Efficiency+of+Machine+Learning+Systems+from+Microwatts+to+Megawatts+for+Sustainable+AI
11. An Empirical Study of Qwen3 Quantization — Xingyu Zheng et al., 2025
https://scholar.google.com/scholar?q=An+Empirical+Study+of+Qwen3+Quantization
12. SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models — Guangxuan Xiao et al., 2023
https://scholar.google.com/scholar?q=SmoothQuant%3A+Accurate+and+Efficient+Post-Training+Quantization+for+Large+Language+Models
13. AlphaEvolve: A coding agent for scientific and algorithmic discovery — Google DeepMind (Novikov et al.), 2025
https://scholar.google.com/scholar?q=AlphaEvolve%3A+A+coding+agent+for+scientific+and+algorithmic+discovery
14. KernelBench: Can LLMs Write Efficient GPU Kernels? — Anne Ouyang et al., 2025
https://scholar.google.com/scholar?q=KernelBench%3A+Can+LLMs+Write+Efficient+GPU+Kernels%3F
15. A graph placement methodology for fast chip design (and the follow-up critique 'Reevaluating Google's Reinforcement Learning for IC Macro Placement') — Azalia Mirhoseini et al. (2021); Igor Markov (2024), 2021/2024
https://scholar.google.com/scholar?q=A+graph+placement+methodology+for+fast+chip+design+%28and+the+follow-up+critique+%27Reevaluating+Google%27s+Reinforcement+Learning+for+IC+Macro+Placement%27%29
16. Gemmini: Enabling Systematic Deep-Learning Architecture Evaluation via Full-Stack Integration — Hasan Genc et al., 2021
https://scholar.google.com/scholar?q=Gemmini%3A+Enabling+Systematic+Deep-Learning+Architecture+Evaluation+via+Full-Stack+Integration

This episode examines "Full-bandwidth transformer," which asks whether a model can feed its entire top-layer hidden state, rather than just the roughly 17-bit sampled token, back into the bottom of the stack at every decoding step. It explains why the KV cache doesn't already solve this: attention is full-bandwidth horizontally, but a state at layer l can only be read by higher layers, and the top layer's output is never cached. The discussion also weighs the design against RNNs and the Feedback Transformer. Nothing is overwritten and the full KV cache is retained, but sequential dependence is the real tension, and it is compared with Coconut, which replaces tokens rather than augmenting them. The hosts then cover how the paper keeps training parallel with multi-pass, Jacobi-style training and a gated fusion of state and token embedding. They cover the pass-mix schedule that keeps decoding stable, where a 75/25 one- and two-pass mix diverges but adding 3% three-pass batches yields a plateau. They set the headline claims aside for scrutiny: about 1.5x effective tokens, matching baselines trained on 2x data, negligible decoding cost, and shorter reasoning traces. Listeners interested in scaling limits, chain-of-thought's depth bottleneck, and getting more out of each training token will find it useful.

Sources:
1. Full-bandwidth transformer — Xi Wang, Ziyang Cai, Zheng Zhan, Harry Dong, Ying Fan, Gustavo de Rosa, Tim Pearce, John Langford, 2026
http://arxiv.org/abs/2608.08888
2. Addressing Some Limitations of Transformers with Feedback Memory — Angela Fan, Thibaut Lavril, Edouard Grave, Armand Joulin, Sainbayar Sukhbaatar, 2020 (arXiv; later versions 2021)
https://scholar.google.com/scholar?q=Addressing+Some+Limitations+of+Transformers+with+Feedback+Memory
3. Training Large Language Models to Reason in a Continuous Latent Space (Coconut) — Shibo Hao, Sainbayar Sukhbaatar, DiJia Su, Xian Li, Zhiting Hu, Jason Weston, Yuandong Tian, 2024
https://scholar.google.com/scholar?q=Training+Large+Language+Models+to+Reason+in+a+Continuous+Latent+Space+%28Coconut%29
4. Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach — Jonas Geiping, Sean McLeish, Neel Jain, John Kirchenbauer, Siddharth Singh, Brian Bartoldson, Bhavya Kailkhura, Abhinav Bhatele, Tom Goldstein, 2025
https://scholar.google.com/scholar?q=Scaling+up+Test-Time+Compute+with+Latent+Reasoning%3A+A+Recurrent+Depth+Approach
5. Chain of Thought Empowers Transformers to Solve Inherently Serial Problems — Zhiyuan Li, Hong Liu, Denny Zhou, Tengyu Ma, 2024
https://scholar.google.com/scholar?q=Chain+of+Thought+Empowers+Transformers+to+Solve+Inherently+Serial+Problems
6. Think before you speak: Training Language Models With Pause Tokens — Sachin Goyal, Ziwei Ji, Ankit Singh Rawat, Aditya Krishna Menon, Sanjiv Kumar, Vaishnavh Nagarajan, 2023
https://scholar.google.com/scholar?q=Think+before+you+speak%3A+Training+Language+Models+With+Pause+Tokens
7. PonderLM-2: Pretraining LLM with Latent Thoughts in Continuous Space — Boyi Zeng et al., 2025
https://scholar.google.com/scholar?q=PonderLM-2%3A+Pretraining+LLM+with+Latent+Thoughts+in+Continuous+Space
8. ParaRNN: Unlocking Parallel Training of Nonlinear RNNs for Large Language Models — Federico Danieli et al., 2025
https://scholar.google.com/scholar?q=ParaRNN%3A+Unlocking+Parallel+Training+of+Nonlinear+RNNs+for+Large+Language+Models
9. Parallelizing Non-linear Sequential Models over the Sequence Length (DEER) — Yi Heng Lim et al., 2024
https://scholar.google.com/scholar?q=Parallelizing+Non-linear+Sequential+Models+over+the+Sequence+Length+%28DEER%29
10. Accelerating Feedforward Computation via Parallel Nonlinear Equation Solving — Yang Song et al., 2021
https://scholar.google.com/scholar?q=Accelerating+Feedforward+Computation+via+Parallel+Nonlinear+Equation+Solving
11. The Expressive Power of Transformers with Chain of Thought — William Merrill, Ashish Sabharwal, 2024
https://scholar.google.com/scholar?q=The+Expressive+Power+of+Transformers+with+Chain+of+Thought
12. Let's Think Dot by Dot: Hidden Computation in Transformer Language Models — Jacob Pfau, William Merrill, Samuel R. Bowman, 2024
https://scholar.google.com/scholar?q=Let%27s+Think+Dot+by+Dot%3A+Hidden+Computation+in+Transformer+Language+Models
13. Reasoning with Latent Thoughts: On the Power of Looped Transformers — Nikunj Saunshi et al., 2025
https://scholar.google.com/scholar?q=Reasoning+with+Latent+Thoughts%3A+On+the+Power+of+Looped+Transformers
14. Scaling Latent Reasoning via Looped Language Models (Ouro) — Rui-Jie Zhu et al., 2025
https://scholar.google.com/scholar?q=Scaling+Latent+Reasoning+via+Looped+Language+Models+%28Ouro%29
15. Scaling Data-Constrained Language Models — Niklas Muennighoff et al., 2023
https://scholar.google.com/scholar?q=Scaling+Data-Constrained+Language+Models
16. Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety — Tomek Korbak et al., 2025
https://scholar.google.com/scholar?q=Chain+of+Thought+Monitorability%3A+A+New+and+Fragile+Opportunity+for+AI+Safety
17. Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting — Miles Turpin et al., 2023
https://scholar.google.com/scholar?q=Language+Models+Don%27t+Always+Say+What+They+Think%3A+Unfaithful+Explanations+in+Chain-of-Thought+Prompting
18. The Illusion of State in State-Space Models — William Merrill, Jackson Petty, Ashish Sabharwal, 2024
https://scholar.google.com/scholar?q=The+Illusion+of+State+in+State-Space+Models
19. Block-Recurrent Transformers — DeLesley Hutchins et al., 2022
https://scholar.google.com/scholar?q=Block-Recurrent+Transformers
Interactive Visualization: Full-Bandwidth Transformer: Feeding Hidden States Back Into the Stack

This episode examines "Learning to Solve Hard Problems in RL for LLMs by Never Giving Up," which finds that RL post-training on math, code, and agentic coding benchmarks disproportionately improves performance on problems models already handle reasonably well, while barely moving the needle on the hardest tasks — a pattern the authors dub the Matthew Effect, after Robert Merton's sociology of science concept. The discussion contrasts two explanations for this skew: the intuitive "signal loss" account, where GRPO's group-relative reward gives zero gradient when every sampled completion fails a hard problem, versus the paper's "signal efficiency" argument, that compute is instead wasted reconfirming easy problems the model has already solved. That distinction matters because it points to different fixes — simply sampling more completions per prompt doesn't help, but reallocating sampling toward unsolved problems does, which motivates the paper's proposed method of persistently re-sampling unsolved prompts rather than discarding them. Listeners get a concrete walkthrough of the controlled GSM8k experiment used to test these competing hypotheses, including a difficulty-tiered evaluation and a K-sweep that challenges conventional assumptions about RL sample scaling. The episode is a useful listen for anyone weighing whether reinforcement learning can actually push language models past their pretrained capability ceiling, or whether it's mainly sharpening skills the model already has.

Sources:
1. Learning to Solve Hard Problems in RL for LLMs by Never Giving Up — Michael Noukhovitch, Hamish Ivison, Nathan Lambert, Aaron Courville, 2026
http://arxiv.org/abs/2609.13443v1
2. Prioritized Experience Replay — Tom Schaul, John Quan, Ioannis Antonoglou, David Silver, 2015
https://scholar.google.com/scholar?q=Prioritized+Experience+Replay
3. DAPO: An Open-Source LLM Reinforcement Learning System at Scale — Qiying Yu, Zheng Zhang, et al. (ByteDance Seed), 2025
https://scholar.google.com/scholar?q=DAPO%3A+An+Open-Source+LLM+Reinforcement+Learning+System+at+Scale
4. Prioritized Level Replay — Minqi Jiang, Edward Grefenstette, Tim Rocktäschel, 2021
https://scholar.google.com/scholar?q=Prioritized+Level+Replay
5. Curriculum Learning — Yoshua Bengio, Jérôme Louradour, Ronan Collobert, Jason Weston, 2009
https://scholar.google.com/scholar?q=Curriculum+Learning
6. Never Give Up: Learning Directed Exploration Strategies — A. P. Badia, P. Sprechmann, A. Vitvitskyi, D. Guo, B. Piot, S. Kapturowski, O. Tieleman, M. Arjovsky, A. Pritzel, A. Bolt, C. Blundell, 2020
https://scholar.google.com/scholar?q=Never+Give+Up%3A+Learning+Directed+Exploration+Strategies
7. Reinforce-Ada: An Adaptive Sampling Framework under Non-linear RL Objectives — W. Xiong, C. Ye, B. Liao, H. Dong, X. Xu, C. Monz, J. Bian, N. Jiang, T. Zhang, 2025
https://scholar.google.com/scholar?q=Reinforce-Ada%3A+An+Adaptive+Sampling+Framework+under+Non-linear+RL+Objectives
8. POLARIS: A Post-Training Recipe for Scaling Reinforcement Learning on Advanced Reasoning Models — C. An, Z. Xie, X. Li, L. Li, J. Zhang, S. Gong, M. Zhong, J. Xu, X. Qiu, M. Wang, L. Kong, 2025
https://scholar.google.com/scholar?q=POLARIS%3A+A+Post-Training+Recipe+for+Scaling+Reinforcement+Learning+on+Advanced+Reasoning+Models
9. RL Grokking Recipe: How Does RL Unlock and Transfer New Algorithms in LLMs? — Y. Sun, Y. Cao, P. Huang, H. Bai, H. Hajishirzi, N. Dziri, D. Song, 2025
https://scholar.google.com/scholar?q=RL+Grokking+Recipe%3A+How+Does+RL+Unlock+and+Transfer+New+Algorithms+in+LLMs%3F
10. The Art of Scaling Reinforcement Learning Compute for LLMs — D. Khatri, L. Madaan, R. Tiwari, R. Bansal, S. S. Duvvuri, M. Zaheer, I. S. Dhillon, D. Brandfonbrener, R. Agarwal, 2025
https://scholar.google.com/scholar?q=The+Art+of+Scaling+Reinforcement+Learning+Compute+for+LLMs
11. The Primacy Bias in Deep Reinforcement Learning — E. Nikishin, M. Schwarzer, P. D'Oro, P.-L. Bacon, A. Courville, 2022
https://scholar.google.com/scholar?q=The+Primacy+Bias+in+Deep+Reinforcement+Learning
12. Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts (GRESO) — H. Zheng, Y. Zhou, B. R. Bartoldson, B. Kailkhura, F. Lai, J. Zhao, B. Chen, 2025
https://scholar.google.com/scholar?q=Act+Only+When+It+Pays%3A+Efficient+Reinforcement+Learning+for+LLM+Reasoning+via+Selective+Rollouts+%28GRESO%29
Interactive Visualization: The Matthew Effect in RL: Rich Get Richer on Hard Problems

This episode examines the NVQLink architecture, a NVIDIA-led collaboration with nine national labs and research institutions that tightly couples high-performance computing with quantum processors. The discussion centers on why reaction time—not just throughput—is critical for quantum error correction, since qubits decohere while waiting for classical correction signals to arrive. It breaks down the system's building blocks (QPU, QSC, PPU, and the real-time host) and explores a counterintuitive design choice: routing the real-time interconnect over commodity, unreliable Ethernet rather than PCIe, trading a seemingly obvious direct connection for datacenter-scale networking. The episode also connects the work to a decade-old 2017 paper by co-author Travis Humble that first framed QPUs as HPC accelerators, showing how NVQLink turns that early concept into a measured, functioning system with sub-4-microsecond round-trip latency. Listeners interested in the intersection of classical computing infrastructure and quantum hardware will find the engineering tradeoffs and vocabulary breakdown a useful entry point into a fast-moving, unfamiliar corner of systems design.

Sources:
1. Platform Architecture for Tight Coupling of High-Performance Computing with Quantum Processors — Shane A. Caldwell, Moein Khazraee, Elena Agostini, Tom Lassiter, Corey Simpson, Omri Kahalon, Mrudula Kanuri, Jin-Sung Kim, Sam Stanwyck, Muyuan Li, Jan Olle, Christopher Chamberland, Ben Howe, Bruno Schmitt, Justin G. Lietz, Alex McCaskey, Jun Ye, Ang Li, Alicia B. Magann, Corey I. Ostrove, Kenneth Rudinger, Robin Blume-Kohout, Kevin Young, Nathan E. Miller, Yilun Xu, Gang Huang, Irfan Siddiqi, John Lange, Christopher Zimmer, Travis Humble, 2025
http://arxiv.org/abs/2510.25213
2. High-Performance Computing with Quantum Processing Units — K. Britt, T. Humble, 2017
https://scholar.google.com/scholar?q=High-Performance+Computing+with+Quantum+Processing+Units
3. CUDA Quantum: The Platform for Integrated Quantum-Classical Computing — A. McCaskey et al. (NVIDIA), 2024
https://scholar.google.com/scholar?q=CUDA+Quantum%3A+The+Platform+for+Integrated+Quantum-Classical+Computing
4. Quantum-Centric Supercomputing for Materials Science: A Perspective on Challenges and Future Directions — Y. Alexeev et al. (IBM, national labs), 2021 (Future Generation Computer Systems)
https://scholar.google.com/scholar?q=Quantum-Centric+Supercomputing+for+Materials+Science%3A+A+Perspective+on+Challenges+and+Future+Directions
5. MLIR: Scaling Compiler Infrastructure for Domain Specific Computation — C. Lattner, M. Amini, U. Bondhugula, A. Cohen, A. Davis, J. Pienaar, R. Riddle, T. Shpeisman, N. Vasilache, O. Zinenko (Google), 2021 (CGO)
https://scholar.google.com/scholar?q=MLIR%3A+Scaling+Compiler+Infrastructure+for+Domain+Specific+Computation
6. QICK: Quantum Instrumentation Control Kit — L. Stefanazzi, K. Treptow, N. Wilcer, et al. (Fermilab), 2022 (Review of Scientific Instruments)
https://scholar.google.com/scholar?q=QICK%3A+Quantum+Instrumentation+Control+Kit
7. Real-Time Processing Systems for Next-Generation Quantum Computers (representative of the QEC-decoder-latency literature, e.g. Google Quantum AI's real-time decoding work) — Google Quantum AI et al., 2023–2024
https://scholar.google.com/scholar?q=Real-Time+Processing+Systems+for+Next-Generation+Quantum+Computers+%28representative+of+the+QEC-decoder-latency+literature%2C+e.g.+Google+Quantum+AI%27s+real-time+decoding+work%29
8. Learning high-accuracy error decoding for quantum processors (AlphaQubit) — Johannes Bausch, Andrew W. Senior, Francisco J. H. Heras, et al., 2024
https://scholar.google.com/scholar?q=Learning+high-accuracy+error+decoding+for+quantum+processors+%28AlphaQubit%29
9. Evaluation of the Classical Hardware Requirements for Large-Scale Quantum Computations — Daan Camps, Ermal Rrapaj, Katherine Klymko, Brian Austin, Nicholas J. Wright, 2024
https://scholar.google.com/scholar?q=Evaluation+of+the+Classical+Hardware+Requirements+for+Large-Scale+Quantum+Computations
10. Parallel window decoding enables scalable fault tolerant quantum computation — Luka Skoric, Dan E. Browne, Kenton M. Barnes, Neil I. Gillespie, Earl T. Campbell, 2023
https://scholar.google.com/scholar?q=Parallel+window+decoding+enables+scalable+fault+tolerant+quantum+computation
11. Quantum error correction below the surface code threshold — Google Quantum AI and Collaborators, 2024
https://scholar.google.com/scholar?q=Quantum+error+correction+below+the+surface+code+threshold
12. A Game of Surface Codes: Large-Scale Quantum Computing with Lattice Surgery — Daniel Litinski, 2019
https://scholar.google.com/scholar?q=A+Game+of+Surface+Codes%3A+Large-Scale+Quantum+Computing+with+Lattice+Surgery
Interactive Visualization: NVQLink: Wiring Supercomputers Directly to Quantum Chips

This episode examines ProgramBench, a new benchmark testing whether frontier language models can rebuild working software from just a compiled binary and its documentation, with no source code, scaffolding, or prescribed architecture to work from. Across nine frontier models and two hundred tasks spanning CLI tools up to FFmpeg, SQLite, and the PHP interpreter, zero tasks were fully resolved, exposing a stark gap between patching existing code and making the upstream architectural decisions—language choice, module boundaries, data structures, error handling—that real software design requires. The discussion unpacks the benchmark's clever self-hosting trick: an LLM agent fuzzes the reference binary to build a behavioral test suite, enabling black-box grading that judges programs by what they do rather than how closely they mimic the original source. Framing the results against Parnas's classic work on information hiding and modular decomposition, the conversation argues that current agents default to monolithic, unstructured code once nobody hands them a skeleton to fill in. It's a sobering data point for anyone assuming coding agents are close to functioning as autonomous software architects rather than sophisticated patch-writers.

Sources:
1. ProgramBench: Can Language Models Rebuild Programs From Scratch? — John Yang, Kilian Lieret, Jeffrey Ma, Parth Thakkar, Dmitrii Pedchenko, Sten Sootla, Emily McMilin, Pengcheng Yin, Rui Hou, Gabriel Synnaeve, Diyi Yang, Ofir Press, 2026
http://arxiv.org/abs/2605.03546
2. On the Criteria To Be Used in Decomposing Systems into Modules — D. L. Parnas, 1972
https://scholar.google.com/scholar?q=On+the+Criteria+To+Be+Used+in+Decomposing+Systems+into+Modules
3. Measuring Coding Challenge Competence With APPS — Dan Hendrycks, Steven Basart, Saurav Kadavath, et al., 2021
https://scholar.google.com/scholar?q=Measuring+Coding+Challenge+Competence+With+APPS
4. SWE-bench: Can Language Models Resolve Real-World GitHub Issues? — Carlos E. Jimenez, John Yang, Alexander Wettig, et al., 2024
https://scholar.google.com/scholar?q=SWE-bench%3A+Can+Language+Models+Resolve+Real-World+GitHub+Issues%3F
5. SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering — John Yang, Carlos E. Jimenez, Alexander Wettig, et al., 2024
https://scholar.google.com/scholar?q=SWE-agent%3A+Agent-Computer+Interfaces+Enable+Automated+Software+Engineering
6. QuickCheck: A Lightweight Tool for Random Testing of Haskell Programs — Koen Claessen, John Hughes, 2000
https://scholar.google.com/scholar?q=QuickCheck%3A+A+Lightweight+Tool+for+Random+Testing+of+Haskell+Programs
7. Finding and Understanding Bugs in C Compilers — Xuejun Yang, Yang Chen, Eric Eide, John Regehr, 2011
https://scholar.google.com/scholar?q=Finding+and+Understanding+Bugs+in+C+Compilers
8. CodeT: Code Generation with Generated Tests — Bei Chen, Fengji Zhang, Anh Nguyen, et al., 2022
https://scholar.google.com/scholar?q=CodeT%3A+Code+Generation+with+Generated+Tests
9. Commit0: Library Generation from Scratch — Wenting Zhao, Nan Jiang, Celine Lee, Justin T Chiu, Claire Cardie, Matthias Gallé, Alexander M Rush, 2024
https://scholar.google.com/scholar?q=Commit0%3A+Library+Generation+from+Scratch
10. DevBench: A Comprehensive Benchmark for Software Development — Bowen Li, Wenhan Wu, Ziwei Tang, et al. (incl. John Yang, Ofir Press), 2024
https://scholar.google.com/scholar?q=DevBench%3A+A+Comprehensive+Benchmark+for+Software+Development
11. NL2Repo-bench: Towards Long-Horizon Repository Generation Evaluation of Coding Agents — Jingzhe Ding et al., 2026
https://scholar.google.com/scholar?q=NL2Repo-bench%3A+Towards+Long-Horizon+Repository+Generation+Evaluation+of+Coding+Agents
12. LLM4Decompile: Decompiling Binary Code with Large Language Models — Hanzhuo Tan, Qi Luo, Jing Li, Yuqun Zhang, 2024
https://scholar.google.com/scholar?q=LLM4Decompile%3A+Decompiling+Binary+Code+with+Large+Language+Models
13. Position: Humans Are Missing from AI Coding Agent Research — Zora Zhiruo Wang, John Yang, Kilian Lieret, et al., 2026
https://scholar.google.com/scholar?q=Position%3A+Humans+Are+Missing+from+AI+Coding+Agent+Research
Interactive Visualization: ProgramBench: Frontier Agents Fail to Rebuild Software from Scratch

This episode examines "An Alien Mind," a single-author essay from an OpenAI research leader arguing that internal results point toward sustained progress and eventually recursive self-improvement, and asks what an outside reader could actually verify. The hosts note that the essay offers no methods, tables, or error bars. They contrast its scaling claims with quantitative work like the Kaplan and Hoffmann scaling-law papers, which give fitted curves and exponents. They also discuss the essay's admission that easy-to-measure capabilities improve faster than hard-to-quantify ones, which makes progress harder to gauge. A large part of the discussion covers how the essay defines alignment: goal alignment versus value alignment, and whether that split is a real testable distinction or a blurry one. The hosts also introduce chain-of-thought monitoring, its fragility when reasoning is trained to look good, and safety cases as a basis for mandated bars. Listeners get a skeptical, evidence-focused look at what a lab insider's claims about AI progress and safety would need in order to hold up.

Sources:
1. Recursive Self-Improvement and Alignment: OpenAI's "An Alien Mind"
https://openai.com/index/an-alien-mind/
2. Artificial Intelligence, Values, and Alignment — Iason Gabriel, 2020
https://scholar.google.com/scholar?q=Artificial+Intelligence%2C+Values%2C+and+Alignment
3. Training language models to follow instructions with human feedback (InstructGPT) — Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, et al. (OpenAI), 2022
https://scholar.google.com/scholar?q=Training+language+models+to+follow+instructions+with+human+feedback+%28InstructGPT%29
4. The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions — Eric Wallace, Kai Xiao, Reimar Leike, Lilian Weng, Johannes Heidecke, Alex Beutel (OpenAI), 2024
https://scholar.google.com/scholar?q=The+Instruction+Hierarchy%3A+Training+LLMs+to+Prioritize+Privileged+Instructions
5. Emergent Misalignment: Narrow Finetuning Can Produce Broadly Misaligned LLMs — Jan Betley, Daniel Tan, Niels Warncke, Anna Sztyber-Betley, Xuchan Bao, Martín Soto, Nathan Labenz, Owain Evans, 2025
https://scholar.google.com/scholar?q=Emergent+Misalignment%3A+Narrow+Finetuning+Can+Produce+Broadly+Misaligned+LLMs
6. Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety — Korbak et al. (multi-lab position paper), 2025
https://scholar.google.com/scholar?q=Chain+of+Thought+Monitorability%3A+A+New+and+Fragile+Opportunity+for+AI+Safety
7. Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation — Baker et al. (OpenAI), 2025
https://scholar.google.com/scholar?q=Monitoring+Reasoning+Models+for+Misbehavior+and+the+Risks+of+Promoting+Obfuscation
8. Clarifying AI alignment — Paul Christiano, 2018
https://scholar.google.com/scholar?q=Clarifying+AI+alignment
9. Deep Reinforcement Learning from Human Preferences — Christiano, Leike, Brown, Martic, Legg, Amodei, 2017
https://scholar.google.com/scholar?q=Deep+Reinforcement+Learning+from+Human+Preferences
10. Alignment Faking in Large Language Models — Greenblatt et al. (Anthropic and Redwood Research), 2024
https://scholar.google.com/scholar?q=Alignment+Faking+in+Large+Language+Models
11. The Persona Selection Model — Anthropic, 2026
https://scholar.google.com/scholar?q=The+Persona+Selection+Model
12. Goal Misgeneralization in Deep Reinforcement Learning — Langosco, Koch, Sharkey, Pfau, Krueger, 2022
https://scholar.google.com/scholar?q=Goal+Misgeneralization+in+Deep+Reinforcement+Learning
13. Risks from Learned Optimization in Advanced Machine Learning Systems — Hubinger, van Merwe, Mikulik, Skalse, Garrabrant, 2019
https://scholar.google.com/scholar?q=Risks+from+Learned+Optimization+in+Advanced+Machine+Learning+Systems
14. Weak-to-Strong Generalization: Eliciting Strong Capabilities with Weak Supervision — Burns et al. (OpenAI), 2023
https://scholar.google.com/scholar?q=Weak-to-Strong+Generalization%3A+Eliciting+Strong+Capabilities+with+Weak+Supervision
15. Training Language Models to Self-Report / Confessions — OpenAI Alignment team, 2025-2026
https://scholar.google.com/scholar?q=Training+Language+Models+to+Self-Report+%2F+Confessions
Interactive Visualization: Recursive Self-Improvement and Alignment: OpenAI's "An Alien Mind"

This episode examines "Dream-RSI: Recursive Self-Improvement through Evolving Worlds," which proposes making the exploration strategy of a discovery system — not the underlying model — the target of recursive self-improvement. Rather than optimizing candidate solutions directly, the system optimizes the policy that decides how to search: which branches to expand, how to parallelize workers, and when to stop, while the underlying coding agent (Gemini, in the paper's experiments) stays fixed. The key innovation is "dreaming": replaying an already-recorded discovery tree of past generate-evaluate attempts as a cheap simulator, letting new exploration policies be scored for free against historical outcomes instead of running costly new agent calls. This produces a three-stage loop — online exploration to grow the tree, constructing a replay simulator from it, then "dreaming" to test and select better policies before redeploying them — addressing the classic problem that policy-level exploration research suffers from painfully delayed feedback. The discussion situates the work against prior exploration methods like bandit algorithms, RL², Never Give Up, and FunSearch, making it a useful listen for anyone interested in how search-strategy meta-optimization, rather than raw model capability, might be the next lever for scaling AI-driven discovery.

Sources:
1. Dream-RSI: Recursive Self-Improvement through Evolving Worlds — Tong Zheng, Xidong Wu, Zheng Zhang, Zhankui He, Chaoyi Zhang, Benjamin Coleman, Ruoqiao Wei, Di Bai, Haolin Liu, Rui Liu, Xue Wang, Yue Zhuan, Wang-Cheng Kang, Renkai Xiang, Heng Huang, Xinwu Cheng, Yunsong Guo, 2026
http://arxiv.org/abs/2609.14858
2. RL²: Fast Reinforcement Learning via Slow Reinforcement Learning — Yan Duan, John Schulman, Xi Chen, Peter L. Bartlett, Ilya Sutskever, Pieter Abbeel, 2016
https://scholar.google.com/scholar?q=RL%C2%B2%3A+Fast+Reinforcement+Learning+via+Slow+Reinforcement+Learning
3. Never Give Up: Learning Directed Exploration Strategies — Adrià Puigdomènech Badia, Pablo Sprechmann, Alex Vitvitskyi, et al. (DeepMind), 2020
https://scholar.google.com/scholar?q=Never+Give+Up%3A+Learning+Directed+Exploration+Strategies
4. Mathematical discoveries from program search with large language models (FunSearch) — Bernardino Romera-Paredes, Mohammadamin Barekatain, Alexander Novikov, et al. (Google DeepMind), 2024
https://scholar.google.com/scholar?q=Mathematical+discoveries+from+program+search+with+large+language+models+%28FunSearch%29
5. Taking the Human Out of the Loop: A Review of Bayesian Optimization — Bobak Shahriari, Kevin Swersky, Ziyu Wang, Ryan P. Adams, Nando de Freitas, 2016
https://scholar.google.com/scholar?q=Taking+the+Human+Out+of+the+Loop%3A+A+Review+of+Bayesian+Optimization
6. AlphaEvolve: A coding agent for scientific and algorithmic discovery — A. Novikov et al., 2025
https://scholar.google.com/scholar?q=AlphaEvolve%3A+A+coding+agent+for+scientific+and+algorithmic+discovery
7. Mastering diverse domains through world models (Dreamer V3) — D. Hafner, J. Pasukonis, J. Ba, T. Lillicrap, 2023
https://scholar.google.com/scholar?q=Mastering+diverse+domains+through+world+models+%28Dreamer+V3%29
8. EvoX: Meta-evolution for automated discovery — S. Liu et al., 2026
https://scholar.google.com/scholar?q=EvoX%3A+Meta-evolution+for+automated+discovery
9. Evaluation-driven scaling for scientific discovery (SimpleTES) — H. Ye et al., 2026
https://scholar.google.com/scholar?q=Evaluation-driven+scaling+for+scientific+discovery+%28SimpleTES%29
Interactive Visualization: Dream-RSI: Teaching AI How to Search, Not Just Solve

This episode examines "Mixture-of-Translators: Translating KV Caches Across Heterogeneous Large Language Models" by Jin-woo Lee and six co-authors from Chungnam National University and KISTI, which tackles the problem of transferring one model's KV cache — its layer-by-layer key/value memory of a processed document — to a completely different model architecture without re-running the original text through it. The discussion explains why this is hard: a KV cache is shaped by the specific depth, width, and head count of the model that produced it, so naive copying fails and a learned "cache translation" is required instead. It surveys prior approaches (Cache-to-Cache, KVComm, Latent Space Communication, and Interlat) and their shared weakness — relying on a single universal mapping or shared latent space — before detailing how Mixture-of-Translators borrows the Mixture-of-Experts routing idea to assign different tokens to different specialized translator modules via a per-token gating network, paired with a Context Correction Loss to correct drift in the target model's own layers. Listeners interested in multi-agent LLM pipelines, cache-augmented generation, or reducing redundant prefill computation across heterogeneous model fleets will find the practical motivation and technical tradeoffs compelling, especially since the paper's honest partial-success framing offers more insight than a clean win would.

Sources:
1. Mixture-of-Translators: Sharing KV Caches Across Different LLMs
https://arxiv.org/pdf/2607.28979
2. Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer — Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, Jeff Dean, 2017
https://scholar.google.com/scholar?q=Outrageously+Large+Neural+Networks%3A+The+Sparsely-Gated+Mixture-of-Experts+Layer
3. Relative Representations Enable Zero-Shot Latent Space Communication — Luca Moschella, Valentino Maiorca, Marco Fumero, Antonio Norelli, Francesco Locatello, Emanuele Rodolà, 2023
https://scholar.google.com/scholar?q=Relative+Representations+Enable+Zero-Shot+Latent+Space+Communication
4. Prompt Cache: Modular Attention Reuse for Low-Latency Inference — In Gim, Guojun Chen, Seung-seob Lee, Nikhil Sarda, Anurag Khandelwal, Lin Zhong, 2024
https://scholar.google.com/scholar?q=Prompt+Cache%3A+Modular+Attention+Reuse+for+Low-Latency+Inference
5. CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving — Yuhan Liu, Hanchen Li, Yihua Cheng, Siddhant Ray, Yuyang Huang, Qizheng Zhang, Kuntai Du, Jiayi Yao, Shan Lu, Ganesh Ananthanarayanan, Michael Maire, Henry Hoffmann, Ari Holtzman, Junchen Jiang, 2024
https://scholar.google.com/scholar?q=CacheGen%3A+KV+Cache+Compression+and+Streaming+for+Fast+Large+Language+Model+Serving
6. Cache-to-cache: Direct semantic communication between large language models — Tianyu Fu, Zihan Min, Hanling Zhang, Jichao Yan, Guohao Dai, Wanli Ouyang, Yu Wang, 2025
https://scholar.google.com/scholar?q=Cache-to-cache%3A+Direct+semantic+communication+between+large+language+models
7. KVComm: Enabling efficient LLM communication through selective KV sharing — Xiangyu Shi, Marco Chiesa, Gerald Q Maguire Jr, Dejan Kostic, 2025
https://scholar.google.com/scholar?q=KVComm%3A+Enabling+efficient+LLM+communication+through+selective+KV+sharing
8. Enabling agents to communicate entirely in latent space (Interlat) — Zhuoyun Du, Runze Wang, Huiyu Bai, Zouying Cao, Xiaoyong Zhu, Yu Cheng, Bo Zheng, Wei Chen, Haochao Ying, 2025
https://scholar.google.com/scholar?q=Enabling+agents+to+communicate+entirely+in+latent+space+%28Interlat%29
9. Latent space communication via KV cache alignment (LSC) — Lucio M Dery, Zohar Yahav, Henry Prior, Qixuan Feng, Jiajun Shen, Arthur Szlam, 2026
https://scholar.google.com/scholar?q=Latent+space+communication+via+KV+cache+alignment+%28LSC%29
10. Fast state restoration in LLM serving with HCache — Shiwei Gao, Youmin Chen, Jiwu Shu, 2025
https://scholar.google.com/scholar?q=Fast+state+restoration+in+LLM+serving+with+HCache
Interactive Visualization: Mixture-of-Translators: Sharing KV Caches Across Different LLMs

This episode explores a Google DeepMind paper proposing that separately trained language models can exchange raw internal state — the key-value cache built during transformer inference — through a shared "global latent space," rather than communicating only through text. The hosts unpack why text is a lossy bottleneck for inter-model communication, and how lightweight, frozen-weight adapter pairs let each model translate its own cache into and out of this common space, keeping training cost linear rather than combinatorial as more models join the pool. A striking result anchors the discussion: translating a model's cache through this shared space can sometimes outperform the model's own untouched cache on the same task. The conversation connects this idea to familiar concepts like prefix-tuning and continuous latent reasoning, framing the cache exchange as a dynamic, evolving version of a static soft prompt. Listeners interested in how models might one day share "trains of thought" instead of finished sentences will find the tension between the approach's architectural simplicity and its surprising performance gains especially compelling.

Sources:
1. Latent Space Communication via K-V Cache Alignment
https://arxiv.org/pdf/2601.06123
2. Prefix-Tuning: Optimizing Continuous Prompts for Generation — Xiang Lisa Li, Percy Liang, 2021
https://scholar.google.com/scholar?q=Prefix-Tuning%3A+Optimizing+Continuous+Prompts+for+Generation
3. The Power of Scale for Parameter-Efficient Prompt Tuning — Brian Lester, Rami Al-Rfou, Noah Constant, 2021
https://scholar.google.com/scholar?q=The+Power+of+Scale+for+Parameter-Efficient+Prompt+Tuning
4. Relative Representations Enable Zero-Shot Latent Space Communication — Luca Moschella, Valentino Maiorca, Marco Fumero, Antonio Norelli, Francesco Locatello, Emanuele Rodolà, 2022
https://scholar.google.com/scholar?q=Relative+Representations+Enable+Zero-Shot+Latent+Space+Communication
5. Git Re-Basin: Merging Models modulo Permutation Symmetries — Samuel K. Ainsworth, Jonathan Hayase, Siddhartha Srinivasa, 2022
https://scholar.google.com/scholar?q=Git+Re-Basin%3A+Merging+Models+modulo+Permutation+Symmetries
6. On the direct alignment of latent spaces — Lähner, Moeller, 2024
https://scholar.google.com/scholar?q=On+the+direct+alignment+of+latent+spaces
7. Harnessing the universal geometry of embeddings — Jha, Zhang, Shmatikov, Morris, 2025
https://scholar.google.com/scholar?q=Harnessing+the+universal+geometry+of+embeddings
8. Training Large Language Models to Reason in a Continuous Latent Space (Coconut) — Hao, Sukhbaatar, Su, Li, Hu, Weston, Tian, 2024
https://scholar.google.com/scholar?q=Training+Large+Language+Models+to+Reason+in+a+Continuous+Latent+Space+%28Coconut%29
9. DiPaCo: Distributed Path Composition — Douillard, Feng, Rusu, Kuncoro, Donchev, Chhaparia, Gog, Ranzato, Shen, Szlam, 2024
https://scholar.google.com/scholar?q=DiPaCo%3A+Distributed+Path+Composition
Interactive Visualization: Latent Space Communication via K-V Cache Alignment

This episode examines CacheBridge, a paper proposing targeted fixes to a training-free method for transferring KV caches between different transformer models in multi-model routing setups. It explains why caches can't simply be handed off — differing residual widths, GQA head counts, and RoPE position encoding make one model's cache unreadable to another — and how a prior affine-mapper approach (FULL-HEADMAPPING) could swing wildly from near-native accuracy on one model pair to catastrophic collapse on another, with no way to predict which. The discussion breaks down the paper's four diagnosed failure causes, spanning head-mixing, mismatched error metrics, layer-count cost scaling, and a GPU implementation bottleneck, then covers the three corresponding repairs: HEAD-LOCAL's narrower one-to-one head mapping, ATTN-REPAIR's attention-aware calibration reweighting, and FUSED-FIT's custom kernel for building the mapper efficiently. Listeners interested in LLM serving infrastructure will find it a concrete look at diagnosing and patching a deployed technique rather than proposing a new architecture from scratch.

Sources:
1. CacheBridge: Efficient Cross-Model KV Cache Transfer — Xingyu Qu, Siyuan Lu, Zhiyu Chen, Sheng Wang, Tao Lin, 2026
http://arxiv.org/abs/2609.00891
2. Cross-model KV cache transfer in LLM families: A closed-form linear mapping for prefill reuse — Heo, T., Shafipour, R., Zhao, R., Golub, M., Kamani, M. M., Borkar, R., Chandran, M. T., Zardoshti, P., Rouhani, B. D., 2026
https://scholar.google.com/scholar?q=Cross-model+KV+cache+transfer+in+LLM+families%3A+A+closed-form+linear+mapping+for+prefill+reuse
3. Cache-to-cache: Direct semantic communication between large language models — Fu, T., Min, Z., Zhang, H., Yan, J., Dai, G., Ouyang, W., Wang, Y., 2026
https://scholar.google.com/scholar?q=Cache-to-cache%3A+Direct+semantic+communication+between+large+language+models
4. Mixture-of-translators: Translating KV caches across heterogeneous large language models — Lee, J.-w., Song, M., Oh, J., Han, S., Park, S., Jang, G., Lim, S., 2026
https://scholar.google.com/scholar?q=Mixture-of-translators%3A+Translating+KV+caches+across+heterogeneous+large+language+models
5. DroidSpeak: KV cache sharing across fine-tuned model variants — Liu, Y., Huang, Y., Yao, J., Feng, S., Gu, Z., Du, K., Li, H., Cheng, Y., Jiang, J., Lu, S., Musuvathi, M., Choukse, E., 2026
https://scholar.google.com/scholar?q=DroidSpeak%3A+KV+cache+sharing+across+fine-tuned+model+variants
6. ICaRus: Identical cache reuse for efficient multi model inference — Woo, S., Kil, J., Kim, H., Kim, M., Kim, J., Seo, A., Lee, S., Jo, M., Ryu, J., Park, B., Kwon, S. J., Lee, D., 2026
https://scholar.google.com/scholar?q=ICaRus%3A+Identical+cache+reuse+for+efficient+multi+model+inference
7. GQA: Training generalized multi-query transformer models from multi-head checkpoints — Ainslie, J., Lee-Thorp, J., de Jong, M., Zemlyanskiy, Y., Lebrón, F., Sanghai, S., 2023
https://scholar.google.com/scholar?q=GQA%3A+Training+generalized+multi-query+transformer+models+from+multi-head+checkpoints
Interactive Visualization: CacheBridge: Fixing Cross-Model KV Cache Transfer Failures

This episode examines a NVIDIA paper on transferring KV cache between different-sized models within the same architecture family — for example Qwen3 14B and 32B — without any gradient training. The hosts explain the core finding: a single layer of a smaller model's cache can explain over half the variance in a larger model's keys, and stacking source layers pushes that correlation even higher. They break down the two practical payoffs — using a small model's cache to bootstrap a larger model mid-conversation for quality upgrades, and the reverse direction, prefilling once on an expensive large model then handing the cache down to a cheap model to skip decode costs entirely. A skeptical exchange probes whether a closed-form ridge-regression mapping can really generalize across the nonlinear depth of transformer layers, with the paper's authors measuring rather than assuming the linear structure holds, and only for "matched-KV pairs" with identical head counts and per-head dimensions. Listeners interested in inference cost reduction, model routing, and cache reuse across model families will find the comparison to trained alternatives like Cache-to-cache and LatentAlign particularly relevant.

Sources:
1. Cross-Model KV Cache Transfer in LLM Families: A Closed-Form Linear Mapping for Prefill Reuse — Taekyung Heo, Rasoul Shafipour, Ritchie Zhao, Maximilian Golub, Mohammad Mahdi Kamani, Ritika Borkar, Makesh Tarun Chandran, Pantea Zardoshti, Bita Darvish Rouhani, 2026
http://arxiv.org/abs/2608.03893
2. Cache-to-cache: Direct semantic communication between large language models (C2C) — Tianyu Fu, Zihan Min, Hanling Zhang, Jichao Yan, Guohao Dai, Wanli Ouyang, Yu Wang, 2026 (ICLR)
https://scholar.google.com/scholar?q=Cache-to-cache%3A+Direct+semantic+communication+between+large+language+models+%28C2C%29
3. Latent space communication via K-V cache alignment (LatentAlign) — Lucio M. Dery, Zohar Yahav, Henry Prior, Qixuan Feng, Jiajun Shen, Arthur Szlam, 2026
https://scholar.google.com/scholar?q=Latent+space+communication+via+K-V+cache+alignment+%28LatentAlign%29
4. DroidSpeak: KV cache sharing across fine-tuned model variants — Yuhan Liu et al., 2026 (NSDI)
https://scholar.google.com/scholar?q=DroidSpeak%3A+KV+cache+sharing+across+fine-tuned+model+variants
5. The Platonic Representation Hypothesis — Minyoung Huh, Brian Cheung, Tongzhou Wang, Phillip Isola, 2024 (ICML)
https://scholar.google.com/scholar?q=The+Platonic+Representation+Hypothesis
6. Nvidia Nemotron 3: Efficient and open intelligence — Aaron Blakeman et al., 2025
https://scholar.google.com/scholar?q=Nvidia+Nemotron+3%3A+Efficient+and+open+intelligence
Interactive Visualization: Cross-Model KV Cache Transfer for Fast LLM Prefill Reuse

This episode examines Semantic Cache Distillation, a technique for reusing KV caches across producer and consumer transformers that share architecture but have different fine-tuned weights. The discussion covers why prefill-decode disaggregation splits compute-bound and memory-bandwidth-bound phases across separate machines, and how naively shipping raw or compressed KV caches between differently-weighted models causes "semantic drift" — a small per-layer mismatch that compounds through deep residual networks and degrades generation quality. The hosts unpack the paper's REUSE mechanism, which uses paired producer-consumer KV traces and low-rank SVD factorization to build a shared latent code, letting a lightweight encoder-decoder pair reconstruct usable cache states instead of forcing a full recompute. Real-world motivations include LoRA-adapter fleets sharing a base model and draft-verifier pairs in speculative decoding. Listeners interested in LLM serving infrastructure will find the reported 2.65x time-to-first-token speedup, and the underlying cross-model cache reconstruction problem, a concrete look at an underexplored bottleneck in production inference systems.

Sources:
1. Semantic Cache Distillation: Efficient State Transfer via Reuse and Selective Patching — Qianli Ma, Zhiqing Tang, Hanshuai Cui, Zhi Yao, Weijia Jia, 2026
http://arxiv.org/abs/2606.07684
2. S-LoRA: Serving Thousands of Concurrent LoRA Adapters — Ying Sheng, Shiyi Cao, Dacheng Li, Coleman Hooper, Nicholas Lee, Shuo Yang, Christopher Chou, Banghua Zhu, Lianmin Zheng, Kurt Keutzer, Joseph E. Gonzalez, Ion Stoica, 2024
https://scholar.google.com/scholar?q=S-LoRA%3A+Serving+Thousands+of+Concurrent+LoRA+Adapters
3. CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving — Yuhan Liu, Hanchen Li, Yihua Cheng, Siddhant Ray, Yuyang Huang, Qizheng Zhang, Kuntai Du, Jiayi Yao, Shan Lu, Ganesh Ananthanarayanan, Michael Maire, Henry Hoffmann, Ari Holtzman, Junchen Jiang, 2024
https://scholar.google.com/scholar?q=CacheGen%3A+KV+Cache+Compression+and+Streaming+for+Fast+Large+Language+Model+Serving
4. CacheBlend: Fast Large Language Model Serving with Cached Knowledge Fusion — Jiayi Yao, Hanchen Li, Yuhan Liu, Siddhant Ray, Yihua Cheng, Qizheng Zhang, Kuntai Du, Shan Lu, Junchen Jiang, 2024
https://scholar.google.com/scholar?q=CacheBlend%3A+Fast+Large+Language+Model+Serving+with+Cached+Knowledge+Fusion
5. Fast Inference from Transformers via Speculative Decoding — Yaniv Leviathan, Matan Kalman, Yossi Matias, 2023
https://scholar.google.com/scholar?q=Fast+Inference+from+Transformers+via+Speculative+Decoding
6. DroidSpeak: KV cache sharing for cross-LLM communication and multi-LLM serving — Liu, Y., Huang, Y., Yao, J., Feng, S., Gu, Z., Du, K., Li, H., Cheng, Y., Jiang, J., Lu, S., et al., 2024
https://scholar.google.com/scholar?q=DroidSpeak%3A+KV+cache+sharing+for+cross-LLM+communication+and+multi-LLM+serving
7. Cache-to-Cache: Direct semantic communication between large language models — Fu, T., Min, Z., Zhang, H., Yan, J., Dai, G., Ouyang, W., and Wang, Y., 2026
https://scholar.google.com/scholar?q=Cache-to-Cache%3A+Direct+semantic+communication+between+large+language+models
8. S-LoRA: Scalable serving of thousands of LoRA adapters — Sheng, Y., Cao, S., Li, D., Hooper, C., Lee, N., Yang, S., Chou, C., Zhu, B., Zheng, L., Keutzer, K., et al., 2024
https://scholar.google.com/scholar?q=S-LoRA%3A+Scalable+serving+of+thousands+of+LoRA+adapters
9. EAGLE: Speculative sampling requires rethinking feature uncertainty — Li, Y., Wei, F., Zhang, C., and Zhang, H., 2024
https://scholar.google.com/scholar?q=EAGLE%3A+Speculative+sampling+requires+rethinking+feature+uncertainty
10. Mooncake: A KVCache-centric disaggregated architecture for LLM serving — Qin, R., Li, Z., He, W., Zhang, M., Wu, Y., Zheng, W., and Xu, X., 2024
https://scholar.google.com/scholar?q=Mooncake%3A+A+KVCache-centric+disaggregated+architecture+for+LLM+serving
Interactive Visualization: Semantic Cache Distillation: Solving Semantic Drift in KV Cache Transfer

This episode examines "An Imitative Reinforcement Learning Framework for Pursuit-Lock-Launch Missions," which trains an unmanned combat aerial vehicle to complete a three-stage dogfighting task by blending TD3-style actor-critic reinforcement learning with a behavior-cloning term drawn from expert trajectories generated in the authors' own simulator. The discussion covers why sparse-reward, multistage combat tasks make good RL benchmarks despite the setting, the classic tradeoffs between pure reinforcement learning (sample inefficiency) and pure imitation learning (compounding error and drifting off-distribution), and how combining both aims to get faster, more reliable learning than either alone. It also flags a notable gap in the paper: it never benchmarks against DAgger, the standard fix for imitation learning's distribution-shift problem, raising open questions about whether the reported near-100% success rate reflects a genuinely better architecture or simply a weak baseline comparison. Listeners interested in robotics, RL/imitation-learning hybrids, or how combat-style testbeds get used for general control research will find the critique of the experimental design as engaging as the headline results.

Sources:
1. An Imitative Reinforcement Learning Framework for Pursuit-Lock-Launch Missions — Siyuan Li, Rongchang Zuo, Bofei Liu, Yaoyu He, Peng Liu, Yingnan Zhao, 2024
http://arxiv.org/abs/2406.11562
2. A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning — Stéphane Ross, Geoffrey Gordon, Drew Bagnell, 2011
https://scholar.google.com/scholar?q=A+Reduction+of+Imitation+Learning+and+Structured+Prediction+to+No-Regret+Online+Learning
3. Generative Adversarial Imitation Learning — Jonathan Ho, Stefano Ermon, 2016
https://scholar.google.com/scholar?q=Generative+Adversarial+Imitation+Learning
4. End to End Learning for Self-Driving Cars — Mariusz Bojarski et al. (NVIDIA), 2016
https://scholar.google.com/scholar?q=End+to+End+Learning+for+Self-Driving+Cars
5. Autonomous Air Combat Maneuvering Decision Making with Deep Reinforcement Learning — Wang et al. (multiple independent groups have published under similar titles), 2019
https://scholar.google.com/scholar?q=Autonomous+Air+Combat+Maneuvering+Decision+Making+with+Deep+Reinforcement+Learning
6. Alpha Dogfight Trials public results and analysis (DARPA program summaries) — DARPA / Heron Systems and other competing teams, 2020
https://scholar.google.com/scholar?q=Alpha+Dogfight+Trials+public+results+and+analysis+%28DARPA+program+summaries%29
7. Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model — Julian Schrittwieser et al. (DeepMind), 2020
https://scholar.google.com/scholar?q=Mastering+Atari%2C+Go%2C+Chess+and+Shogi+by+Planning+with+a+Learned+Model
8. Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor — Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, Sergey Levine, 2018
https://scholar.google.com/scholar?q=Soft+Actor-Critic%3A+Off-Policy+Maximum+Entropy+Deep+Reinforcement+Learning+with+a+Stochastic+Actor
9. Deep Q-learning from Demonstrations — Todd Hester et al. (DeepMind), 2018
https://scholar.google.com/scholar?q=Deep+Q-learning+from+Demonstrations
10. A Minimalist Approach to Offline Reinforcement Learning (TD3+BC) — Scott Fujimoto, Shixiang Shane Gu, 2021
https://scholar.google.com/scholar?q=A+Minimalist+Approach+to+Offline+Reinforcement+Learning+%28TD3%2BBC%29
11. A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning (DAgger) — Stéphane Ross, Geoffrey Gordon, Drew Bagnell, 2011
https://scholar.google.com/scholar?q=A+Reduction+of+Imitation+Learning+and+Structured+Prediction+to+No-Regret+Online+Learning+%28DAgger%29
12. Discovering Expert-Level Air Combat Knowledge via Deep Excitatory-Inhibitory Factorized Reinforcement Learning — Hai Yin Piao, Shengqi Yang, Hechang Chen, et al., 2024
https://scholar.google.com/scholar?q=Discovering+Expert-Level+Air+Combat+Knowledge+via+Deep+Excitatory-Inhibitory+Factorized+Reinforcement+Learning
13. Multi-Dimensional Decision-Making for UAV Air Combat Based on Hierarchical Reinforcement Learning — Jiandong Zhang, Dinghan Wang, Qiming Yang, et al., 2023
https://scholar.google.com/scholar?q=Multi-Dimensional+Decision-Making+for+UAV+Air+Combat+Based+on+Hierarchical+Reinforcement+Learning
Interactive Visualization: Imitative Reinforcement Learning for UCAV Pursuit-Lock-Launch Dogfights

This episode explores "Deep Drone Acrobatics," which trains a quadrotor to fly extreme maneuvers — a Power Loop, Barrel Roll, and Matty Flip — using only an onboard camera and IMU, with no external motion capture. The discussion centers on how the policy is trained entirely in simulation via DAgger imitation learning, where a privileged model-predictive controller with perfect ground-truth state acts as an expert that a vision-limited student imitates, rather than through reinforcement learning or reward shaping. A key focus is the sim-to-real gap: at high accelerations, motion blur degrades vision-based state estimation, so the paper's "input abstraction" approach feeds the network geometry-based feature tracks instead of raw pixels, drawing on prior work showing that shared abstractions between simulated and real observations shrink the performance gap. The conversation also traces the paper's intellectual lineage, connecting it to "Does Computer Vision Matter for Action?" and "Learning by Cheating," while highlighting why acrobatic flight is a harder version of the sim-to-real problem than driving, since a flipping drone has no margin for hesitation. Listeners interested in robotics, sim-to-real transfer, or imitation learning will find a concrete, technically grounded case study of zero-shot policy transfer under extreme physical constraints.

Sources:
1. Deep Drone Acrobatics: Vision-Only Zero-Shot Sim-to-Real Flight
https://roboticsproceedings.org/rss16/p040.pdf
2. Learning by Cheating — Dian Chen, Brady Zhou, Vladlen Koltun, Philipp Krähenbühl, 2019
https://scholar.google.com/scholar?q=Learning+by+Cheating
3. Deep Drone Racing: From Simulation to Reality with Domain Randomization — Antonio Loquercio, Elia Kaufmann, René Ranftl, Alexey Dosovitskiy, Vladlen Koltun, Davide Scaramuzza, 2020
https://scholar.google.com/scholar?q=Deep+Drone+Racing%3A+From+Simulation+to+Reality+with+Domain+Randomization
4. Does computer vision matter for action? — Brady Zhou, Philipp Krähenbühl, Vladlen Koltun, 2019
https://scholar.google.com/scholar?q=Does+computer+vision+matter+for+action%3F
5. Driving Policy Transfer via Modularity and Abstraction — Matthias Müller, Alexey Dosovitskiy, Bernard Ghanem, Vladlen Koltun, 2018
https://scholar.google.com/scholar?q=Driving+Policy+Transfer+via+Modularity+and+Abstraction
6. Agile Autonomous Driving Using End-to-End Deep Imitation Learning — Yunpeng Pan, Ching-An Cheng, Kamil Saigol, Keuntaek Lee, Xinyan Yan, Evangelos Theodorou, Byron Boots, 2018
https://scholar.google.com/scholar?q=Agile+Autonomous+Driving+Using+End-to-End+Deep+Imitation+Learning
7. A reduction of imitation learning and structured prediction to no-regret online learning — Stéphane Ross, Geoffrey Gordon, Drew Bagnell, 2011
https://scholar.google.com/scholar?q=A+reduction+of+imitation+learning+and+structured+prediction+to+no-regret+online+learning
Interactive Visualization: Deep Drone Acrobatics: Vision-Only Zero-Shot Sim-to-Real Flight

This episode examines Programmatically Interpretable Reinforcement Learning (PIRL), a 2018 framework from Rice University, Google Brain, and DeepMind researchers that forces RL policies to be expressed as short, human-readable programs rather than opaque neural network weights. The discussion centers on why formal verification—proving properties like bounded steering output in a self-driving car—is tractable for small domain-specific programs but essentially impossible for networks with millions of parameters. Using the paper's driving example, the hosts unpack "policy sketches" (a switch statement branching on track position, with PID controllers filling each branch) and Neurally Directed Program Search (NDPS), which trains a conventional deep RL policy as an oracle and then searches program space to imitate its outputs via smooth regression rather than fighting a jagged, non-differentiable reward landscape. They draw out the connection to DAgger's iterative imitation-learning approach from Ross, Gordon, and Bagnell, while flagging a subtle mismatch between matching an expert's actions and matching reward through an imitation proxy. Listeners interested in AI safety, control theory, or the tension between interpretability and performance will find the concrete TORCS driving case a clear entry point into verifiable reinforcement learning.

Sources:
1. Programmatically Interpretable Reinforcement Learning — Abhinav Verma, Vijayaraghavan Murali, Rishabh Singh, Pushmeet Kohli, Swarat Chaudhuri, 2018
http://arxiv.org/abs/1804.02477
2. A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning — Stéphane Ross, Geoffrey Gordon, J. Andrew Bagnell, 2011
https://scholar.google.com/scholar?q=A+Reduction+of+Imitation+Learning+and+Structured+Prediction+to+No-Regret+Online+Learning
3. Is Imitation Learning the Route to Humanoid Robots? — Stefan Schaal, 1999
https://scholar.google.com/scholar?q=Is+Imitation+Learning+the+Route+to+Humanoid+Robots%3F
4. ALVINN: An Autonomous Land Vehicle in a Neural Network — Dean Pomerleau, 1989
https://scholar.google.com/scholar?q=ALVINN%3A+An+Autonomous+Land+Vehicle+in+a+Neural+Network
5. Generative Adversarial Imitation Learning — Jonathan Ho, Stefano Ermon, 2016
https://scholar.google.com/scholar?q=Generative+Adversarial+Imitation+Learning
6. Verifiable Reinforcement Learning via Policy Extraction — Osbert Bastani, Yewen Pu, Armando Solar-Lezama, 2018
https://scholar.google.com/scholar?q=Verifiable+Reinforcement+Learning+via+Policy+Extraction
7. Programmatically Interpretable Reinforcement Learning — Abhinav Verma, Vijayaraghavan Murali, Rishabh Singh, Pushmeet Kohli, Swarat Chaudhuri, 2018
https://scholar.google.com/scholar?q=Programmatically+Interpretable+Reinforcement+Learning
8. Optimization Methods for Interpretable Differentiable Decision Trees Applied to Reinforcement Learning — Andrew Silva, Matthew Gombolay, Taylor Killian, Ivan Jimenez, Sung-Hyun Son, 2020
https://scholar.google.com/scholar?q=Optimization+Methods+for+Interpretable+Differentiable+Decision+Trees+Applied+to+Reinforcement+Learning
9. Distilling a Neural Network Into a Soft Decision Tree — Nicholas Frosst, Geoffrey Hinton, 2017
https://scholar.google.com/scholar?q=Distilling+a+Neural+Network+Into+a+Soft+Decision+Tree
10. A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning (DAgger) — Stéphane Ross, Geoffrey J. Gordon, J. Andrew Bagnell, 2011
https://scholar.google.com/scholar?q=A+Reduction+of+Imitation+Learning+and+Structured+Prediction+to+No-Regret+Online+Learning+%28DAgger%29
11. Continuous Control with Deep Reinforcement Learning (DDPG) — Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, Daan Wierstra, 2015
https://scholar.google.com/scholar?q=Continuous+Control+with+Deep+Reinforcement+Learning+%28DDPG%29
12. Reluplex: An Efficient SMT Solver for Verifying Deep Neural Networks — Guy Katz, Clark Barrett, David L. Dill, Kyle Julian, Mykel J. Kochenderfer, 2017
https://scholar.google.com/scholar?q=Reluplex%3A+An+Efficient+SMT+Solver+for+Verifying+Deep+Neural+Networks
13. The Sketching Approach to Program Synthesis — Armando Solar-Lezama, 2009
https://scholar.google.com/scholar?q=The+Sketching+Approach+to+Program+Synthesis
14. Syntax-Guided Synthesis (SyGuS) — Rajeev Alur, Rastislav Bodík, Eric Dallal, Dana Fisman, Pranav Garg, Ghila Juniwal, Hadas Kress-Gazit, P. Madhusudan, Milo M. K. Martin, Mukund Raghothaman, Shambwaditya Saha, Sanjit A. Seshia, Rishabh Singh, Armando Solar-Lezama, Emina Torlak, Abhishek Udupa, 2015
https://scholar.google.com/scholar?q=Syntax-Guided+Synthesis+%28SyGuS%29
Interactive Visualization: Programmatically Interpretable Reinforcement Learning: Readable Policies

This episode examines "Hierarchical Reinforcement Learning for Air Combat at DARPA's AlphaDogfight Trials," in which the PHANG-MAN agent swept a graduate of the USAF Weapons Instructor Course 5-0 in simulated dogfighting. The discussion traces the DARPA ACE program's rationale for building trust incrementally toward AI-assisted piloted aircraft, and contrasts this work with prior systems like Nick Ernest's genetic fuzzy tree ALPHA, highlighting how PHANG-MAN operates with genuinely continuous stick-and-rudder control in the high-fidelity JSBSim F-16 simulator rather than a maneuver library. It unpacks the two-layer architecture — three frozen, independently-trained low-level Soft Actor-Critic specialist policies (Control Zone, Aggressive Shooter, Conservative Shooter) governed by a higher-frequency policy selector — and explains supporting concepts like curriculum learning and maximum-entropy RL that make the training tractable. Listeners interested in reinforcement learning architecture, autonomous systems trust-building, or the gap between simulated and real-world control will find the technical breakdown of temporally-extended specialist routing especially compelling.

Sources:
1. Hierarchical Reinforcement Learning for Air Combat at DARPA's AlphaDogfight Trials — Adrian P. Pope, Jaime S. Ide, Daria Micovic, Henry Diaz, David Rosenbluth, Lee Ritholtz, Jason C. Twedt, Thayne T. Walker, Kevin Alcedo, Daniel Javorsek, 2021
http://arxiv.org/abs/2105.00990
2. Feudal Reinforcement Learning — Peter Dayan, Geoffrey Hinton, 1993
https://scholar.google.com/scholar?q=Feudal+Reinforcement+Learning
3. Hierarchical Reinforcement Learning with the MAXQ Value Function Decomposition — Thomas G. Dietterich, 2000
https://scholar.google.com/scholar?q=Hierarchical+Reinforcement+Learning+with+the+MAXQ+Value+Function+Decomposition
4. FeUdal Networks for Hierarchical Reinforcement Learning — Alexander Sasha Vezhnevets, Simon Osindero, Tom Schaul, Nicolas Heess, Max Jaderberg, David Silver, Koray Kavukcuoglu, 2017
https://scholar.google.com/scholar?q=FeUdal+Networks+for+Hierarchical+Reinforcement+Learning
5. The Option-Critic Architecture — Pierre-Luc Bacon, Jean Harb, Doina Precup, 2017
https://scholar.google.com/scholar?q=The+Option-Critic+Architecture
6. Curriculum Learning — Yoshua Bengio, Jerome Louradour, Ronan Collobert, Jason Weston, 2009
https://scholar.google.com/scholar?q=Curriculum+Learning
7. Reverse Curriculum Generation for Reinforcement Learning — Carlos Florensa, David Held, Markus Wulfmeier, Michael Zhang, Pieter Abbeel, 2017
https://scholar.google.com/scholar?q=Reverse+Curriculum+Generation+for+Reinforcement+Learning
8. Automatic Goal Generation for Reinforcement Learning Agents — Carlos Florensa, David Held, Xinyang Geng, Pieter Abbeel, 2018
https://scholar.google.com/scholar?q=Automatic+Goal+Generation+for+Reinforcement+Learning+Agents
9. Dota 2 with Large Scale Deep Reinforcement Learning — OpenAI (Christopher Berner, Greg Brockman, Brooke Chan, et al.), 2019
https://scholar.google.com/scholar?q=Dota+2+with+Large+Scale+Deep+Reinforcement+Learning
10. Reinforcement Learning with Deep Energy-Based Policies — Tuomas Haarnoja, Haoran Tang, Pieter Abbeel, Sergey Levine, 2017
https://scholar.google.com/scholar?q=Reinforcement+Learning+with+Deep+Energy-Based+Policies
11. Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor — Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, Sergey Levine, 2018
https://scholar.google.com/scholar?q=Soft+Actor-Critic%3A+Off-Policy+Maximum+Entropy+Deep+Reinforcement+Learning+with+a+Stochastic+Actor
12. Soft Actor-Critic Algorithms and Applications — Tuomas Haarnoja, Aurick Zhou, Kristian Hartikainen, George Tucker, Sehoon Ha, Jie Tan, Vikash Kumar, Henry Zhu, Abhishek Gupta, Pieter Abbeel, Sergey Levine, 2018
https://scholar.google.com/scholar?q=Soft+Actor-Critic+Algorithms+and+Applications
13. Maximum Entropy Inverse Reinforcement Learning — Brian D. Ziebart, Andrew Maas, J. Andrew Bagnell, Anind K. Dey, 2008
https://scholar.google.com/scholar?q=Maximum+Entropy+Inverse+Reinforcement+Learning
14. Meta Learning Shared Hierarchies — K. Frans, J. Ho, X. Chen, P. Abbeel, J. Schulman, 2018
https://scholar.google.com/scholar?q=Meta+Learning+Shared+Hierarchies
15. Data-Efficient Hierarchical Reinforcement Learning (HIRO) — O. Nachum, S. Gu, H. Lee, S. Levine, 2018
https://scholar.google.com/scholar?q=Data-Efficient+Hierarchical+Reinforcement+Learning+%28HIRO%29
16. Genetic Fuzzy based Artificial Intelligence for Unmanned Combat Aerial Vehicle Control in Simulated Air Combat Missions — N. Ernest, D. Carroll, C. Schumacher, M. Clark, K. Cohen, G. Lee, 2016
https://scholar.google.com/scholar?q=Genetic+Fuzzy+based+Artificial+Intelligence+for+Unmanned+Combat+Aerial+Vehicle+Control+in+Simulated+Air+Combat+Missions
17. Multi-agent hierarchical policy gradient for air combat tactics emergence via self-play — Z. Sun, H. Piao, Z. Yang, Y. Zhao, et al., 2021
https://scholar.google.com/scholar?q=Multi-agent+hierarchical+policy+gradient+for+air+combat+tactics+emergence+via+self-play
18. Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World — J. Tobin, R. Fong, A. Ray, J. Schneider, W. Zaremba, P. Abbeel, 2017
https://scholar.google.com/scholar?q=Domain+Randomization+for+Transferring+Deep+Neural+Networks+from+Simulation+to+the+Real+World
Interactive Visualization: Hierarchical RL Beats an F-16 Instructor Pilot 5-0

This episode examines MPC-Net, a 2019/2020 paper from ETH Zürich's Robotic Systems Lab that trains a fast neural policy to replace expensive model predictive control on the ANYmal quadruped, cutting per-step evaluation from 38 milliseconds to roughly 0.125 milliseconds using less than ten minutes of demonstration data. The discussion centers on why the method learns by minimizing the control Hamiltonian — the optimality condition MPC itself solves internally — rather than copying the expert's chosen actions, arguing this teaches the network the underlying reasoning rather than surface behavior. It contrasts this approach with classical Guided Policy Search, where the teacher adapts toward the student over training, versus MPC-Net's fixed, non-adaptive teacher that keeps solving the same optimal control problem regardless of the learner's progress. The hosts debate the tradeoffs of adaptive versus static teachers in imitation learning, weighing convergence speed against the validity and reusability of generated trajectories. Listeners interested in legged robotics, optimal control theory, or the mechanics of imitation learning will find a detailed technical walkthrough of how theory-grounded objectives can outperform standard behavioral cloning.

Sources:
1. MPC-Net: A First Principles Guided Policy Search — Jan Carius, Farbod Farshidian, Marco Hutter, 2019
http://arxiv.org/abs/1909.05197
2. ALVINN: An Autonomous Land Vehicle in a Neural Network — Dean Pomerleau, 1989
https://scholar.google.com/scholar?q=ALVINN%3A+An+Autonomous+Land+Vehicle+in+a+Neural+Network
3. A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning (DAgger) — Stéphane Ross, Geoffrey Gordon, J. Andrew Bagnell, 2011
https://scholar.google.com/scholar?q=A+Reduction+of+Imitation+Learning+and+Structured+Prediction+to+No-Regret+Online+Learning+%28DAgger%29
4. Guided Policy Search — Sergey Levine, Vladlen Koltun, 2013
https://scholar.google.com/scholar?q=Guided+Policy+Search
5. Generative Adversarial Imitation Learning — Jonathan Ho, Stefano Ermon, 2016
https://scholar.google.com/scholar?q=Generative+Adversarial+Imitation+Learning
6. Learning Agile and Dynamic Motor Skills for Legged Robots — Jemin Hwangbo, Joonho Lee, Alexey Dosovitskiy, Dario Bellicoso, Vassilios Tsounis, Vladlen Koltun, Marco Hutter, 2019
https://scholar.google.com/scholar?q=Learning+Agile+and+Dynamic+Motor+Skills+for+Legged+Robots
7. Sim-to-Real: Learning Agile Locomotion for Quadruped Robots — Jie Tan, Tingnan Zhang, Erwin Coumans, Atil Iscen, Yunfei Bai, Danijar Hafner, Steven Bohez, Vincent Vanhoucke, 2018
https://scholar.google.com/scholar?q=Sim-to-Real%3A+Learning+Agile+Locomotion+for+Quadruped+Robots
8. Learning Quadrupedal Locomotion over Challenging Terrain — Joonho Lee, Jemin Hwangbo, Lorenz Wellhausen, Vladlen Koltun, Marco Hutter, 2020
https://scholar.google.com/scholar?q=Learning+Quadrupedal+Locomotion+over+Challenging+Terrain
9. High-Slope Terrain Locomotion for Torque-Controlled Quadruped Robots (representative MPC/whole-body baseline) — Marco Hutter, Christian Gehring, and colleagues, ETH Zurich Robotic Systems Lab, 2016-2018 (various)
https://scholar.google.com/scholar?q=High-Slope+Terrain+Locomotion+for+Torque-Controlled+Quadruped+Robots+%28representative+MPC%2Fwhole-body+baseline%29
10. Adaptive mixtures of local experts — R. A. Jacobs, M. I. Jordan, S. J. Nowlan, G. E. Hinton, 1991
https://scholar.google.com/scholar?q=Adaptive+mixtures+of+local+experts
11. An efficient optimal planning and control framework for quadrupedal locomotion — F. Farshidian, M. Neunert, A. W. Winkler, G. Rey, J. Buchli, 2017
https://scholar.google.com/scholar?q=An+efficient+optimal+planning+and+control+framework+for+quadrupedal+locomotion
Interactive Visualization: MPC-Net: Learning Optimal Control via the Hamiltonian

This episode examines Guided Policy Search, a 2013 method from Sergey Levine and Vladlen Koltun that lets flexible neural-network policies control robots without falling into the poor local optima that plague direct policy search over high-dimensional parameter spaces. The discussion traces the paper's teacher-student structure: differential dynamic programming (DDP), a model-based trajectory optimizer rooted in 1960s optimal control theory, generates high-reward example trajectories for specific starting conditions, and the neural-network student learns to match and generalize this behavior via policy gradients and importance sampling rather than naive imitation. A key distinction drawn out is why this differs from imitation learning approaches like DAGGER — DDP's guidance is only locally valid, so the method needs an objective built to maximize reward everywhere, not just mimic a narrow expert trajectory. The conversation connects DDP's backward pass to Bellman recursion and the broader LQR/Kalman-filter lineage, and explains how importance sampling lets the same batch of guiding samples be reused across many gradient steps, which matters when real-hardware data collection is expensive. Listeners interested in the historical roots of modern reinforcement learning — and how classical control theory was fused with neural networks years before this became standard practice — will find the episode's walkthrough of the underlying mechanics clarifying.

Sources:
1. Guided Policy Search: Teaching Neural Nets via Trajectory Optimization
https://proceedings.mlr.press/v28/levine13.pdf
2. Learning Neural Network Policies with Guided Policy Search under Unknown Dynamics — Sergey Levine, Pieter Abbeel, 2014
https://scholar.google.com/scholar?q=Learning+Neural+Network+Policies+with+Guided+Policy+Search+under+Unknown+Dynamics
3. End-to-End Training of Deep Visuomotor Policies — Sergey Levine, Chelsea Finn, Trevor Darrell, Pieter Abbeel, 2016
https://scholar.google.com/scholar?q=End-to-End+Training+of+Deep+Visuomotor+Policies
4. Trust Region Policy Optimization — John Schulman, Sergey Levine, Philipp Moritz, Michael Jordan, Pieter Abbeel, 2015
https://scholar.google.com/scholar?q=Trust+Region+Policy+Optimization
5. Reinforcement Learning of Motor Skills with Policy Gradients — Jan Peters, Stefan Schaal, 2008
https://scholar.google.com/scholar?q=Reinforcement+Learning+of+Motor+Skills+with+Policy+Gradients
6. Differential Dynamic Programming — David Jacobson, David Mayne, 1970
https://scholar.google.com/scholar?q=Differential+Dynamic+Programming
7. A Generalized Iterative LQG Method for Locally-Optimal Feedback Control of Constrained Nonlinear Stochastic Systems — Emanuel Todorov, Weiwei Li, 2005
https://scholar.google.com/scholar?q=A+Generalized+Iterative+LQG+Method+for+Locally-Optimal+Feedback+Control+of+Constrained+Nonlinear+Stochastic+Systems
8. Synthesis and Stabilization of Complex Behaviors through Online Trajectory Optimization — Yuval Tassa, Tom Erez, Emanuel Todorov, 2012
https://scholar.google.com/scholar?q=Synthesis+and+Stabilization+of+Complex+Behaviors+through+Online+Trajectory+Optimization
9. Aggressive Driving with Model Predictive Path Integral Control — Grady Williams, Paul Drews, Brian Goldfain, James Rehg, Evangelos Theodorou, 2016
https://scholar.google.com/scholar?q=Aggressive+Driving+with+Model+Predictive+Path+Integral+Control
10. Eligibility Traces for Off-Policy Policy Evaluation — Doina Precup, Richard Sutton, Satinder Singh, 2000
https://scholar.google.com/scholar?q=Eligibility+Traces+for+Off-Policy+Policy+Evaluation
11. Learning from Scarce Experience — Leonid Peshkin, Christian Shelton, 2002
https://scholar.google.com/scholar?q=Learning+from+Scarce+Experience
12. Proximal Policy Optimization Algorithms — John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, Oleg Klimov, 2017
https://scholar.google.com/scholar?q=Proximal+Policy+Optimization+Algorithms
13. PILCO: A Model-Based and Data-Efficient Approach to Policy Search — Deisenroth, M. and Rasmussen, C., 2011
https://scholar.google.com/scholar?q=PILCO%3A+A+Model-Based+and+Data-Efficient+Approach+to+Policy+Search
14. A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning (DAGGER) — Ross, S., Gordon, G., and Bagnell, A., 2011
https://scholar.google.com/scholar?q=A+Reduction+of+Imitation+Learning+and+Structured+Prediction+to+No-Regret+Online+Learning+%28DAGGER%29
15. On a Connection Between Importance Sampling and the Likelihood Ratio Policy Gradient — Tang, J. and Abbeel, P., 2010
https://scholar.google.com/scholar?q=On+a+Connection+Between+Importance+Sampling+and+the+Likelihood+Ratio+Policy+Gradient
16. Approximately Optimal Approximate Reinforcement Learning — Kakade, S. and Langford, J., 2002
https://scholar.google.com/scholar?q=Approximately+Optimal+Approximate+Reinforcement+Learning
17. SIMBICON: Simple Biped Locomotion Control — Yin, K., Loken, K., and van de Panne, M., 2007
https://scholar.google.com/scholar?q=SIMBICON%3A+Simple+Biped+Locomotion+Control
Interactive Visualization: Guided Policy Search: Teaching Neural Nets via Trajectory Optimization

This episode examines a 2018 survey, "An Algorithmic Perspective on Imitation Learning," which frames robot skill acquisition as an alternative to brittle manual programming or fragile reward engineering. It contrasts two core approaches: behavioral cloning, which treats the problem as supervised learning but suffers from compounding errors when the policy drifts into states the expert never demonstrated, and inverse reinforcement learning, which recovers the expert's underlying reward function before solving for a policy, trading computational cost for better generalization. Concrete examples like the ALVINN self-driving system, AlphaGo's use of expert-game pretraining, and Dynamic Movement Primitives illustrate how these ideas played out in practice, with DMPs offered as a hand-structured counterpoint to fully learned neural approaches. Listeners interested in the tradeoffs between hand-designed structure and end-to-end learning, or in how robotics tackled these problems just before deep learning reshaped the field, will find the historical framing useful for understanding today's imitation-learning methods.

Sources:
1. An Algorithmic Perspective on Imitation Learning — Takayuki Osa, Joni Pajarinen, Gerhard Neumann, J. Andrew Bagnell, Pieter Abbeel, Jan Peters, 2018
http://arxiv.org/abs/1811.06711
2. A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning (DAgger) — Stéphane Ross, Geoffrey Gordon, J. Andrew Bagnell, 2011
https://scholar.google.com/scholar?q=A+Reduction+of+Imitation+Learning+and+Structured+Prediction+to+No-Regret+Online+Learning+%28DAgger%29
3. End to End Learning for Self-Driving Cars — Mariusz Bojarski, et al. (NVIDIA), 2016
https://scholar.google.com/scholar?q=End+to+End+Learning+for+Self-Driving+Cars
4. Generative Adversarial Imitation Learning (GAIL) — Jonathan Ho, Stefano Ermon, 2016
https://scholar.google.com/scholar?q=Generative+Adversarial+Imitation+Learning+%28GAIL%29
5. ALVINN: An Autonomous Land Vehicle in a Neural Network — Dean Pomerleau, 1989
https://scholar.google.com/scholar?q=ALVINN%3A+An+Autonomous+Land+Vehicle+in+a+Neural+Network
6. Diffusion Policy: Visuomotor Policy Learning via Action Diffusion — Cheng Chi, Siyuan Feng, Yilun Du, Zhenjia Xu, Eric Cousineau, Benjamin Burchfiel, Shuran Song, 2023
https://scholar.google.com/scholar?q=Diffusion+Policy%3A+Visuomotor+Policy+Learning+via+Action+Diffusion
7. Algorithms for Inverse Reinforcement Learning — Andrew Ng, Stuart Russell, 2000
https://scholar.google.com/scholar?q=Algorithms+for+Inverse+Reinforcement+Learning
8. Maximum Entropy Inverse Reinforcement Learning — Brian Ziebart, Andrew Maas, J. Andrew Bagnell, Anind Dey, 2008
https://scholar.google.com/scholar?q=Maximum+Entropy+Inverse+Reinforcement+Learning
9. Deep Reinforcement Learning from Human Preferences — Paul Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, Dario Amodei, 2017
https://scholar.google.com/scholar?q=Deep+Reinforcement+Learning+from+Human+Preferences
10. Dynamical Movement Primitives: Learning Attractor Models for Motor Behaviors — Auke Ijspeert, Jun Nakanishi, Heiko Hoffmann, Peter Pastor, Stefan Schaal, 2013 (Neural Computation; building on the authors' earlier 2002/2003 conference papers)
https://scholar.google.com/scholar?q=Dynamical+Movement+Primitives%3A+Learning+Attractor+Models+for+Motor+Behaviors
11. Probabilistic Movement Primitives — Alexandros Paraschos, Christian Daniel, Jan Peters, Gerhard Neumann, 2013
https://scholar.google.com/scholar?q=Probabilistic+Movement+Primitives
12. Learning and Generalization of Motor Skills by Learning from Demonstration — Peter Pastor, Heiko Hoffmann, Tamim Asfour, Stefan Schaal, 2009
https://scholar.google.com/scholar?q=Learning+and+Generalization+of+Motor+Skills+by+Learning+from+Demonstration
13. Generative Adversarial Imitation Learning — J. Ho, S. Ermon, 2016
https://scholar.google.com/scholar?q=Generative+Adversarial+Imitation+Learning
14. Trust Region Policy Optimization — J. Schulman, S. Levine, P. Moritz, M. Jordan, P. Abbeel, 2015
https://scholar.google.com/scholar?q=Trust+Region+Policy+Optimization
15. Cooperative Inverse Reinforcement Learning — D. Hadfield-Menell, S. J. Russell, P. Abbeel, A. Dragan, 2016
https://scholar.google.com/scholar?q=Cooperative+Inverse+Reinforcement+Learning
16. Time-Contrastive Networks: Self-Supervised Learning from Video — P. Sermanet, C. Lynch, Y. Chebotar, J. Hsu, E. Jang, S. Schaal, S. Levine, 2017
https://scholar.google.com/scholar?q=Time-Contrastive+Networks%3A+Self-Supervised+Learning+from+Video
17. Deep Q-learning from Demonstrations / Sun et al. on sample-complexity of imitation vs RL — W. Sun, A. Venkatraman, G. Gordon, B. Boots, J. A. Bagnell, 2017
https://scholar.google.com/scholar?q=Deep+Q-learning+from+Demonstrations+%2F+Sun+et+al.+on+sample-complexity+of+imitation+vs+RL
Interactive Visualization: An Algorithmic Perspective on Imitation Learning

This episode examines "A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning" (Ross, Gordon, and Bagnell, AISTATS 2011), unpacking why naively training a policy via supervised learning on expert demonstrations produces errors that compound quadratically with task horizon rather than linearly. It traces this compounding-error problem through concrete analogues — a self-driving policy drifting off the expert's trajectory into unseen states, and the exposure bias later rediscovered in sequence-to-sequence language models trained with teacher forcing — showing the same closed-loop failure mode recurring across robotics, structured prediction (like part-of-speech tagging), and NLP. It also revisits the authors' own earlier fix, SMILe (2010), explaining why its stochastic mixture-of-policies approach required an impractically large number of iterations to approach linear regret. Listeners interested in the theoretical foundations connecting imitation learning, sequence generation, and online learning reductions will find this a clear walkthrough of a foundational result that anticipated problems later rediscovered independently in deep learning.

Sources:
1. A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning — Stephane Ross, Geoffrey J. Gordon, J. Andrew Bagnell, 2010
http://arxiv.org/abs/1011.0686
2. Efficient Reductions for Imitation Learning — Stéphane Ross, J. Andrew Bagnell, 2010
https://scholar.google.com/scholar?q=Efficient+Reductions+for+Imitation+Learning
3. Search-based Structured Prediction (SEARN) — Hal Daumé III, John Langford, Daniel Marcu, 2009
https://scholar.google.com/scholar?q=Search-based+Structured+Prediction+%28SEARN%29
4. Apprenticeship Learning via Inverse Reinforcement Learning — Pieter Abbeel, Andrew Y. Ng, 2004
https://scholar.google.com/scholar?q=Apprenticeship+Learning+via+Inverse+Reinforcement+Learning
5. Generative Adversarial Imitation Learning — Jonathan Ho, Stefano Ermon, 2016
https://scholar.google.com/scholar?q=Generative+Adversarial+Imitation+Learning
6. Conditional Random Fields: Probabilistic Models for Segmenting and Labeling Sequence Data — John Lafferty, Andrew McCallum, Fernando Pereira, 2001
https://scholar.google.com/scholar?q=Conditional+Random+Fields%3A+Probabilistic+Models+for+Segmenting+and+Labeling+Sequence+Data
7. Scheduled Sampling for Sequence Prediction with Recurrent Neural Networks — Samy Bengio, Oriol Vinyals, Navdeep Jaitly, Noam Shazeer, 2015
https://scholar.google.com/scholar?q=Scheduled+Sampling+for+Sequence+Prediction+with+Recurrent+Neural+Networks
8. Approximately Optimal Approximate Reinforcement Learning — Sham Kakade, John Langford, 2002
https://scholar.google.com/scholar?q=Approximately+Optimal+Approximate+Reinforcement+Learning
9. Error-Correcting Tournaments / Error Limiting Reductions — Alina Beygelzimer, Varsha Dani, Tom Hayes, John Langford, Bianca Zadrozny (various papers in this line, ~2005), 2005
https://scholar.google.com/scholar?q=Error-Correcting+Tournaments+%2F+Error+Limiting+Reductions
10. Relating Reinforcement Learning Performance to Classification Performance — John Langford, Bianca Zadrozny, 2005
https://scholar.google.com/scholar?q=Relating+Reinforcement+Learning+Performance+to+Classification+Performance
11. Efficient Reductions for Imitation Learning (SMILe, Forward Training) — Stéphane Ross, J. Andrew Bagnell, 2010
https://scholar.google.com/scholar?q=Efficient+Reductions+for+Imitation+Learning+%28SMILe%2C+Forward+Training%29
12. Approximately Optimal Approximate Reinforcement Learning (CPI) — Sham Kakade, John Langford, 2002
https://scholar.google.com/scholar?q=Approximately+Optimal+Approximate+Reinforcement+Learning+%28CPI%29
13. Error Limiting Reductions Between Classification Tasks — Alina Beygelzimer, Varsha Dani, Tom Hayes, John Langford, Bianca Zadrozny, 2005
https://scholar.google.com/scholar?q=Error+Limiting+Reductions+Between+Classification+Tasks
14. Max-Margin Markov Networks — Ben Taskar, Carlos Guestrin, Daphne Koller, 2003
https://scholar.google.com/scholar?q=Max-Margin+Markov+Networks
15. On the Generalization Ability of Online Strongly Convex Programming Algorithms — Sham Kakade, Ambuj Tewari, 2009
https://scholar.google.com/scholar?q=On+the+Generalization+Ability+of+Online+Strongly+Convex+Programming+Algorithms
Interactive Visualization: Reduction of Imitation Learning to No-Regret Online Learning

This episode examines Piper, a training system from Oak Ridge National Laboratory designed to fix catastrophic GPU underutilization in large-scale Mixture-of-Experts training, where the leading framework X-MoE hits only about 5% utilization on a 545-billion-parameter model. The discussion traces MoE's evolution from GShard and Switch Transformer's coarse-grained experts to DeepSeek-MoE's fine-grained approach with hundreds of small experts, and explains why expert parallelism's all-to-all communication becomes a severe bottleneck on Frontier's Dragonfly network topology, where bandwidth varies sharply with GPU distance. Piper's core innovation is repurposing pipeline parallelism, normally used only to split layers across dense models, to also confine expensive expert-parallel communication within small, physically local GPU groups arranged in a pipeline-by-expert-parallel grid. The conversation details how an analytical resource model prunes infeasible configurations for memory and communication cost before a micro-benchmarking pass measures real hardware throughput to select the optimal setup, claiming a two-to-three-and-a-half-times utilization improvement. Listeners interested in the practical gap between theoretical FLOPs and real supercomputer throughput will find this a concrete look at what it takes to make trillion-parameter training economically viable on shared HPC infrastructure rather than purpose-built AI clusters.

Sources:
1. Piper: Efficient Large-Scale MoE Training via Resource Modeling and Pipelined Hybrid Parallelism — Sajal Dash, Feiyi Wang, 2026
http://arxiv.org/abs/2605.05049
2. GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding — Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, Zhifeng Chen, 2020
https://scholar.google.com/scholar?q=GShard%3A+Scaling+Giant+Models+with+Conditional+Computation+and+Automatic+Sharding
3. Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity — William Fedus, Barret Zoph, Noam Shazeer, 2021
https://scholar.google.com/scholar?q=Switch+Transformers%3A+Scaling+to+Trillion+Parameter+Models+with+Simple+and+Efficient+Sparsity
4. Tutel: Adaptive Mixture-of-Experts at Scale — Changho Hwang, Wei Cui, Yifan Xiong, Ziyue Yang, Ze Liu, Han Hu, Zilong Wang, Rafael Salas, Jithin Jose, Prabhat Ram, Joe Chau, Peng Cheng, Fan Yang, Mao Yang, Yongqiang Xiong, 2022 (updated 2023)
https://scholar.google.com/scholar?q=Tutel%3A+Adaptive+Mixture-of-Experts+at+Scale
5. FasterMoE: Modeling and Optimizing Training of Large-Scale Dynamic Pre-Trained Models — Jiaao He, Jiezhong Qiu, Aohan Zeng, Zhilin Yang, Jidong Zhai, Jie Tang, 2022
https://scholar.google.com/scholar?q=FasterMoE%3A+Modeling+and+Optimizing+Training+of+Large-Scale+Dynamic+Pre-Trained+Models
6. FlexMoE: Scaling Large-scale Sparse Pre-trained Model Training via Dynamic Device Placement — Xiaonan Nie, Xupeng Miao, Zilong Wang, Zichao Yang, Jilong Xue, Lingxiao Ma, Gang Cao, Bin Cui, 2023 (SIGMOD/VLDB)
https://scholar.google.com/scholar?q=FlexMoE%3A+Scaling+Large-scale+Sparse+Pre-trained+Model+Training+via+Dynamic+Device+Placement
7. SmartMoE: Efficiently Training Sparsely-Activated Models through Combining Offline and Online Parallelization — Mingshu Zhai, Jiaao He, Zixuan Ma, Zan Zong, Runqing Zhang, Jidong Zhai, 2023 (USENIX ATC)
https://scholar.google.com/scholar?q=SmartMoE%3A+Efficiently+Training+Sparsely-Activated+Models+through+Combining+Offline+and+Online+Parallelization
8. Lina: Enabling Sparse-Aware Distributed Training via Efficient Communication Scheduling — Jiamin Li, Yimin Jiang, Yibo Zhu, Cong Wang, Hong Xu, 2023 (SIGCOMM/ASPLOS)
https://scholar.google.com/scholar?q=Lina%3A+Enabling+Sparse-Aware+Distributed+Training+via+Efficient+Communication+Scheduling
9. ZeRO-Infinity: Breaking the GPU Memory Wall for Extreme Scale Deep Learning — Samyam Rajbhandari, Olatunji Ruwase, Jeff Rasley, Shaden Smith, Yuxiong He, 2021 (SC)
https://scholar.google.com/scholar?q=ZeRO-Infinity%3A+Breaking+the+GPU+Memory+Wall+for+Extreme+Scale+Deep+Learning
Interactive Visualization: Piper: Fixing GPU Utilization in MoE Training on Frontier

This episode explores MoM (Mixture-of-Memories), a linear sequence modeling architecture from researchers at Shanghai AI Laboratory and collaborating universities that tackles a core weakness in efficient Transformer alternatives: their tendency to forget information from earlier in a sequence. The discussion traces the trade-off at the heart of the field — Transformers preserve every token via a growing key-value cache at quadratic cost, while linear models like Mamba and RWKV compress everything into a single fixed-size memory state, trading recall precision for constant-time efficiency. MoM's proposed fix draws on two distinct sources: a neuroscience-inspired analogy to how the hippocampus uses separate oscillatory channels to keep simultaneous memories from blending together, and the Mixture-of-Experts routing mechanism, applied here to memory states rather than feed-forward layers. The result is an architecture with multiple independent memory slots plus a shared accumulating memory, with a lightweight router directing each token to the appropriate slot. Listeners interested in efficient sequence modeling, long-context recall, or the cross-pollination between neuroscience and deep learning architecture design will find the mechanics of this capacity-versus-interference problem — and its proposed solution — a compelling deep dive.

Sources:
1. MoM: Linear Sequence Modeling with Mixture-of-Memories — Jusen Du, Weigao Sun, Disen Lan, Jiaxi Hu, Yu Cheng, 2025
http://arxiv.org/abs/2502.13685v4
2. Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention — Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François Fleuret, 2020
https://scholar.google.com/scholar?q=Transformers+are+RNNs%3A+Fast+Autoregressive+Transformers+with+Linear+Attention
3. Mamba: Linear-Time Sequence Modeling with Selective State Spaces — Albert Gu, Tri Dao, 2023 (updated 2024)
https://scholar.google.com/scholar?q=Mamba%3A+Linear-Time+Sequence+Modeling+with+Selective+State+Spaces
4. Gated Linear Attention Transformers with Hardware-Efficient Training — Songlin Yang, Bailin Wang, Yikang Shen, Rameswar Panda, Yoon Kim, 2023
https://scholar.google.com/scholar?q=Gated+Linear+Attention+Transformers+with+Hardware-Efficient+Training
5. Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer — Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, Jeff Dean, 2017
https://scholar.google.com/scholar?q=Outrageously+Large+Neural+Networks%3A+The+Sparsely-Gated+Mixture-of-Experts+Layer
6. Gated Delta Networks: Improving Mamba2 with Delta Rule — Songlin Yang, Jan Kautz, Ali Hatamizadeh, 2024
https://scholar.google.com/scholar?q=Gated+Delta+Networks%3A+Improving+Mamba2+with+Delta+Rule
7. Titans: Learning to Memorize at Test Time — Ali Behrouz, Peilin Zhong, Vahab Mirrokni, 2024
https://scholar.google.com/scholar?q=Titans%3A+Learning+to+Memorize+at+Test+Time
8. Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity — William Fedus, Barret Zoph, Noam Shazeer, 2022
https://scholar.google.com/scholar?q=Switch+Transformers%3A+Scaling+to+Trillion+Parameter+Models+with+Simple+and+Efficient+Sparsity
9. Simple Linear Attention Language Models Balance the Recall-Throughput Tradeoff (BASED) — Simran Arora, Sabri Eyuboglu, Michael Zhang, et al., 2024
https://scholar.google.com/scholar?q=Simple+Linear+Attention+Language+Models+Balance+the+Recall-Throughput+Tradeoff+%28BASED%29
10. Gated Slot Attention for Efficient Linear-Time Sequence Modeling — Yu Zhang, Songlin Yang, Ruijie Zhu, et al., 2024
https://scholar.google.com/scholar?q=Gated+Slot+Attention+for+Efficient+Linear-Time+Sequence+Modeling
Interactive Visualization: MoM: Mixture-of-Memories Fixes Linear Attention Recall

This episode examines a Lightmatter-authored paper claiming photonic interconnects can cut inference prefill latency by up to 8.5x for long-context, Mixture-of-Experts workloads. The hosts unpack why prefill has become a dominant cost center as agentic coding pushes median prompt lengths toward 96K tokens, and explain the technical distinction between compute-bound prefill and memory-bandwidth-bound decode. They dig into the physics behind the claim: copper's one-meter reach limit at 224 Gbps per lane forces multi-rack scale-out, while 3D-integrated photonics decouples I/O from a chip's shoreline, enabling far higher bandwidth density. Throughout, the co-hosts push back on taking the headline multiplier at face value, stressing that the bandwidth specs come from Lightmatter's own published sheet and the performance gains from a simulator the company itself built and controls. It's a useful listen for anyone wanting a grounded, skeptical walkthrough of interconnect physics versus vendor-reported benchmarks in AI infrastructure claims.

Sources:
1. Scaling Inference Prefill with High-Radix Photonic Interconnects — Arulselvan Madhavan, Peter Carson, Taylor Groves, Thomas Graham, 2026
http://arxiv.org/abs/2609.01821
2. Accelerating Frontier MoE Training with 3D Integrated Optics — M. Bernadskiy, P. Carson, T. Graham, T. Groves, H. J. Lee, E. Yeh, 2025
https://scholar.google.com/scholar?q=Accelerating+Frontier+MoE+Training+with+3D+Integrated+Optics
3. DistServe: Disaggregating Prefill and Decoding for Goodput-Optimized LLM Serving — Y. Zhong, S. Liu, et al., 2024
https://scholar.google.com/scholar?q=DistServe%3A+Disaggregating+Prefill+and+Decoding+for+Goodput-Optimized+LLM+Serving
4. Sarathi-Serve: Taming Throughput-Latency Tradeoff in LLM Inference — A. Agrawal, N. Kedia, A. Panwar, et al., 2024 (OSDI)
https://scholar.google.com/scholar?q=Sarathi-Serve%3A+Taming+Throughput-Latency+Tradeoff+in+LLM+Inference
5. Flash Communication: Reducing Tensor Parallelization Bottleneck for Fast LLM Inference — Q. Li, B. Zhang, L. Ye, et al., 2024
https://scholar.google.com/scholar?q=Flash+Communication%3A+Reducing+Tensor+Parallelization+Bottleneck+for+Fast+LLM+Inference
6. DeepSeek-V3 Technical Report — DeepSeek-AI, 2024
https://scholar.google.com/scholar?q=DeepSeek-V3+Technical+Report
Interactive Visualization: Photonic Interconnects and the Race to Cut Prefill Latency

This episode revisits "The Case for Learned Index Structures" by Tim Kraska and coauthors from MIT and Google, which proposes replacing classic data structures like B-Trees, hash maps, and Bloom filters with small trained neural networks. The discussion covers how a B-Tree traversal is mathematically equivalent to estimating a cumulative distribution function, and how a tiny two-layer model can predict a key's position directly, turning a branchy O(log N) search into near-constant-time arithmetic suited to SIMD and GPU hardware. It also digs into how the same idea extends to hash functions tuned to real key distributions and to Bloom filters reframed as classifiers, complete with a backup filter to catch the model's false negatives. The hosts push back on each other over the paper's headline claims of 70% faster lookups and order-of-magnitude memory savings, debating whether results measured on static, read-only, in-memory workloads generalize or overstate the case. Listeners interested in database internals, the intersection of machine learning and systems design, or the tradeoffs between hand-engineered and learned structures will find the back-and-forth a useful gut check on a widely cited but contested idea.

Sources:
1. The Case for Learned Index Structures — Tim Kraska, Alex Beutel, Ed H. Chi, Jeffrey Dean, Neoklis Polyzotis, 2017
http://arxiv.org/abs/1712.01208
2. SOSD: A Benchmark for Learned Indexes — Andreas Kipf, Ryan Marcus, Alexander van Renen, Mihail Stoian, Alfons Kemper, Tim Kraska, Thomas Neumann, 2019
https://scholar.google.com/scholar?q=SOSD%3A+A+Benchmark+for+Learned+Indexes
3. ALEX: An Updatable Adaptive Learned Index — Jialin Ding, Umar Farooq Minhas, Jia Yu, Chi Wang, Jaeyoung Do, Yinan Li, Hantian Zhang, Badrish Chandramouli, Johannes Gehrke, Donald Kossmann, David Lomet, Tim Kraska, 2020
https://scholar.google.com/scholar?q=ALEX%3A+An+Updatable+Adaptive+Learned+Index
4. The PGM-index: A Fully-Dynamic Compressed Learned Index with Provable Worst-Case Bounds — Paolo Ferragina, Giorgio Vinciguerra, 2020
https://scholar.google.com/scholar?q=The+PGM-index%3A+A+Fully-Dynamic+Compressed+Learned+Index+with+Provable+Worst-Case+Bounds
5. FITing-Tree: A Data-aware Index Structure — Alex Galakatos, Michael Markovitch, Carsten Binnig, Rodrigo Fonseca, Tim Kraska, 2019
https://scholar.google.com/scholar?q=FITing-Tree%3A+A+Data-aware+Index+Structure
6. Organization and Maintenance of Large Ordered Indexes — Rudolf Bayer, Edward M. McCreight, 1972
https://scholar.google.com/scholar?q=Organization+and+Maintenance+of+Large+Ordered+Indexes
7. The Ubiquitous B-Tree — Douglas Comer, 1979
https://scholar.google.com/scholar?q=The+Ubiquitous+B-Tree
8. Making B+-Trees Cache Conscious in Main Memory — Jun Rao, Kenneth A. Ross, 2000
https://scholar.google.com/scholar?q=Making+B%2B-Trees+Cache+Conscious+in+Main+Memory
9. Space/Time Trade-offs in Hash Coding with Allowable Errors — Burton H. Bloom, 1970
https://scholar.google.com/scholar?q=Space%2FTime+Trade-offs+in+Hash+Coding+with+Allowable+Errors
10. Network Applications of Bloom Filters: A Survey — Andrei Broder, Michael Mitzenmacher, 2004
https://scholar.google.com/scholar?q=Network+Applications+of+Bloom+Filters%3A+A+Survey
11. A Model for Learned Bloom Filters and Optimizing by Sandwiching — Michael Mitzenmacher, 2018
https://scholar.google.com/scholar?q=A+Model+for+Learned+Bloom+Filters+and+Optimizing+by+Sandwiching
12. Cuckoo Filter: Practically Better Than Bloom — Bin Fan, Dave G. Andersen, Michael Kaminsky, Michael D. Mitzenmacher, 2014
https://scholar.google.com/scholar?q=Cuckoo+Filter%3A+Practically+Better+Than+Bloom
13. A model for learned bloom filters and related structures — M. Mitzenmacher, 2018
https://scholar.google.com/scholar?q=A+model+for+learned+bloom+filters+and+related+structures
14. A-Tree: A Bounded Approximate Index Structure — A. Galakatos, M. Markovitch, C. Binnig, R. Fonseca, T. Kraska, 2018
https://scholar.google.com/scholar?q=A-Tree%3A+A+Bounded+Approximate+Index+Structure
15. FAST: Fast Architecture Sensitive Tree Search on Modern CPUs and GPUs — C. Kim, J. Chhugani, N. Satish, et al., 2010
https://scholar.google.com/scholar?q=FAST%3A+Fast+Architecture+Sensitive+Tree+Search+on+Modern+CPUs+and+GPUs
16. Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer — N. Shazeer, A. Mirhoseini, K. Maziarz, et al., 2017
https://scholar.google.com/scholar?q=Outrageously+Large+Neural+Networks%3A+The+Sparsely-Gated+Mixture-of-Experts+Layer
Interactive Visualization: The Case for Learned Index Structures Revisited

This episode examines "AI Coaching for Accelerating Human Skill Development with Reinforcement Learning," a University of Pennsylvania and Johns Hopkins paper that challenges the assumption that AI assistance always benefits learners, arguing that copilots optimized for immediate task success can quietly prevent people from ever mastering a skill independently. The discussion traces the tension between over-assistance, which produces clean performance but no learning, and under-assistance, which produces uninstructive failure, connecting this to Manu Kapur's productive-failure research and decades of shared-control work from Dragan, Srinivasa, Reddy, and Levine. The core innovation discussed is framing coaching as a non-cooperative dynamic game rather than a cooperative one: the learner optimizes for immediate performance while the coach is trained on Value of Independence, a counterfactual measure of how well the human would perform if the AI were removed entirely. The hosts debate whether this framing is truly adversarial or just a shared goal on different timescales, concluding the reward structures can genuinely conflict moment-to-moment. The episode also situates the work against closest prior art, including the Cyber Racing Coach's fixed assistance-decay schedule, highlighting why a skill-aware, game-theoretic approach marks a meaningful departure from prior fading-assistance methods.

Sources:
1. AI Coaching for Accelerating Human Skill Development with Reinforcement Learning — Wei Wang, Enlin Gu, Antonio Loquercio, Haimin Hu, Rahul Mangharam, 2026
http://arxiv.org/abs/2606.25337v1
2. A Policy-Blending Formalism for Shared Control — Anca D. Dragan, Siddhartha S. Srinivasa, 2013
https://scholar.google.com/scholar?q=A+Policy-Blending+Formalism+for+Shared+Control
3. Shared Autonomy via Hindsight Optimization — Shervin Javdani, Siddhartha S. Srinivasa, J. Andrew Bagnell, 2015
https://scholar.google.com/scholar?q=Shared+Autonomy+via+Hindsight+Optimization
4. Shared Autonomy via Deep Reinforcement Learning — Siddharth Reddy, Anca D. Dragan, Sergey Levine, 2018
https://scholar.google.com/scholar?q=Shared+Autonomy+via+Deep+Reinforcement+Learning
5. Highway Driving with a Semi-Autonomous Alliance of Human and Machine (Shared Control for Highway Driving) — Jake Brawer / Vaibhav Gupta / related shared-control-for-driving lineage (e.g., Broad, Arkin, Ratliff, Howard, Argall), 2017-2019
https://scholar.google.com/scholar?q=Highway+Driving+with+a+Semi-Autonomous+Alliance+of+Human+and+Machine+%28Shared+Control+for+Highway+Driving%29
6. Dynamic Noncooperative Game Theory — Tamer Basar, Geert Jan Olsder, 1982 (2nd ed. 1999)
https://scholar.google.com/scholar?q=Dynamic+Noncooperative+Game+Theory
7. Planning for Autonomous Cars that Leverage Effects on Human Actions — Dorsa Sadigh, S. Shankar Sastry, Sanjit A. Seshia, Anca D. Dragan, 2016
https://scholar.google.com/scholar?q=Planning+for+Autonomous+Cars+that+Leverage+Effects+on+Human+Actions
8. Efficient Iterative Linear-Quadratic Approximations for Nonlinear Multi-Player General-Sum Differential Games (ILQGames) — David Fridovich-Keil, Ellis Ratner, Lasse Peters, Anca D. Dragan, Claire J. Tomlin, 2020
https://scholar.google.com/scholar?q=Efficient+Iterative+Linear-Quadratic+Approximations+for+Nonlinear+Multi-Player+General-Sum+Differential+Games+%28ILQGames%29
9. Who Plays First? Optimizing the Order of Play in Stackelberg Games with Many Robots (and related Stackelberg human-AV interaction work) — Haimin Hu, Zixu Zhang, Kensuke Nakamura, Andrea Bajcsy, Jaime F. Fisac (and related Fisac/Dragan lineage), 2023
https://scholar.google.com/scholar?q=Who+Plays+First%3F+Optimizing+the+Order+of+Play+in+Stackelberg+Games+with+Many+Robots+%28and+related+Stackelberg+human-AV+interaction+work%29
10. Legibility and Predictability of Robot Motion — Anca D. Dragan, Kenton C.T. Lee, Siddhartha S. Srinivasa, 2013
https://scholar.google.com/scholar?q=Legibility+and+Predictability+of+Robot+Motion
11. The Physical Presence of a Robot Tutor Increases Cognitive Learning Gains — Daniel Leyzberg, Samuel Spaulding, Mariya Toneva, Brian Scassellati, 2012
https://scholar.google.com/scholar?q=The+Physical+Presence+of+a+Robot+Tutor+Increases+Cognitive+Learning+Gains
12. Trajectory Deformations from Physical Human-Robot Interaction — Dylan P. Losey, Marcia K. O'Malley, 2017
https://scholar.google.com/scholar?q=Trajectory+Deformations+from+Physical+Human-Robot+Interaction
13. Productive Failure — Manu Kapur, 2008
https://scholar.google.com/scholar?q=Productive+Failure
14. Challenge Point: A Framework for Conceptualizing the Effects of Various Practice Conditions in Motor Learning — Mark A. Guadagnoli, Timothy D. Lee, 2004
https://scholar.google.com/scholar?q=Challenge+Point%3A+A+Framework+for+Conceptualizing+the+Effects+of+Various+Practice+Conditions+in+Motor+Learning
15. The Role of Tutoring in Problem Solving — David Wood, Jerome S. Bruner, Gail Ross, 1976
https://scholar.google.com/scholar?q=The+Role+of+Tutoring+in+Problem+Solving
16. The Relative Effectiveness of Human Tutoring, Intelligent Tutoring Systems, and Other Tutoring Systems — Kurt VanLehn, 2011
https://scholar.google.com/scholar?q=The+Relative+Effectiveness+of+Human+Tutoring%2C+Intelligent+Tutoring+Systems%2C+and+Other+Tutoring+Systems
17. Cyber Racing Coach: A Haptic Shared Control Framework for Teaching Advanced Driving Skills — C. Shen et al., 2025
https://scholar.google.com/scholar?q=Cyber+Racing+Coach%3A+A+Haptic+Shared+Control+Framework+for+Teaching+Advanced+Driving+Skills
18. Shared Autonomy for Proximal Teaching (Z-COACH) — M. Srivastava et al., 2025
https://scholar.google.com/scholar?q=Shared+Autonomy+for+Proximal+Teaching+%28Z-COACH%29
19. AssistanceZero: Scalably Solving Assistance Games — C. Laidlaw et al., 2025
https://scholar.google.com/scholar?q=AssistanceZero%3A+Scalably+Solving+Assistance+Games
20. Gradual Disempowerment: Systemic Existential Risks from Incremental AI Development — J. Kulveit et al., 2025
https://scholar.google.com/scholar?q=Gradual+Disempowerment%3A+Systemic+Existential+Risks+from+Incremental+AI+Development
21. Champion-Level Drone Racing Using Deep Reinforcement Learning — E. Kaufmann, L. Bauersfeld, A. Loquercio, M. Müller, V. Koltun, D. Scaramuzza, 2023
https://scholar.google.com/scholar?q=Champion-Level+Drone+Racing+Using+Deep+Reinforcement+Learning
22. Human-Robot Mutual Adaptation in Collaborative Tasks: Models and Experiments — S. Nikolaidis, D. Hsu, S. Srinivasa, 2017
https://scholar.google.com/scholar?q=Human-Robot+Mutual+Adaptation+in+Collaborative+Tasks%3A+Models+and+Experiments
23. Does Using Artificial Intelligence Assistance Accelerate Skill Decay and Hinder Skill Development Without Performers' Awareness? — B. N. Macnamara et al., 2024
https://scholar.google.com/scholar?q=Does+Using+Artificial+Intelligence+Assistance+Accelerate+Skill+Decay+and+Hinder+Skill+Development+Without+Performers%27+Awareness%3F
Interactive Visualization: AI Coaching That Preserves Human Skill Development

This episode surveys "Recursive Self-Improvement in AI," a paper by Mingguang Chen and colleagues that classifies 1,250 papers on how AI systems attempt to improve themselves. The discussion establishes precise definitions distinguishing agents, harnesses, and evaluators, then draws a critical line between bounded self-refinement (improvement against a fixed external evaluator) and open-ended recursive self-improvement (where the system also modifies its own criteria for success). It traces the intellectual lineage from I.J. Good's 1965 "intelligence explosion" concept through Schmidhuber's provably-optimal but practically unusable Gödel machines, showing how the field traded mathematical proof for empirically checkable but weaker signals like benchmarks and tests. The episode also introduces a verification hierarchy ranking formal verifiers above execution feedback, learned judges, and self-assessment, citing key findings that scoring reasoning steps beats scoring final answers, and that language models largely cannot self-correct without external feedback. Listeners interested in AI safety, agent architectures, or the theoretical limits of self-improving systems will find this a rigorous framework for cutting through loose talk about "self-improving AI."

Sources:
1. Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops — Mingguang Chen, Licheng Wang, Bo Qu, 2026
http://arxiv.org/abs/2607.07663
2. Speculations Concerning the First Ultraintelligent Machine — I. J. Good, 1965
https://scholar.google.com/scholar?q=Speculations+Concerning+the+First+Ultraintelligent+Machine
3. Gödel Machines: Self-Referential Universal Problem Solvers Making Provably Optimal Self-Improvements — Jürgen Schmidhuber, 2003 (extended 2006)
https://scholar.google.com/scholar?q=G%C3%B6del+Machines%3A+Self-Referential+Universal+Problem+Solvers+Making+Provably+Optimal+Self-Improvements
4. Self-Rewarding Language Models — Weizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Sainbayar Sukhbaatar, Jing Xu, Jason Weston (Meta AI/NYU), 2024
https://scholar.google.com/scholar?q=Self-Rewarding+Language+Models
5. Darwin Gödel Machine: Open-Ended Evolution of Self-Improving Agents — Jenny Zhang, Shengran Hu, Cong Lu, Robert Lange, Jeff Clune (Sakana AI / UBC), 2025
https://scholar.google.com/scholar?q=Darwin+G%C3%B6del+Machine%3A+Open-Ended+Evolution+of+Self-Improving+Agents
6. AI models collapse when trained on recursively generated data — Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Nicolas Papernot, Ross Anderson, Yarin Gal, 2024
https://scholar.google.com/scholar?q=AI+models+collapse+when+trained+on+recursively+generated+data
7. Large Language Models Cannot Self-Correct Reasoning Yet — Jie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng, Adams Wei Yu, Xinying Song, Denny Zhou, 2024 (ICLR)
https://scholar.google.com/scholar?q=Large+Language+Models+Cannot+Self-Correct+Reasoning+Yet
8. Let's Verify Step by Step — Hunter Lightman et al., 2023
https://scholar.google.com/scholar?q=Let%27s+Verify+Step+by+Step
9. Scaling Laws for Reward Model Overoptimization — Leo Gao, John Schulman, Jacob Hilton, 2022/2023
https://scholar.google.com/scholar?q=Scaling+Laws+for+Reward+Model+Overoptimization
10. Recursive Self-Improvement (Anthropic Institute blog post) — Anthropic, 2026
https://scholar.google.com/scholar?q=Recursive+Self-Improvement+%28Anthropic+Institute+blog+post%29
Interactive Visualization: The Great Debate: Bounded Refinement vs Open-Ended Self-Improvement

This episode examines "Post-Training Language Models for Gold-Medal Performance in Coding Competitions" from NVIDIA, whose Ultra-CC system scored 535.4 out of 600 at the 2026 International Olympiad in Informatics — beating not just the gold cutoff of 361.12 but the top human competitor's 498.27, live and under real contest conditions. The discussion breaks down why IOI-style problems are a harder test than typical coding benchmarks, since they demand inventing novel algorithms rather than recognizing familiar patterns, with partial credit across subtasks revealing the difference between no idea, the right idea with wrong complexity, and a fully correct solution. It covers the four-stage training pipeline behind the result: curating 22,000 competitive programming problems, distilling teacher-model reasoning traces for supervised fine-tuning, applying reinforcement learning from verifiable code-execution rewards via GRPO, and using a test-time strategy called GenCorrect to refine candidate solutions before submission. The episode also contrasts the two models built, a smaller mixture-of-experts Nano-CC with RL and a much larger Ultra-CC trained only with SFT, tracing the underlying architecture and training ideas back to foundational work like Shazeer's mixture-of-experts paper, Switch Transformer, InstructGPT, and DeepSeek's R1. Listeners interested in how far AI reasoning has come against elite human problem-solvers will find the specific mechanics behind this milestone result compelling.

Sources:
1. Post-Training Language Models for Gold-Medal Performance in Coding Competitions — Aleksander Ficek, Sean Narenthiran, Mehrzad Samadi, Somshubra Majumdar, Boris Ginsburg, 2026
http://arxiv.org/abs/2609.02849
2. Competitive programming with large reasoning models (o1-ioi, o3) — OpenAI et al., 2025
https://scholar.google.com/scholar?q=Competitive+programming+with+large+reasoning+models+%28o1-ioi%2C+o3%29
3. Scaling test-time compute to achieve IOI gold medal with open-weight models (GenCluster) — Mehrzad Samadi, Aleksander Ficek, Sean Narenthiran, et al., 2026
https://scholar.google.com/scholar?q=Scaling+test-time+compute+to+achieve+IOI+gold+medal+with+open-weight+models+%28GenCluster%29
4. Nemotron-cascade 2: Post-training LLMs with cascade RL and multi-domain on-policy distillation — Zhuolin Yang et al., 2026
https://scholar.google.com/scholar?q=Nemotron-cascade+2%3A+Post-training+LLMs+with+cascade+RL+and+multi-domain+on-policy+distillation
5. Large Language Models Cannot Self-Correct Reasoning Yet — Jie Huang, Xinyun Chen, Swaroop Mishra, et al., 2024
https://scholar.google.com/scholar?q=Large+Language+Models+Cannot+Self-Correct+Reasoning+Yet
6. DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models (DeepSeek-V3.2-Speciale) — DeepSeek-AI, 2025
https://scholar.google.com/scholar?q=DeepSeek-V3.2%3A+Pushing+the+Frontier+of+Open+Large+Language+Models+%28DeepSeek-V3.2-Speciale%29
Interactive Visualization: NVIDIA Ultra-CC Beats Top Human at IOI 2026

This episode examines MegaTrain, a method for full-precision training of 100-billion-parameter-plus language models on a single H200 GPU paired with 1.5 terabytes of host RAM, from a Notre Dame and Lehigh University team. The discussion centers on why memory, not compute, is the real bottleneck for most researchers, citing a survey showing only two of 167 surveyed U.S. universities average more than one H100 per student, while post-training work like instruction tuning and alignment increasingly demands full parameter and optimizer states without full pretraining-scale hardware. The hosts walk through the GPU memory hierarchy — from on-chip SRAM through HBM, host DDR5, and NVMe — and the 12-bytes-per-parameter cost of Adam optimizer state that makes a 70B model require 840 gigabytes of persistent storage. They contrast MegaTrain's approach with prior offloading systems like ZeRO-Offload and ZeRO-Infinity, highlighting the key architectural inversion: host memory becomes the authoritative store for all parameters and optimizer state, while GPU HBM is reduced to a transient scratchpad streaming one layer at a time across PCIe. Listeners interested in democratizing large-model training on constrained hardware will find the systems-level tradeoffs and pointed debate over whether this is genuinely novel or a repackaging of known offloading techniques particularly engaging.

Sources:
1. MegaTrain: Full Precision Training of 100B+ Parameter Large Language Models on a Single GPU — Zhengqing Yuan, Hanchi Sun, Lichao Sun, Yanfang Ye, 2026
http://arxiv.org/abs/2604.05091
2. vDNN: Virtualized Deep Neural Networks for Scalable, Memory-Efficient Neural Network Design — Minsoo Rhu, Natalia Gimelshein, Jason Clemons, Arslan Zulfiqar, Stephen W. Keckler, 2016
https://scholar.google.com/scholar?q=vDNN%3A+Virtualized+Deep+Neural+Networks+for+Scalable%2C+Memory-Efficient+Neural+Network+Design
3. ZeRO-Offload: Democratizing Billion-Scale Model Training — Jie Ren, Samyam Rajbhandari, Reza Yazdani Aminabadi, Olatunji Ruwase, Shuangyan Yang, Minjia Zhang, Dong Li, Yuxiong He, 2021
https://scholar.google.com/scholar?q=ZeRO-Offload%3A+Democratizing+Billion-Scale+Model+Training
4. ZeRO-Infinity: Breaking the GPU Memory Wall for Extreme Scale Deep Learning — Samyam Rajbhandari, Olatunji Ruwase, Jeff Rasley, Shaden Smith, Yuxiong He, 2021
https://scholar.google.com/scholar?q=ZeRO-Infinity%3A+Breaking+the+GPU+Memory+Wall+for+Extreme+Scale+Deep+Learning
5. JAX: composable transformations of Python+NumPy programs — James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, Skye Wanderman-Milne, 2018
https://scholar.google.com/scholar?q=JAX%3A+composable+transformations+of+Python%2BNumPy+programs
6. Chainer: A Next-Generation Open Source Framework for Deep Learning — Seiya Tokui, Kenta Oono, Shohei Hido, Justin Clayton, 2015
https://scholar.google.com/scholar?q=Chainer%3A+A+Next-Generation+Open+Source+Framework+for+Deep+Learning
7. Automatic Differentiation in PyTorch — Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, et al., 2017
https://scholar.google.com/scholar?q=Automatic+Differentiation+in+PyTorch
8. Training Deep Nets with Sublinear Memory Cost — Tianqi Chen, Bing Xu, Chiyuan Zhang, Carlos Guestrin, 2016
https://scholar.google.com/scholar?q=Training+Deep+Nets+with+Sublinear+Memory+Cost
9. Reducing Activation Recomputation in Large Transformer Models — Vijay Korthikanti, Jared Casper, Sangkug Lym, Lawrence McAfee, Michael Andersch, Mohammad Shoeybi, Bryan Catanzaro, 2022
https://scholar.google.com/scholar?q=Reducing+Activation+Recomputation+in+Large+Transformer+Models
10. Efficient Rematerialization for Deep Networks — Ravi Kumar, Manish Purohit, Zoya Svitkina, Erik Vee, Joshua Wang, 2019
https://scholar.google.com/scholar?q=Efficient+Rematerialization+for+Deep+Networks
11. Ratel: Optimizing Holistic Data Movement to Fine-Tune 100B Model on a Consumer GPU — Changyue Liao, Mo Sun, Zihan Yang, Jun Xie, Kaiqi Chen, Binhang Yuan, Fei Wu, Zeke Wang, 2025
https://scholar.google.com/scholar?q=Ratel%3A+Optimizing+Holistic+Data+Movement+to+Fine-Tune+100B+Model+on+a+Consumer+GPU
12. FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU — Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, et al., 2023
https://scholar.google.com/scholar?q=FlexGen%3A+High-Throughput+Generative+Inference+of+Large+Language+Models+with+a+Single+GPU
13. PatrickStar: Parallel Training of Pre-trained Models via a Chunk-based Memory Management — Jiarui Fang, Zilin Zhu, Shenggui Li, Hui Su, Yang Yu, Jie Zhou, Yang You, 2022
https://scholar.google.com/scholar?q=PatrickStar%3A+Parallel+Training+of+Pre-trained+Models+via+a+Chunk-based+Memory+Management
14. GaLore / other optimizer-state-aware memory-efficient training methods — Jiawei Zhao, Zhenyu Zhang, Beidi Chen, Zhangyang Wang, Anima Anandkumar, Yuandong Tian, 2024
https://scholar.google.com/scholar?q=GaLore+%2F+other+optimizer-state-aware+memory-efficient+training+methods
Interactive Visualization: MegaTrain: Training 120B-Parameter Models on a Single GPU

This episode examines a comparative study of three KV cache management strategies for LLM inference — vLLM's PagedAttention memory management, H2O's static sparsification, and InfiniGen's dynamic CPU-offload selection — tested side by side on identical hardware for the first time. The standout finding: both H2O and InfiniGen hit out-of-memory errors around 10,000 tokens, less than 10% of the 128K context window modern models claim to support, revealing that many eviction-based approaches can't even survive prefill on long documents. The discussion traces why KV caches exist at all (avoiding quadratic recomputation cost), how Grouped Query Attention reduces steady-state cache size but does nothing for the transient attention-score matrix that must be materialized during prefill to decide what to evict, and why that structural gap explains the paradigms' divergent failure modes. Testing spans Llama-3.1-8B and 70B, GPT-OSS-20B, and multiple benchmark datasets across four H100 GPUs. Listeners interested in the practical limits of long-context LLM serving — and why architectural tricks like GQA don't fully solve the memory problem — will find the paper's empirical exposure of these failure points compelling.

Sources:
1. Comparative Characterization of KV Cache Management Strategies for LLM Inference — Oteo Mamo, Olga Kogiou, Hyunjin Yi, Weikuan Yu, 2026
http://arxiv.org/abs/2604.05012
2. Efficient Memory Management for Large Language Model Serving with PagedAttention — W. Kwon et al., 2023
https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention
3. H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models — Z. Zhang et al., 2023
https://scholar.google.com/scholar?q=H2O%3A+Heavy-Hitter+Oracle+for+Efficient+Generative+Inference+of+Large+Language+Models
4. InfiniGen: Efficient Generative Inference of Large Language Models with Dynamic KV Cache Management — W. Lee, J. Lee, J. Seo, J. Sim, 2024
https://scholar.google.com/scholar?q=InfiniGen%3A+Efficient+Generative+Inference+of+Large+Language+Models+with+Dynamic+KV+Cache+Management
5. Characterizing the Behavior and Impact of KV Caching on Transformer Inferences under Concurrency — J. Ye, J. Cernuda, A. Maurya, X.-H. Sun, A. Kougkas, B. Nicolae, 2025
https://scholar.google.com/scholar?q=Characterizing+the+Behavior+and+Impact+of+KV+Caching+on+Transformer+Inferences+under+Concurrency
6. Efficient Streaming Language Models with Attention Sinks (StreamingLLM) — G. Xiao et al., 2023
https://scholar.google.com/scholar?q=Efficient+Streaming+Language+Models+with+Attention+Sinks+%28StreamingLLM%29
7. Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference — J. Tang et al., 2024
https://scholar.google.com/scholar?q=Quest%3A+Query-Aware+Sparsity+for+Efficient+Long-Context+LLM+Inference
8. KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization — C. Hooper et al., 2024
https://scholar.google.com/scholar?q=KVQuant%3A+Towards+10+Million+Context+Length+LLM+Inference+with+KV+Cache+Quantization
Interactive Visualization: KV Cache Management Faces Its First Head-to-Head Test

This episode examines "Language Models Can Control Their Own Attention" from KAIST AI, which tackles the memory bottleneck of long-context inference: at a million tokens, generating each token requires hauling roughly 15 gigabytes of key-value cache through memory, comparable to reloading the entire model's active parameters. The discussion traces prior fixes—StreamingLLM's attention-sink heuristic, H2O's cumulative attention scoring, and Quest's query-aware page selection—all grouped as "extrinsic scoring" methods that still require a full pass over context statistics before discarding anything. The paper's proposed alternative, Declarative Attention, repurposes chain-of-thought so the model states in its own reasoning which parts of the context it needs, letting the inference engine skip loading the rest instead of relying on an external scorer. The hosts debate whether self-reported relevance is trustworthy compared to an independent extrinsic estimate, since errors here directly create blind spots in what the model can see rather than showing up as recoverable noise. Listeners interested in long-context efficiency, KV-cache management, or the mechanics behind sparse attention will find the back-and-forth over whether this approach is elegant or quietly risky especially engaging.

Sources:
1. Language Models Can Control Their Own Attention — Namgyu Ho, Huzama Ahmad, Woosung Koh, Se-Young Yun, Tal Schuster, Cicero Nogueira dos Santos, 2026
http://arxiv.org/abs/2609.02737
2. Generating Long Sequences with Sparse Transformers — Rewon Child, Scott Gray, Alec Radford, Ilya Sutskever, 2019
https://scholar.google.com/scholar?q=Generating+Long+Sequences+with+Sparse+Transformers
3. H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models — Zhenyu Zhang, Ying Sheng, Tianyi Zhou, et al., 2023
https://scholar.google.com/scholar?q=H2O%3A+Heavy-Hitter+Oracle+for+Efficient+Generative+Inference+of+Large+Language+Models
4. Efficient Streaming Language Models with Attention Sinks — Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, Mike Lewis, 2023/2024
https://scholar.google.com/scholar?q=Efficient+Streaming+Language+Models+with+Attention+Sinks
5. Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference — Jiaming Tang, Yilong Zhao, Kan Zhu, Guangxuan Xiao, Baris Kasikci, Song Han, 2024
https://scholar.google.com/scholar?q=Quest%3A+Query-Aware+Sparsity+for+Efficient+Long-Context+LLM+Inference
6. Self-selected attention span for accelerating large language model inference — Tian Jin, Wanzin Yazar, Zifei Xu, Sayeh Sharify, Xin Wang, 2024
https://scholar.google.com/scholar?q=Self-selected+attention+span+for+accelerating+large+language+model+inference
7. System 2 Attention (is something you might need too) — Jason Weston, Sainbayar Sukhbaatar, 2023
https://scholar.google.com/scholar?q=System+2+Attention+%28is+something+you+might+need+too%29
8. Native sparse attention: Hardware-aligned and natively trainable sparse attention — Jingyang Yuan et al., 2025
https://scholar.google.com/scholar?q=Native+sparse+attention%3A+Hardware-aligned+and+natively+trainable+sparse+attention
Interactive Visualization: Language Models Learn to Fetch Only What They Need

This episode examines when it actually pays to split LLM inference hardware into four specialized pools rather than the now-standard two-way prefill/decode split, based on a paper proposing PDAF (prefill-attention, prefill-FFN, decode-attention, decode-FFN). It traces the reasoning from why agentic workloads — which can hit context sizes of 100,000+ tokens through repeated tool calls — strain hardware differently than chatbot traffic, through the compute-bound nature of prefill versus the memory-bandwidth-bound nature of decode, and why attention and FFN sublayers batch so differently that combining them on one device forces a similar compromise. It covers prior production systems (DistServe, Splitwise, StepFun's Step-3) that motivated these splits, and introduces the authors' HeteroPanacea simulator, validated against a real 8-node NVIDIA B200 cluster, which they use to search for when the reported up to 2.06x throughput gain actually materializes versus when the added complexity isn't worth it. Listeners interested in LLM serving infrastructure will find a grounded, skeptical take on a systems paper that resists overselling its own headline number.

Sources:
1. When Does Disaggregation Pay? Simulating Prefill--Decode--Attention--FFN Specialization for Agentic LLM Inference — Przemyslaw Forys, Haoran Wu, Can Xiao, Jiayi Nie, Tony Liu, Rika Antonova, Timothy Jones, Robert Mullins, Wayne Luk, Aaron Zhao, George A. Constantinides, 2026
http://arxiv.org/abs/2608.03741
2. DistServe: Disaggregating Prefill and Decoding for Goodput-Optimized Large Language Model Serving — Y. Zhong, S. Liu, J. Chen, J. Hu, Y. Zhu, X. Liu, X. Jin, H. Zhang, 2024
https://scholar.google.com/scholar?q=DistServe%3A+Disaggregating+Prefill+and+Decoding+for+Goodput-Optimized+Large+Language+Model+Serving
3. Splitwise: Efficient Generative LLM Inference Using Phase Splitting — P. Patel, E. Choukse, C. Zhang, A. Shah, I. Goiri, S. Maleki, R. Bianchini, 2024
https://scholar.google.com/scholar?q=Splitwise%3A+Efficient+Generative+LLM+Inference+Using+Phase+Splitting
4. Step-3 Is Large yet Affordable: Model-System Co-Design for Cost-Effective Decoding — StepFun, 2025
https://scholar.google.com/scholar?q=Step-3+Is+Large+yet+Affordable%3A+Model-System+Co-Design+for+Cost-Effective+Decoding
5. Mooncake: Trading More Storage for Less Computation — A KVCache-Centric Architecture for Serving LLM Chatbot — R. Qin, Z. Li, W. He, J. Cui, F. Ren, M. Zhang, Y. Wu, W. Zheng, X. Xu, 2025
https://scholar.google.com/scholar?q=Mooncake%3A+Trading+More+Storage+for+Less+Computation+%E2%80%94+A+KVCache-Centric+Architecture+for+Serving+LLM+Chatbot
6. MemExplorer: Navigating the Heterogeneous Memory Design Space for Agentic Inference NPUs — H. Wu, Z. Cao, Y. Lai, B. Lou, J. Nie, C. Xiao, T. Adeniran, P. Forys, et al., 2026
https://scholar.google.com/scholar?q=MemExplorer%3A+Navigating+the+Heterogeneous+Memory+Design+Space+for+Agentic+Inference+NPUs
Interactive Visualization: Disaggregating Prefill, Decode, Attention, and FFN for Agentic Inference

This episode examines TimesFM, Google Research's decoder-only foundation model for time-series forecasting, and its central claim that a single pretrained 200-million-parameter model can forecast unfamiliar datasets zero-shot, without fine-tuning, at accuracy close to models trained specifically on each dataset. The hosts trace the architecture's lineage from patching, borrowed from the Vision Transformer's image-patch approach and specifically from PatchTST's time-series adaptation, to TimesFM's own contribution of pairing patched inputs with a causal, GPT-style autoregressive setup that naturally handles variable context lengths. They contrast this against DeepAR's RNN-based forecasting, which still required target series in training, and against a 2023 NeurIPS trick of feeding raw numbers as text into large language models, which TimesFM claims to beat at a fraction of the cost. A key surprise is the training corpus itself: since real time-series data is far scarcer online than text, the roughly 100 billion timepoints come largely from Google Trends and Wikipedia pageviews, supplemented by synthetic ARMA and seasonal processes engineered to fill coverage gaps. Listeners interested in foundation models, forecasting infrastructure, or how architectural ideas transfer across modalities will find the discussion's skepticism about benchmark claims and evaluation rigor especially engaging as the hosts preview a closer look at the paper's actual scoring methodology.

Sources:
1. TimesFM: A Decoder-Only Foundation Model for Zero-Shot Time-Series Forecasting
https://arxiv.org/pdf/2310.10688
2. https://research.google/blog/timesfm-3-a-zero-shot-foundation-model-for-multivariate-forecasting/
https://research.google/blog/timesfm-3-a-zero-shot-foundation-model-for-multivariate-forecasting/
3. DeepAR: Probabilistic Forecasting with Autoregressive Recurrent Networks — David Salinas, Valentin Flunkert, Jan Gasthaus, Tim Januschowski, 2017 (arXiv), 2020 (Intl. J. Forecasting)
https://scholar.google.com/scholar?q=DeepAR%3A+Probabilistic+Forecasting+with+Autoregressive+Recurrent+Networks
4. N-BEATS: Neural Basis Expansion Analysis for Interpretable Time Series Forecasting — Boris Oreshkin, Dmitri Carpov, Nicolas Chapados, Yoshua Bengio, 2019 (arXiv), ICLR 2020
https://scholar.google.com/scholar?q=N-BEATS%3A+Neural+Basis+Expansion+Analysis+for+Interpretable+Time+Series+Forecasting
5. Temporal Fusion Transformers for Interpretable Multi-horizon Time Series Forecasting — Bryan Lim, Sercan Arik, Nicolas Loeff, Tomas Pfister, 2019 (arXiv), 2021 (Intl. J. Forecasting)
https://scholar.google.com/scholar?q=Temporal+Fusion+Transformers+for+Interpretable+Multi-horizon+Time+Series+Forecasting
6. Chronos: Learning the Language of Time Series — Abdul Fatir Ansari, Lorenzo Stella, et al. (Amazon), 2024
https://scholar.google.com/scholar?q=Chronos%3A+Learning+the+Language+of+Time+Series
7. Large Language Models Are Zero-Shot Time Series Forecasters — Nate Gruver, Marc Finzi, Shikai Qiu, Andrew Gordon Wilson, 2023 (NeurIPS)
https://scholar.google.com/scholar?q=Large+Language+Models+Are+Zero-Shot+Time+Series+Forecasters
8. One Fits All: Power General Time Series Analysis by Pretrained LM — Tian Zhou, Peisong Niu, Xue Wang, Liang Sun, Rong Jin, 2023 (NeurIPS)
https://scholar.google.com/scholar?q=One+Fits+All%3A+Power+General+Time+Series+Analysis+by+Pretrained+LM
9. Moirai: Unified Training of Universal Time Series Forecasting Transformers — Gerald Woo, Chenghao Liu, Akshat Kumar, Caiming Xiong, Silvio Savarese, Doug Arnold (Salesforce), 2024 (ICML)
https://scholar.google.com/scholar?q=Moirai%3A+Unified+Training+of+Universal+Time+Series+Forecasting+Transformers
10. Lag-Llama: Towards Foundation Models for Probabilistic Time Series Forecasting — Kashif Rasul, Arjun Ashok, Andrew Robert Williams, et al., 2024
https://scholar.google.com/scholar?q=Lag-Llama%3A+Towards+Foundation+Models+for+Probabilistic+Time+Series+Forecasting
11. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale — Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, et al. (Google), 2020
https://scholar.google.com/scholar?q=An+Image+is+Worth+16x16+Words%3A+Transformers+for+Image+Recognition+at+Scale
12. A Time Series is Worth 64 Words: Long-term Forecasting with Transformers — Yuqi Nie, Nam H. Nguyen, Phanwadee Sinthong, Jayant Kalagnanam, 2023 (ICLR)
https://scholar.google.com/scholar?q=A+Time+Series+is+Worth+64+Words%3A+Long-term+Forecasting+with+Transformers
13. Informer: Beyond Efficient Transformer for Long Sequence Time-Series Forecasting — Haoyi Zhou, Shanghang Zhang, Jieqi Peng, et al., 2021 (AAAI, Best Paper)
https://scholar.google.com/scholar?q=Informer%3A+Beyond+Efficient+Transformer+for+Long+Sequence+Time-Series+Forecasting
14. Masked Autoencoders Are Scalable Vision Learners — Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, Ross Girshick, 2021
https://scholar.google.com/scholar?q=Masked+Autoencoders+Are+Scalable+Vision+Learners
15. A Time Series is Worth 64 Words: Long-Term Forecasting with Transformers (PatchTST) — Yuqi Nie, Nam H. Nguyen, Phanwadee Sinthong, Jayant Kalagnanam, 2022
https://scholar.google.com/scholar?q=A+Time+Series+is+Worth+64+Words%3A+Long-Term+Forecasting+with+Transformers+%28PatchTST%29
16. TimeGPT-1 — Azul Garza, Max Mergenthaler-Canseco, 2023
https://scholar.google.com/scholar?q=TimeGPT-1
17. Training Compute-Optimal Large Language Models (Chinchilla) — Jordan Hoffmann et al., 2022
https://scholar.google.com/scholar?q=Training+Compute-Optimal+Large+Language+Models+%28Chinchilla%29
Interactive Visualization: TimesFM: A Decoder-Only Foundation Model for Zero-Shot Time-Series Forecasting

This episode examines TreeWY, a proposed method for speculative decoding verification in hybrid language models that mix standard attention with Gated DeltaNet linear-attention layers. The discussion explains why current systems like vLLM and SGLang must snapshot the full recurrent state at every draft position before verification, since GDN's decay-and-overwrite state update can't be partially rolled back — a cost that multiplies across branches and makes wide speculative draft trees prohibitively memory-expensive. It traces the problem to its root, from the memory-bandwidth bottleneck that motivates speculative decoding in the first place to the mathematical mechanics of the gated delta rule that make hybrid-model states lossy and irreversible. The paper's proposed fix reframes the state update as decayed additive attention with a corrected value, hinting at a way to verify an entire draft tree with a single triangular solve rather than exhaustive snapshotting. Listeners interested in LLM inference efficiency will find the episode's central claim striking: a roughly 128x reduction in per-node memory without sacrificing correctness guarantees, potentially unlocking much more aggressive tree-based speculation on hybrid architectures.

Sources:
1. TreeWY: Speculative Verification for Gated DeltaNet Hybrids
https://arxiv.org/pdf/2608.20961
2. Bole: Efficient Tree Speculation for Hybrid-Attention Language Models — L. Wang et al., 2026
https://scholar.google.com/scholar?q=Bole%3A+Efficient+Tree+Speculation+for+Hybrid-Attention+Language+Models
3. ReplaySSM: Cache SSM Inputs, Not State — Dao AI Lab and NVIDIA, 2026
https://scholar.google.com/scholar?q=ReplaySSM%3A+Cache+SSM+Inputs%2C+Not+State
4. STree: Speculative Tree Decoding for Hybrid State-Space Models — Y. Wu et al., 2025
https://scholar.google.com/scholar?q=STree%3A+Speculative+Tree+Decoding+for+Hybrid+State-Space+Models
5. Parallelizing Linear Transformers with the Delta Rule over Sequence Length — S. Yang et al., 2024
https://scholar.google.com/scholar?q=Parallelizing+Linear+Transformers+with+the+Delta+Rule+over+Sequence+Length
6. Gated Delta Networks: Improving Mamba2 with Delta Rule — S. Yang, J. Kautz, A. Hatamizadeh, 2025
https://scholar.google.com/scholar?q=Gated+Delta+Networks%3A+Improving+Mamba2+with+Delta+Rule

This episode examines "On First-Order Meta-Learning Algorithms" by Alex Nichol, Joshua Achiam, and John Schulman, which challenges the assumption that MAML's expensive second-derivative computation is essential for effective few-shot learning. The discussion traces the lineage from MAML's nested optimization — where an outer loop backpropagates through an inner loop's gradient steps via the Hessian — through First-Order MAML's approximation, to Reptile, a stripped-down algorithm that simply runs SGD on sampled tasks and nudges the initialization toward the result, with no meta-gradient or train-test split required. A central tension drives the conversation: why pulling an initialization toward "wherever SGD landed" produces a genuinely different target than plain joint training across tasks, rather than just averaging into one generic model. The hosts set up a Taylor-expansion argument to explain which gradient terms MAML, FOMAML, and Reptile weight differently, revealing the mathematical reason the cheaper approximation retains nearly all the useful signal. Listeners interested in the mechanics of meta-learning, gradient-based optimization tradeoffs, or the history of few-shot learning approaches will find the paper's practical implications for scaling meta-learning algorithms especially relevant.

Sources:
1. On First-Order Meta-Learning Algorithms — Alex Nichol, Joshua Achiam, John Schulman, 2018
http://arxiv.org/abs/1803.02999
2. Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks — Chelsea Finn, Pieter Abbeel, Sergey Levine, 2017
https://scholar.google.com/scholar?q=Model-Agnostic+Meta-Learning+for+Fast+Adaptation+of+Deep+Networks
3. How to train your MAML — Antreas Antoniou, Harrison Edwards, Amos Storkey, 2019
https://scholar.google.com/scholar?q=How+to+train+your+MAML
4. Meta-Learning with Implicit Gradients — Aravind Rajeswaran, Chelsea Finn, Sham Kakade, Sergey Levine, 2019
https://scholar.google.com/scholar?q=Meta-Learning+with+Implicit+Gradients
5. Optimization as a Model for Few-Shot Learning — Sachin Ravi, Hugo Larochelle, 2017
https://scholar.google.com/scholar?q=Optimization+as+a+Model+for+Few-Shot+Learning
6. Learning to Learn by Gradient Descent by Gradient Descent — Marcin Andrychowicz, Misha Denil, Sergio Gomez, Matthew W. Hoffman, David Pfau, Tom Schaul, Nando de Freitas, 2016
https://scholar.google.com/scholar?q=Learning+to+Learn+by+Gradient+Descent+by+Gradient+Descent
7. Using Fast Weights to Deblur Old Memories — Geoffrey E. Hinton, David C. Plaut, 1987
https://scholar.google.com/scholar?q=Using+Fast+Weights+to+Deblur+Old+Memories
8. Parallelized Stochastic Gradient Descent — Martin Zinkevich, Markus Weimer, Lihong Li, Alex J. Smola, 2010
https://scholar.google.com/scholar?q=Parallelized+Stochastic+Gradient+Descent

Browse by month

Sep 2026 Aug 2026 Jul 2026 Jun 2026 May 2026 Apr 2026 Mar 2026 Feb 2026 Jan 2026 Dec 2025 Nov 2025 Oct 2025 Sep 2025 Aug 2025