This episode takes up the thread from the published episode "MAML and the Basics of Meta-Learning" and shows how those ideas reappear in a much messier setting: a live agent that has to keep improving while it is already deployed. Instead of treating meta-learning as a clean laboratory exercise, the discussion follows MetaClaw as a continual agent system built for changing real workloads, where coding assistants, research agents, and other LLM-based tools face drift in tasks, tools, and failure modes. The hosts frame the paper as a concrete answer to a practical question: how an agent can keep learning on the job rather than waiting for the next full retraining cycle. The conversation focuses on MetaClaw’s two-speed adaptation design. The fast path updates behavior immediately through an external skill library, where failures are distilled into reusable behavioral instructions that can be injected at inference time; the slow path consolidates some of those lessons later through lightweight parameter updates. The hosts unpack the paper’s core formulation of the meta-model as base parameters plus skills, and they explain why that split matters for continual meta-learning: the agent is not only learning facts or storing transcripts, but improving its ability to adapt across a stream of tasks. They also dig into the process reward model, which scores intermediate reasoning and action steps, and the paper’s support-query separation, which keeps skill creation and later reinforcement updates from collapsing into stale self-training. A large part of the episode is about the systems implications of making that loop work in the wild. The hosts examine the paper’s zero-downtime claim in its narrower sense: skill updates can land during live use, while LoRA-based policy optimization is pushed into idle windows detected through sleep schedules, keyboard inactivity, and calendar availability, then swapped back into service later. That makes this episode a useful bridge not only from "MAML and the Basics of Meta-Learning" but, secondarily, from "Doc-to-LoRA: Internalizing Context as LoRA," because the slow adaptation path is explicitly about compressing recurring lessons into lightweight weight changes. The result is a detailed discussion of how MetaClaw tries to turn adaptation into an operational loop rather than a one-shot training event.

Sources:
1. MetaClaw: Just Talk -- An Agent That Meta-Learns and Evolves in the Wild — Peng Xia, Jianwen Chen, Xinyu Yang, Haoqin Tu, Jiaqi Liu, Kaiwen Xiong, Siwei Han, Shi Qiu, Haonian Ji, Yuyin Zhou, Zeyu Zheng, Cihang Xie, Huaxiu Yao, 2026
http://arxiv.org/abs/2603.17187
2. Reflexion: Language Agents with Verbal Reinforcement Learning — Noah Shinn, Federico Cassano, Edward Berman, Ashwin Gopinath, Karthik Narasimhan, Shunyu Yao, 2023
https://scholar.google.com/scholar?q=Reflexion%3A+Language+Agents+with+Verbal+Reinforcement+Learning
3. Voyager: An Open-Ended Embodied Agent with Large Language Models — Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, Anima Anandkumar, 2023
https://scholar.google.com/scholar?q=Voyager%3A+An+Open-Ended+Embodied+Agent+with+Large+Language+Models
4. ExpeL: LLM Agents Are Experiential Learners — Andrew Zhao, Daniel Huang, Quentin Xu, Matthieu Lin, Yong-Jin Liu, Gao Huang, 2023
https://scholar.google.com/scholar?q=ExpeL%3A+LLM+Agents+Are+Experiential+Learners
5. Agent Lightning: Train ANY AI Agents with Reinforcement Learning — Xufang Luo, Yuge Zhang, Zhiyuan He, Zilong Wang, Siyun Zhao, Dongsheng Li, Luna K. Qiu, Yuqing Yang, 2025
https://scholar.google.com/scholar?q=Agent+Lightning%3A+Train+ANY+AI+Agents+with+Reinforcement+Learning
6. ReAct: Synergizing Reasoning and Acting in Language Models — Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, Yuan Cao, 2022
https://scholar.google.com/scholar?q=ReAct%3A+Synergizing+Reasoning+and+Acting+in+Language+Models
7. Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks — Chelsea Finn, Pieter Abbeel, Sergey Levine, 2017
https://scholar.google.com/scholar?q=Model-Agnostic+Meta-Learning+for+Fast+Adaptation+of+Deep+Networks
8. LoRA: Low-Rank Adaptation of Large Language Models — Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, 2021
https://scholar.google.com/scholar?q=LoRA%3A+Low-Rank+Adaptation+of+Large+Language+Models
9. Who is introducing the failure? Automatically attributing failures of multi-agent systems via spectrum analysis — not verified from snippet, recent (exact year not verified from snippet)
https://scholar.google.com/scholar?q=Who+is+introducing+the+failure%3F+Automatically+attributing+failures+of+multi-agent+systems+via+spectrum+analysis
10. Weak-to-strong generalization with failure trajectories: A tree-based approach to elicit optimal policy in strong models — not verified from snippet, recent (exact year not verified from snippet)
https://scholar.google.com/scholar?q=Weak-to-strong+generalization+with+failure+trajectories%3A+A+tree-based+approach+to+elicit+optimal+policy+in+strong+models
11. Understanding Code Agent Behaviour: An Empirical Study of Success and Failure Trajectories — not verified from snippet, recent (exact year not verified from snippet)
https://scholar.google.com/scholar?q=Understanding+Code+Agent+Behaviour%3A+An+Empirical+Study+of+Success+and+Failure+Trajectories
12. Twosome: An efficient online framework to align LLMs with embodied environments via reinforcement learning — not verified from snippet, recent (exact year not verified from snippet)
https://scholar.google.com/scholar?q=Twosome%3A+An+efficient+online+framework+to+align+LLMs+with+embodied+environments+via+reinforcement+learning
13. AI Post Transformers: MAML and the Basics of Meta-Learning — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-29-maml-and-the-basics-of-meta-learning-7d449f.mp3
14. AI Post Transformers: Experiential Reinforcement Learning: Internalizing Reflection for Better Policy Training — Hal Turing & Dr. Ada Shannon, Fri,
https://podcast.do-not-panic.com/episodes/experiential-reinforcement-learning-internalizing-reflection-for-better-policy-t/
15. AI Post Transformers: Mem0: Scalable Long-Term Memory for AI Agents — Hal Turing & Dr. Ada Shannon, Tue,
https://podcast.do-not-panic.com/episodes/mem0-scalable-long-term-memory-for-ai-agents/
16. AI Post Transformers: NeurIPS 2025: A-Mem: Agentic Memory for LLM Agents — Hal Turing & Dr. Ada Shannon, Sat,
https://podcast.do-not-panic.com/episodes/neurips-2025-a-mem-agentic-memory-for-llm-agents/
17. AI Post Transformers: Evolving Language Models Without Labels: EVOL-RL — Hal Turing & Dr. Ada Shannon, Fri,
https://podcast.do-not-panic.com/episodes/evolving-language-models-without-labels-evol-rl/
18. AI Post Transformers: NeurIPS 2025: Reward Reasoning Model — Hal Turing & Dr. Ada Shannon, Sat,
https://podcast.do-not-panic.com/episodes/neurips-2025-reward-reasoning-model/
19. AI Post Transformers: Generalist Reward Modeling with Inference-Time Scaling — Hal Turing & Dr. Ada Shannon, Tue,
https://podcast.do-not-panic.com/episodes/generalist-reward-modeling-with-inference-time-scaling/
20. AI Post Transformers: LLM Benchmark Robustness to Linguistic Variation — Hal Turing & Dr. Ada Shannon, Tue,
https://podcast.do-not-panic.com/episodes/llm-benchmark-robustness-to-linguistic-variation/
Interactive Visualization: MetaClaw: Just Talk and Continual Agent Adaptation

This episode explores Doc-to-LoRA, a method for turning an entire document into a lightweight LoRA adapter so a language model can answer later questions without repeatedly rereading the source text. It explains how the paper combines context distillation, LoRA fine-tuning, and a Perceiver-style hypernetwork that ingests variable-length documents and emits fixed-size parameter updates, using chunking to handle longer inputs. The discussion highlights reported results such as near-perfect zero-shot performance on synthetic long-context retrieval beyond 32K tokens and improved efficiency on long-document question answering through lower update latency, lower peak memory use, and reduced KV-cache costs at inference time. It also digs into the systems argument behind the work, framing reusable internalized memory as a different primitive from prompting, while questioning how well the approach holds up outside limited-query evaluations and whether its benefits persist against alternatives like prompt compression or keeping context externally.

Sources:
1. Doc-to-LoRA: Internalizing Context as LoRA
https://arxiv.org/pdf/2602.15902
2. 2603.13875
https://arxiv.org/abs/2603.13875
3. 2510.03215
https://arxiv.org/abs/2510.03215
4. LoRA: Low-Rank Adaptation of Large Language Models — Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, 2022
https://scholar.google.com/scholar?q=LoRA%3A+Low-Rank+Adaptation+of+Large+Language+Models
5. QLoRA: Efficient Finetuning of Quantized LLMs — Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke Zettlemoyer, 2023
https://scholar.google.com/scholar?q=QLoRA%3A+Efficient+Finetuning+of+Quantized+LLMs
6. AdaLoRA: Adaptive Budget Allocation for Parameter-Efficient Fine-Tuning — Qingru Zhang, Minshuo Chen, Alexander Bukharin, Nikos Karampatziakis, Pengcheng He, Yu Cheng, Weizhu Chen, Tuo Zhao, 2023
https://scholar.google.com/scholar?q=AdaLoRA%3A+Adaptive+Budget+Allocation+for+Parameter-Efficient+Fine-Tuning
7. DoRA: Weight-Decomposed Low-Rank Adaptation — Shih-Yang Liu, Chien-Yi Wang, Hongxu Yin, Pavlo Molchanov, Yu-Chiang Frank Wang, Kwang-Ting Cheng, Min-Hung Chen, 2024
https://scholar.google.com/scholar?q=DoRA%3A+Weight-Decomposed+Low-Rank+Adaptation
8. HyperNetworks — David Ha, Andrew Dai, Quoc V. Le, 2016
https://scholar.google.com/scholar?q=HyperNetworks
9. Parameter-efficient Multi-task Fine-tuning for Transformers via Shared Hypernetworks — Rabeeh Karimi Mahabadi, Sebastian Ruder, Mostafa Dehghani, James Henderson, 2021
https://scholar.google.com/scholar?q=Parameter-efficient+Multi-task+Fine-tuning+for+Transformers+via+Shared+Hypernetworks
10. HyperPrompt: Prompt-based Task-Conditioning of Transformers — Yun He, Huaixiu Steven Zheng, Yi Tay, Jai Gupta, Yu Du, Vamsi Aribandi, Zhe Zhao, Yaguang Li, Zhao Chen, Donald Metzler, Heng-Tze Cheng, Ed H. Chi, 2022
https://scholar.google.com/scholar?q=HyperPrompt%3A+Prompt-based+Task-Conditioning+of+Transformers
11. Doc-to-LoRA: Learning to Instantly Internalize Contexts — Rujikorn Charakorn, Edoardo Cetin, Shinnosuke Uesaka, Robert Tjarko Lange, 2026
https://scholar.google.com/scholar?q=Doc-to-LoRA%3A+Learning+to+Instantly+Internalize+Contexts
12. Text-to-LoRA: Instant Transformer Adaption — Rujikorn Charakorn, Edoardo Cetin, Yujin Tang, Robert Tjarko Lange, 2025
https://scholar.google.com/scholar?q=Text-to-LoRA%3A+Instant+Transformer+Adaption
13. Generative Adapter: Contextualizing Language Models in Parameters with a Single Forward Pass — Tianyu Chen, Huanran Fang, Patrick Xia, Xiaodong Liu, Benjamin Van Durme, Luke Zettlemoyer, Jianfeng Gao, Hao Cheng, 2025
https://scholar.google.com/scholar?q=Generative+Adapter%3A+Contextualizing+Language+Models+in+Parameters+with+a+Single+Forward+Pass
14. Cartridges: Lightweight and General-Purpose Long Context Representations via Self-Study — Sabri Eyuboglu, Ryan Saul Ehrlich, Simran Arora, Neel Guha, Dylan Zinsley, Emily Ruoyu Liu, William Tennien, Atri Rudra, James Zou, Azalia Mirhoseini, Christopher Re, 2025
https://scholar.google.com/scholar?q=Cartridges%3A+Lightweight+and+General-Purpose+Long+Context+Representations+via+Self-Study
15. Propagating Knowledge Updates to LMs through Distillation — Suchin Padmanabhan, Yoon Kim Onoe, Michael Zhang, Greg Durrett, Eunsol Choi, 2023
https://scholar.google.com/scholar?q=Propagating+Knowledge+Updates+to+LMs+through+Distillation
16. LLMLingua-2: Data Distillation for Efficient and Faithful Task-Agnostic Prompt Compression — Zefan Pan, Qipeng Wu, Hao Jiang, Mengzhou Xia, Xuefei Luo, Jiaqi Zhang, Qingyu Lin, Viktor Ruhle, Yi Yang, Chin-Yew Lin, H. Vicky Zhao, Lidong Qiu, Dongmei Zhang, 2024
https://scholar.google.com/scholar?q=LLMLingua-2%3A+Data+Distillation+for+Efficient+and+Faithful+Task-Agnostic+Prompt+Compression
17. RazorAttention: Efficient KV Cache Compression Through Retrieval Heads — Hanlin Tang et al., 2024
https://scholar.google.com/scholar?q=RazorAttention%3A+Efficient+KV+Cache+Compression+Through+Retrieval+Heads
18. Not All Heads Matter: A Head-Level KV Cache Compression Method with Integrated Retrieval and Reasoning — Yu Fu et al., 2024/2025
https://scholar.google.com/scholar?q=Not+All+Heads+Matter%3A+A+Head-Level+KV+Cache+Compression+Method+with+Integrated+Retrieval+and+Reasoning
19. How Much Knowledge Can You Pack into a LoRA Adapter without Harming LLM? — Sergey Pletenev et al., 2025
https://scholar.google.com/scholar?q=How+Much+Knowledge+Can+You+Pack+into+a+LoRA+Adapter+without+Harming+LLM%3F
20. Can Fine-Tuning Erase Your Edits? On the Fragile Coexistence of Knowledge Editing and Adaptation — Yinjie Cheng et al., 2025
https://scholar.google.com/scholar?q=Can+Fine-Tuning+Erase+Your+Edits%3F+On+the+Fragile+Coexistence+of+Knowledge+Editing+and+Adaptation
21. Memorization in In-Context Learning — Shahriar Golchin et al., 2024
https://scholar.google.com/scholar?q=Memorization+in+In-Context+Learning
22. In-Context Learning can Perform Continual Learning Like Humans — Liuwang Kang et al., 2025
https://scholar.google.com/scholar?q=In-Context+Learning+can+Perform+Continual+Learning+Like+Humans
23. AI Post Transformers: LoRA: Low-Rank Adaptation of Large Language Models — Hal Turing & Dr. Ada Shannon, Fri,
https://podcast.do-not-panic.com/episodes/lora-low-rank-adaptation-of-large-language-models/
24. AI Post Transformers: ShadowKV: High-Throughput Long-Context LLM Inference — Hal Turing & Dr. Ada Shannon, Wed,
https://podcast.do-not-panic.com/episodes/shadowkv-high-throughput-long-context-llm-inference/
25. AI Post Transformers: Lookahead Q-Cache for Consistent KV Eviction — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-25-lookahead-q-cache-for-consistent-kv-evic-d97b09.mp3
26. AI Post Transformers: Mem0: Scalable Long-Term Memory for AI Agents — Hal Turing & Dr. Ada Shannon, Tue,
https://podcast.do-not-panic.com/episodes/mem0-scalable-long-term-memory-for-ai-agents/
27. AI Post Transformers: Kimi Linear: Efficient Expressive Attention Architecture — Hal Turing & Dr. Ada Shannon, Sun,
https://podcast.do-not-panic.com/episodes/kimi-linear-efficient-expressive-attention-architecture/
28. AI Post Transformers: ComoRAG: Cognitively Inspired Narrative Reasoning — Hal Turing & Dr. Ada Shannon, Tue,
https://podcast.do-not-panic.com/episodes/comorag-cognitively-inspired-narrative-reasoning/
Interactive Visualization: Doc-to-LoRA: Internalizing Context as LoRA

This episode explores meta-learning through the lens of MAML, explaining how it differs from ordinary supervised learning and standard transfer learning by explicitly training models to adapt quickly to new tasks after just one or a few gradient updates. It walks through the core idea of optimizing for post-update performance, including the role of second-order meta-gradients and the simpler first-order approximation, while placing MAML within the broader landscape of few-shot and gradient-based meta-learning. The discussion also highlights why the paper mattered across multiple domains, covering not just classification benchmarks like Omniglot and MiniImagenet but also regression with sinusoid fitting and reinforcement learning with fast-adapting policies. A listener would find it interesting because it turns a buzzword-heavy area into a concrete framework for thinking about how models can learn to learn, setting up deeper discussions about newer systems built on these ideas.

Interactive Visualization: MAML and the Basics of Meta-Learning
Sources:
1. MAML and the Basics of Meta-Learning
https://arxiv.org/pdf/1703.03400
2. https://par.nsf.gov/servlets/purl/10427895
https://par.nsf.gov/servlets/purl/10427895
3. Optimization as a Model for Few-Shot Learning — Sachin Ravi, Hugo Larochelle, 2017
https://scholar.google.com/scholar?q=Optimization+as+a+Model+for+Few-Shot+Learning
4. Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks — Chelsea Finn, Pieter Abbeel, Sergey Levine, 2017
https://scholar.google.com/scholar?q=Model-Agnostic+Meta-Learning+for+Fast+Adaptation+of+Deep+Networks
5. RL^2: Fast Reinforcement Learning via Slow Reinforcement Learning — Yan Duan, John Schulman, Xi Chen, Peter L. Bartlett, Ilya Sutskever, Pieter Abbeel, 2016
https://scholar.google.com/scholar?q=RL%5E2%3A+Fast+Reinforcement+Learning+via+Slow+Reinforcement+Learning
6. Meta-Learning in Neural Networks: A Survey — Timothy M. Hospedales, Antreas Antoniou, Paul Micaelli, Amos J. Storkey, 2021
https://scholar.google.com/scholar?q=Meta-Learning+in+Neural+Networks%3A+A+Survey
7. Matching Networks for One Shot Learning — Oriol Vinyals, Charles Blundell, Timothy Lillicrap, Koray Kavukcuoglu, Daan Wierstra, 2016
https://scholar.google.com/scholar?q=Matching+Networks+for+One+Shot+Learning
8. Prototypical Networks for Few-shot Learning — Jake Snell, Kevin Swersky, Richard Zemel, 2017
https://scholar.google.com/scholar?q=Prototypical+Networks+for+Few-shot+Learning
9. A Closer Look at Few-shot Classification — Wei-Yu Chen, Yen-Cheng Liu, Zsolt Kira, Yu-Chiang Frank Wang, Jia-Bin Huang, 2019
https://scholar.google.com/scholar?q=A+Closer+Look+at+Few-shot+Classification
10. Generalizing from a Few Examples — Yaqing Wang, Quanming Yao, James T. Kwok, Lionel M. Ni, 2020
https://scholar.google.com/scholar?q=Generalizing+from+a+Few+Examples
11. Meta-SGD: Learning to Learn Quickly for Few-Shot Learning — Zhenguo Li, Fengwei Zhou, Fei Chen, Hang Li, 2017
https://scholar.google.com/scholar?q=Meta-SGD%3A+Learning+to+Learn+Quickly+for+Few-Shot+Learning
12. On First-Order Meta-Learning Algorithms — Alex Nichol, Joshua Achiam, John Schulman, 2018
https://scholar.google.com/scholar?q=On+First-Order+Meta-Learning+Algorithms
13. How to Train Your MAML — Antreas Antoniou, Harrison Edwards, Amos Storkey, 2019
https://scholar.google.com/scholar?q=How+to+Train+Your+MAML
14. Meta-learning with Differentiable Closed-Form Solvers — Luca Bertinetto, Joao F. Henriques, Philip H. S. Torr, Andrea Vedaldi, 2018
https://scholar.google.com/scholar?q=Meta-learning+with+Differentiable+Closed-Form+Solvers
15. Efficient Off-Policy Meta-Reinforcement Learning via Probabilistic Context Variables — Kate Rakelly, Aurick Zhou, Chelsea Finn, Sergey Levine, Deirdre Quillen, 2019
https://scholar.google.com/scholar?q=Efficient+Off-Policy+Meta-Reinforcement+Learning+via+Probabilistic+Context+Variables
16. Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning — Ronald J. Williams, 1992
https://scholar.google.com/scholar?q=Simple+Statistical+Gradient-Following+Algorithms+for+Connectionist+Reinforcement+Learning
17. Trust Region Policy Optimization — John Schulman, Sergey Levine, Philipp Moritz, Michael Jordan, Pieter Abbeel, 2015
https://scholar.google.com/scholar?q=Trust+Region+Policy+Optimization
18. Proximal Policy Optimization Algorithms — John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, Oleg Klimov, 2017
https://scholar.google.com/scholar?q=Proximal+Policy+Optimization+Algorithms
19. Meta-Learning with Memory-Augmented Neural Networks — Adam Santoro, Sergey Bartunov, Matthew Botvinick, Daan Wierstra, Timothy Lillicrap, 2016
https://scholar.google.com/scholar?q=Meta-Learning+with+Memory-Augmented+Neural+Networks
20. Learning to Learn by Gradient Descent by Gradient Descent — Marcin Andrychowicz, Misha Denil, Sergio Gomez, Matthew W. Hoffman, David Pfau, Tom Schaul, Nando de Freitas, 2016
https://scholar.google.com/scholar?q=Learning+to+Learn+by+Gradient+Descent+by+Gradient+Descent
21. Transformers learn in-context by gradient descent — Johannes von Oswald et al., 2022
https://scholar.google.com/scholar?q=Transformers+learn+in-context+by+gradient+descent
22. In-context Learning and Gradient Descent Revisited — Gilad Deutch et al., 2023
https://scholar.google.com/scholar?q=In-context+Learning+and+Gradient+Descent+Revisited
23. Low-Rank Few-Shot Adaptation of Vision-Language Models — Maxime Zanella and Ismail Ben Ayed, 2024
https://scholar.google.com/scholar?q=Low-Rank+Few-Shot+Adaptation+of+Vision-Language+Models
24. Meta-Adapter: An Online Few-shot Learner for Vision-Language Model — Cheng Cheng et al., 2023
https://scholar.google.com/scholar?q=Meta-Adapter%3A+An+Online+Few-shot+Learner+for+Vision-Language+Model
25. Cross-Domain Few-Shot Learning via Adaptive Transformer Networks — Naeem Paeedeh et al., 2024
https://scholar.google.com/scholar?q=Cross-Domain+Few-Shot+Learning+via+Adaptive+Transformer+Networks
26. Few-shot Adaptation of Multi-modal Foundation Models: A Survey — Fan Liu et al., 2024
https://scholar.google.com/scholar?q=Few-shot+Adaptation+of+Multi-modal+Foundation+Models%3A+A+Survey
27. AI Post Transformers: In-Context Learning as Implicit Learning Algorithms — Hal Turing & Dr. Ada Shannon, Wed,
https://podcast.do-not-panic.com/episodes/in-context-learning-as-implicit-learning-algorithms/
28. AI Post Transformers: NVIDIA: TTT-E2E: Unlocking Long-Context Learning via End-to-End Test-Time Training — Hal Turing & Dr. Ada Shannon, Sat,
https://podcast.do-not-panic.com/episodes/nvidia-ttt-e2e-unlocking-long-context-learning-via-end-to-end-test-time-training/
29. AI Post Transformers: Zero-Shot Context Generalization in Reinforcement Learning from Few Training Contexts — Hal Turing & Dr. Ada Shannon, Tue,
https://podcast.do-not-panic.com/episodes/zero-shot-context-generalization-in-reinforcement-learning-from-few-training-con/
30. AI Post Transformers: A 2024 Survey Analyzing Generalization in Deep Reinforcement Learning — Hal Turing & Dr. Ada Shannon, Fri,
https://podcast.do-not-panic.com/episodes/a-2024-survey-analyzing-generalization-in-deep-reinforcement-learning/
Interactive Visualization: MAML and the Basics of Meta-Learning

This episode explores a perspective paper arguing that the next major leap in AI may come less from scaling a single model and more from organizing intelligence across agents, tools, humans, and institutions. It explains key ideas including agentic AI, multi-agent reasoning, human-AI centaurs, and “societies of thought,” where useful reasoning may emerge through internal dialogue among specialized perspectives rather than just longer single-threaded outputs. The discussion contrasts straightforward parameter scaling with the harder problem of organizational design, emphasizing that collective intelligence only works under specific conditions such as good communication, balanced participation, and careful aggregation. Listeners would find it interesting because it reframes the usual singularity story into a concrete debate about coordination, role design, and whether intelligence scales socially as much as technically.

Sources:
1. Agentic AI and the next intelligence explosion — James Evans, Benjamin Bratton, Blaise Agüera y Arcas, 2026
http://arxiv.org/abs/2603.20639
2. Evidence for a Collective Intelligence Factor in the Performance of Human Groups — Anita Williams Woolley, Christopher F. Chabris, Alex Pentland, Nada Hashmi, Thomas W. Malone, 2010
https://scholar.google.com/scholar?q=Evidence+for+a+Collective+Intelligence+Factor+in+the+Performance+of+Human+Groups
3. AI-enhanced Collective Intelligence — Hao Cui, Taha Yasseri, 2024
https://scholar.google.com/scholar?q=AI-enhanced+Collective+Intelligence
4. Artificial Intelligence for Collective Intelligence: a National-scale Research Strategy — Seth Bullock and many coauthors, 2024
https://scholar.google.com/scholar?q=Artificial+Intelligence+for+Collective+Intelligence%3A+a+National-scale+Research+Strategy
5. Artificial Intelligence versus Collective Intelligence — Harry Halpin, 2025
https://scholar.google.com/scholar?q=Artificial+Intelligence+versus+Collective+Intelligence
6. Man-Computer Symbiosis — J. C. R. Licklider, 1960
https://scholar.google.com/scholar?q=Man-Computer+Symbiosis
7. Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of AI on Knowledge Worker Productivity and Quality — Fabrizio Dell'Acqua, Edward McFowland III, Ethan Mollick, Hila Lifshitz-Assaf, Katherine C. Kellogg, Saran Rajendran, Lisa Krayer, Francois Candelon, Karim R. Lakhani, 2023
https://scholar.google.com/scholar?q=Navigating+the+Jagged+Technological+Frontier%3A+Field+Experimental+Evidence+of+the+Effects+of+AI+on+Knowledge+Worker+Productivity+and+Quality
8. When Combinations of Humans and AI are Useful: A Systematic Review and Meta-analysis — Michelle Vaccaro, Abdullah Almaatouq, Thomas W. Malone, 2024
https://scholar.google.com/scholar?q=When+Combinations+of+Humans+and+AI+are+Useful%3A+A+Systematic+Review+and+Meta-analysis
9. Effective Generative AI: The Human-Algorithm Centaur — Soroush Saghafian, Lihi Idan, 2024
https://scholar.google.com/scholar?q=Effective+Generative+AI%3A+The+Human-Algorithm+Centaur
10. Reasoning Models Generate Societies of Thought — Junsol Kim, Shiyang Lai, Nino Scherrer, Blaise Aguera y Arcas, James Evans, 2026
https://scholar.google.com/scholar?q=Reasoning+Models+Generate+Societies+of+Thought
11. Self-Consistency Improves Chain of Thought Reasoning in Language Models — Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, Denny Zhou, 2022
https://scholar.google.com/scholar?q=Self-Consistency+Improves+Chain+of+Thought+Reasoning+in+Language+Models
12. Improving Factuality and Reasoning in Language Models through Multiagent Debate — Yilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum, Igor Mordatch, 2024
https://scholar.google.com/scholar?q=Improving+Factuality+and+Reasoning+in+Language+Models+through+Multiagent+Debate
13. DeepSeek-R1 Incentivizes Reasoning in LLMs through Reinforcement Learning — Daya Guo and many coauthors, 2025
https://scholar.google.com/scholar?q=DeepSeek-R1+Incentivizes+Reasoning+in+LLMs+through+Reinforcement+Learning
14. CAMEL: Communicative Agents for "Mind" Exploration of Large Scale Language Model Society — Guohao Li, Hasan Abed Al Kader Hammoud, Hani Itani, Dmitrii Khizbullin, Bernard Ghanem, 2023
https://scholar.google.com/scholar?q=CAMEL%3A+Communicative+Agents+for+%22Mind%22+Exploration+of+Large+Scale+Language+Model+Society
15. AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation — Qingyun Wu, Gagan Bansal, Jieyu Zhang, Yiran Wu, Beibin Li, Erkang Zhu, Li Jiang, Xiaoyun Zhang, Shaokun Zhang, Jiale Liu, Ahmed Hassan Awadallah, Ryen W. White, Doug Burger, Chi Wang, 2024
https://scholar.google.com/scholar?q=AutoGen%3A+Enabling+Next-Gen+LLM+Applications+via+Multi-Agent+Conversation
16. A Survey on LLM-based Multi-agent Systems: Workflow, Infrastructure, and Challenges — Xinyi Li, Sai Wang, Siqi Zeng, Yu Wu, Yi Yang, 2024
https://scholar.google.com/scholar?q=A+Survey+on+LLM-based+Multi-agent+Systems%3A+Workflow%2C+Infrastructure%2C+and+Challenges
17. Deep Reinforcement Learning from Human Preferences — P. F. Christiano et al., 2017
https://scholar.google.com/scholar?q=Deep+Reinforcement+Learning+from+Human+Preferences
18. Constitutional AI: Harmlessness from AI Feedback — Y. Bai et al., 2022
https://scholar.google.com/scholar?q=Constitutional+AI%3A+Harmlessness+from+AI+Feedback
19. Large AI Models Are Cultural and Social Technologies — H. Farrell, A. Gopnik, C. Shalizi, J. Evans, 2025
https://scholar.google.com/scholar?q=Large+AI+Models+Are+Cultural+and+Social+Technologies
20. Governing the Commons: The Evolution of Institutions for Collective Action — E. Ostrom, 1990
https://scholar.google.com/scholar?q=Governing+the+Commons%3A+The+Evolution+of+Institutions+for+Collective+Action
21. Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them — Mirac Suzgun, Nathan Scales, Nathanael Scharli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc V. Le, Ed H. Chi, Denny Zhou, Jason Wei, 2022
https://scholar.google.com/scholar?q=Challenging+BIG-Bench+Tasks+and+Whether+Chain-of-Thought+Can+Solve+Them
22. SoftCoT++: Test-Time Scaling with Soft Chain-of-Thought Reasoning — Yige Xu, Xu Guo, Zhiwei Zeng, Chunyan Miao, 2025
https://scholar.google.com/scholar?q=SoftCoT%2B%2B%3A+Test-Time+Scaling+with+Soft+Chain-of-Thought+Reasoning
23. Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning — Chengqi Lyu, Songyang Gao, Yuzhe Gu, Wenwei Zhang, Jianfei Gao and others, 2025
https://scholar.google.com/scholar?q=Exploring+the+Limit+of+Outcome+Reward+for+Learning+Mathematical+Reasoning
24. Linking Process to Outcome: Conditional Reward Modeling for LLM Reasoning — Zheng Zhang, Ziwei Shan, Kaitao Song, Yexin Li, Kan Ren, 2025
https://scholar.google.com/scholar?q=Linking+Process+to+Outcome%3A+Conditional+Reward+Modeling+for+LLM+Reasoning
25. SophiaVL-R1: Reinforcing MLLMs Reasoning with Thinking Reward — Kaixuan Fan, Kaituo Feng, Haoming Lyu, Dongzhan Zhou, Xiangyu Yue, 2025
https://scholar.google.com/scholar?q=SophiaVL-R1%3A+Reinforcing+MLLMs+Reasoning+with+Thinking+Reward
26. Parsel: Algorithmic Reasoning with Language Models by Composing Decompositions — Eric Zelikman, Qian Huang and others, 2022
https://scholar.google.com/scholar?q=Parsel%3A+Algorithmic+Reasoning+with+Language+Models+by+Composing+Decompositions
27. Emergent Hierarchical Reasoning in LLMs through Reinforcement Learning — Haozhe Wang, Qixin Xu, Che Liu, Junhong Wu, Fangzhen Lin, Wenhu Chen, 2025
https://scholar.google.com/scholar?q=Emergent+Hierarchical+Reasoning+in+LLMs+through+Reinforcement+Learning
28. Debate4MATH: Multi-Agent Debate for Fine-Grained Reasoning in Math — Shaowei Zhang, Deyi Xiong, 2025
https://scholar.google.com/scholar?q=Debate4MATH%3A+Multi-Agent+Debate+for+Fine-Grained+Reasoning+in+Math
29. Learning to Break: Knowledge-Enhanced Reasoning in Multi-Agent Debate System — Haotian Wang, Xiyuan Du, Weijiang Yu, Qianglong Chen, Kun Zhu, Zheng Chu, Lian Yan, Yi Guan, 2025
https://scholar.google.com/scholar?q=Learning+to+Break%3A+Knowledge-Enhanced+Reasoning+in+Multi-Agent+Debate+System
30. AI Post Transformers: Reasoning Models Generate Societies of Thought — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/reasoning-models-generate-societies-of-thought/
31. AI Post Transformers: HyperAgents and Metacognitive Self-Improvement — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-26-hyperagents-and-metacognitive-self-impro-de711a.mp3
32. AI Post Transformers: Bloom: an open source tool for automated behavioral evaluations — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/bloom-an-open-source-tool-for-automated-behavioral-evaluations/
33. AI Post Transformers: NeurIPS 2025: Reward Reasoning Model — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/neurips-2025-reward-reasoning-model/
34. AI Post Transformers: MASA: Meta-Awareness via Self-Alignment Reinforcement Learning — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/masa-meta-awareness-via-self-alignment-reinforcement-learning/
35. AI Post Transformers: Evolving Language Models Without Labels: EVOL-RL — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/evolving-language-models-without-labels-evol-rl/
36. AI Post Transformers: LeCun's AMI Energy-Based Models and the Path to Autonomous Intelligence — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/lecuns-ami-energy-based-models-and-the-path-to-autonomous-intelligence/
Interactive Visualization: Agentic AI and the Next Intelligence Explosion

This episode explores a wireless communications paper that reframes multi-antenna receive combining as a distributionally robust estimation problem rather than a collection of separate techniques like MMSE, Capon beamforming, and diagonal loading. It explains how the paper uses the language of robust statistics and distributionally robust optimization to handle uncertainty in channels, covariance estimates, impulsive noise, hardware distortions, and limited pilot data, including the provocative claim that explicit channel estimation may not always be necessary. The discussion also connects this framework to integrated sensing and communication, where transmitted signals can be structured, correlated, and complex-valued enough to require richer estimation methods such as kernel ridge regression and potentially neural receivers. A listener would find it interesting because it ties together classical signal processing and modern machine learning ideas into a single view of how receivers can stay effective when real-world assumptions break down.

Sources:
1. Distributionally Robust Receive Combining for Wireless Receivers
https://arxiv.org/abs/2401.12345
2. Distributionally Robust Optimization: A Review — Hamed Rahimian and Sanjay Mehrotra, 2019
https://scholar.google.com/scholar?q=Distributionally+Robust+Optimization%3A+A+Review
3. Data-Driven Distributionally Robust Optimization Using the Wasserstein Metric: Performance Guarantees and Tractable Reformulations — Peyman Mohajerin Esfahani and Daniel Kuhn, 2018
https://scholar.google.com/scholar?q=Data-Driven+Distributionally+Robust+Optimization+Using+the+Wasserstein+Metric%3A+Performance+Guarantees+and+Tractable+Reformulations
4. From Data to Decisions: Distributionally Robust Optimization is Optimal — Daniel Kuhn, Peyman Mohajerin Esfahani, and Bart Van Parys, 2017
https://scholar.google.com/scholar?q=From+Data+to+Decisions%3A+Distributionally+Robust+Optimization+is+Optimal
5. Certifying Some Distributional Robustness with Principled Adversarial Training — Aman Sinha, Hongseok Namkoong, and John Duchi, 2018
https://scholar.google.com/scholar?q=Certifying+Some+Distributional+Robustness+with+Principled+Adversarial+Training
6. High-Resolution Frequency-Wavenumber Spectrum Analysis — John Capon, 1969
https://scholar.google.com/scholar?q=High-Resolution+Frequency-Wavenumber+Spectrum+Analysis
7. Massive MIMO Detection Techniques: A Survey — Mahmoud A. M. Albreem, Markku Juntti, and Shahriar Shahabuddin, 2019
https://scholar.google.com/scholar?q=Massive+MIMO+Detection+Techniques%3A+A+Survey
8. Model-Driven Deep Learning for MIMO Detection — Hao He, Chao-Kai Wen, Shi Jin, and Geoffrey Ye Li, 2020
https://scholar.google.com/scholar?q=Model-Driven+Deep+Learning+for+MIMO+Detection
9. Adaptive Neural Signal Detection for Massive MIMO — Mehrdad Khani, Mohammad Alizadeh, Jakob Hoydis, and Phil Fleming, 2020
https://scholar.google.com/scholar?q=Adaptive+Neural+Signal+Detection+for+Massive+MIMO
10. Robust Estimation of a Location Parameter — Peter J. Huber, 1964
https://scholar.google.com/scholar?q=Robust+Estimation+of+a+Location+Parameter
11. The Influence Curve and its Role in Robust Estimation — Frank R. Hampel, 1974
https://scholar.google.com/scholar?q=The+Influence+Curve+and+its+Role+in+Robust+Estimation
12. Robust Estimation in Signal Processing: A Tutorial-Style Treatment of Fundamental Concepts — Abdelhak M. Zoubir, Visa Koivunen, Youssef Chakhchoukh, and Michael Muma, 2012
https://scholar.google.com/scholar?q=Robust+Estimation+in+Signal+Processing%3A+A+Tutorial-Style+Treatment+of+Fundamental+Concepts
13. A Robust Learning Approach for Regression Models Based on Distributionally Robust Optimization — Ruidi Chen and Ioannis Ch. Paschalidis, 2018
https://scholar.google.com/scholar?q=A+Robust+Learning+Approach+for+Regression+Models+Based+on+Distributionally+Robust+Optimization
14. Survey of RF Communications and Sensing Convergence Research — Binoj Paul, Anil R. Chiriyath, and Daniel W. Bliss, 2017
https://scholar.google.com/scholar?q=Survey+of+RF+Communications+and+Sensing+Convergence+Research
15. Integrated Sensing and Communication in 6G: Motivations, Use Cases, Requirements, Challenges and Future Directions — D. K. Pin Tan, J. He, Y. Li, A. Bayesteh, Y. Chen, P. Zhu, and W. Tong, 2021
https://scholar.google.com/scholar?q=Integrated+Sensing+and+Communication+in+6G%3A+Motivations%2C+Use+Cases%2C+Requirements%2C+Challenges+and+Future+Directions
16. Integrated Sensing and Communication Signals Toward 5G-A and 6G: A Survey — Zhiqing Wei, Hanyang Qu, Yuan Wang, Xin Yuan, Huici Wu, Ying Du, Kaifeng Han, Ning Zhang, and Zhiyong Feng, 2023
https://scholar.google.com/scholar?q=Integrated+Sensing+and+Communication+Signals+Toward+5G-A+and+6G%3A+A+Survey
17. A Survey on Machine Learning Enhanced Integrated Sensing and Communication Systems: Architectures, Algorithms, and Applications — M. A. K. Respati and coauthors, 2024
https://scholar.google.com/scholar?q=A+Survey+on+Machine+Learning+Enhanced+Integrated+Sensing+and+Communication+Systems%3A+Architectures%2C+Algorithms%2C+and+Applications
18. Regularization Networks and Support Vector Machines — Theodoros Evgeniou, Massimiliano Pontil, and Tomaso Poggio, 2000
https://scholar.google.com/scholar?q=Regularization+Networks+and+Support+Vector+Machines
19. Divide and Conquer Kernel Ridge Regression — Yuchen Zhang, John Duchi, and Martin Wainwright, 2013
https://scholar.google.com/scholar?q=Divide+and+Conquer+Kernel+Ridge+Regression
20. Random Fourier Features for Kernel Ridge Regression: Approximation Bounds and Statistical Guarantees — Haim Avron, Michael Kapralov, Cameron Musco, Christopher Musco, Ameya Velingker, and Amir Zandieh, 2017
https://scholar.google.com/scholar?q=Random+Fourier+Features+for+Kernel+Ridge+Regression%3A+Approximation+Bounds+and+Statistical+Guarantees
21. On the Optimality of Misspecified Kernel Ridge Regression — Haobo Zhang, Yicheng Li, Weihao Lu, and Qian Lin, 2023
https://scholar.google.com/scholar?q=On+the+Optimality+of+Misspecified+Kernel+Ridge+Regression
22. On Robust Capon Beamforming and Diagonal Loading — Jian Li, Petre Stoica, and Zhisong Wang, 2003
https://scholar.google.com/scholar?q=On+Robust+Capon+Beamforming+and+Diagonal+Loading
23. Robust Minimum Variance Beamforming — R. G. Lorenz and Stephen P. Boyd, 2005
https://scholar.google.com/scholar?q=Robust+Minimum+Variance+Beamforming
24. Distributionally Robust Optimization and Generalization in Kernel Methods — Maximilian Staib and Stefanie Jegelka, 2019
https://scholar.google.com/scholar?q=Distributionally+Robust+Optimization+and+Generalization+in+Kernel+Methods
25. Regularization via Mass Transportation — Soroosh Shafieezadeh-Abadeh, Daniel Kuhn, and Peyman Mohajerin Esfahani, 2019
https://scholar.google.com/scholar?q=Regularization+via+Mass+Transportation
26. Toward Dual-Functional Radar-Communication Systems: Optimal Waveform Design — Fan Liu, Longfei Zhou, Christos Masouros, Ang Li, Wei Luo, and Athina Petropulu, 2018
https://scholar.google.com/scholar?q=Toward+Dual-Functional+Radar-Communication+Systems%3A+Optimal+Waveform+Design
27. Channel estimation for reconfigurable intelligent surface aided multi-user mmWave MIMO systems — approx. multiple authors in RIS/mmWave MIMO literature, recent
https://scholar.google.com/scholar?q=Channel+estimation+for+reconfigurable+intelligent+surface+aided+multi-user+mmWave+MIMO+systems
28. Channel estimation for movable antenna communication systems: A framework based on compressed sensing — approx. multiple authors in movable-antenna communications, recent
https://scholar.google.com/scholar?q=Channel+estimation+for+movable+antenna+communication+systems%3A+A+framework+based+on+compressed+sensing
29. Deep learning-based channel estimation for wideband hybrid mmWave massive MIMO — approx. multiple authors in hybrid mmWave massive MIMO, recent
https://scholar.google.com/scholar?q=Deep+learning-based+channel+estimation+for+wideband+hybrid+mmWave+massive+MIMO
30. Structured channel covariance estimation from limited samples for large antenna arrays — approx. multiple authors in massive-MIMO covariance estimation, recent
https://scholar.google.com/scholar?q=Structured+channel+covariance+estimation+from+limited+samples+for+large+antenna+arrays
31. Robust estimation of angular power spectrum in massive MIMO under covariance estimation errors: Learning centers and scales of Gaussians — approx. multiple authors in massive-MIMO APS estimation, recent
https://scholar.google.com/scholar?q=Robust+estimation+of+angular+power+spectrum+in+massive+MIMO+under+covariance+estimation+errors%3A+Learning+centers+and+scales+of+Gaussians
32. Self-Supervised Learning Enhanced Channel Estimation in Massive MIMO System With Low-Resolution ADCs — approx. multiple authors in self-supervised massive MIMO, recent
https://scholar.google.com/scholar?q=Self-Supervised+Learning+Enhanced+Channel+Estimation+in+Massive+MIMO+System+With+Low-Resolution+ADCs
33. Zero-Shot Self-Supervised Channel Estimation in Massive MIMO LEO Satellites Systems — approx. multiple authors in satellite massive MIMO, recent
https://scholar.google.com/scholar?q=Zero-Shot+Self-Supervised+Channel+Estimation+in+Massive+MIMO+LEO+Satellites+Systems

This episode explores a paper on long-lived AI agents that keep adapting to changing real-world tasks without being taken offline. It explains the paper’s central idea of combining two learning timescales: fast updates through an evolving skill library and slower policy improvement through parameter-efficient weight tuning such as LoRA. The discussion unpacks why agent learning is harder than ordinary one-shot language modeling, since failure happens across whole action trajectories involving tools, recovery strategies, and multi-step decisions. Listeners would find it interesting because the episode connects this proposal to broader debates about memory retrieval, skill libraries, and continual meta-learning, while questioning whether dynamic skill evolution alone can already deliver substantial behavioral improvement.

Sources:
1. MetaClaw: Just Talk -- An Agent That Meta-Learns and Evolves in the Wild — Peng Xia, Jianwen Chen, Xinyu Yang, Haoqin Tu, Jiaqi Liu, Kaiwen Xiong, Siwei Han, Shi Qiu, Haonian Ji, Yuyin Zhou, Zeyu Zheng, Cihang Xie, Huaxiu Yao, 2026
http://arxiv.org/abs/2603.17187
2. Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks — Chelsea Finn, Pieter Abbeel, Sergey Levine, 2017
https://scholar.google.com/scholar?q=Model-Agnostic+Meta-Learning+for+Fast+Adaptation+of+Deep+Networks
3. Meta-Learning with Sparse Experience Replay for Lifelong Language Learning — Nithin Holla, Pushkar Mishra, Helen Yannakoudakis, Ekaterina Shutova, 2020
https://scholar.google.com/scholar?q=Meta-Learning+with+Sparse+Experience+Replay+for+Lifelong+Language+Learning
4. La-MAML: Look-ahead Meta Learning for Continual Learning — Gunshi Gupta, Karmesh Yadav, Liam Paull, 2020
https://scholar.google.com/scholar?q=La-MAML%3A+Look-ahead+Meta+Learning+for+Continual+Learning
5. When Meta-Learning Meets Online and Continual Learning: A Survey — Jaehyeon Son, Soochan Lee, Gunhee Kim, 2025
https://scholar.google.com/scholar?q=When+Meta-Learning+Meets+Online+and+Continual+Learning%3A+A+Survey
6. Reflexion: Language Agents with Verbal Reinforcement Learning — Noah Shinn, Federico Cassano, Beck Labash, Ashwin Gopinath, Karthik Narasimhan, Shunyu Yao, 2023
https://scholar.google.com/scholar?q=Reflexion%3A+Language+Agents+with+Verbal+Reinforcement+Learning
7. Voyager: An Open-Ended Embodied Agent with Large Language Models — Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, Anima Anandkumar, 2023
https://scholar.google.com/scholar?q=Voyager%3A+An+Open-Ended+Embodied+Agent+with+Large+Language+Models
8. ExpeL: LLM Agents Are Experiential Learners — Andrew Zhao, Daniel Huang, Quentin Xu, Matthieu Lin, Yong-Jin Liu, Gao Huang, 2023
https://scholar.google.com/scholar?q=ExpeL%3A+LLM+Agents+Are+Experiential+Learners
9. Solving Math Word Problems with Process- and Outcome-Based Feedback — Jonathan Uesato, Nate Kushman, Ramana Kumar, Francis Song, Noah Siegel, Lisa Wang, Antonia Creswell, Geoffrey Irving, Irina Higgins, 2022
https://scholar.google.com/scholar?q=Solving+Math+Word+Problems+with+Process-+and+Outcome-Based+Feedback
10. Let's Verify Step by Step — Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, Karl Cobbe, 2023
https://scholar.google.com/scholar?q=Let%27s+Verify+Step+by+Step
11. Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations — Peiyi Wang, Lei Li, Zhihong Shao, Runxin Xu, Damai Dai, Yifei Li, Deli Chen, Yu Wu, Zhifang Sui, 2024
https://scholar.google.com/scholar?q=Math-Shepherd%3A+Verify+and+Reinforce+LLMs+Step-by-step+without+Human+Annotations
12. Process Reward Models for LLM Agents: Practical Framework and Directions — Sanjiban Choudhury, 2025
https://scholar.google.com/scholar?q=Process+Reward+Models+for+LLM+Agents%3A+Practical+Framework+and+Directions
13. ReAct: Synergizing Reasoning and Acting in Language Models — Shunyu Yao, Jeffrey Zhao, Dian Yu, et al., 2022
https://scholar.google.com/scholar?q=ReAct%3A+Synergizing+Reasoning+and+Acting+in+Language+Models
14. LoRA: Low-Rank Adaptation of Large Language Models — Edward J. Hu, Yelong Shen, Phillip Wallis, et al., 2021
https://scholar.google.com/scholar?q=LoRA%3A+Low-Rank+Adaptation+of+Large+Language+Models
15. Proximal Policy Optimization Algorithms — John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, Oleg Klimov, 2017
https://scholar.google.com/scholar?q=Proximal+Policy+Optimization+Algorithms
16. Trajectory-Informed Memory Generation for Self-Improving Agent Systems — Gaodan Fang et al., 2026
https://scholar.google.com/scholar?q=Trajectory-Informed+Memory+Generation+for+Self-Improving+Agent+Systems
17. Reward Under Attack: Analyzing the Robustness and Hackability of Process Reward Models — Rishabh Tiwari et al., 2026
https://scholar.google.com/scholar?q=Reward+Under+Attack%3A+Analyzing+the+Robustness+and+Hackability+of+Process+Reward+Models
18. Process Reinforcement through Implicit Rewards — Ganqu Cui et al., 2025
https://scholar.google.com/scholar?q=Process+Reinforcement+through+Implicit+Rewards
19. Reinforcement Learning for Self-Improving Agent with Skill Library — Jiongxiao Wang et al., 2025
https://scholar.google.com/scholar?q=Reinforcement+Learning+for+Self-Improving+Agent+with+Skill+Library
20. When Single-Agent with Skills Replace Multi-Agent Systems and When They Fail — Xiaoxiao Li, 2026
https://scholar.google.com/scholar?q=When+Single-Agent+with+Skills+Replace+Multi-Agent+Systems+and+When+They+Fail
21. AI Post Transformers: Agentic Context Engineering: Evolving Contexts for Self-Improving LLMs — Hal Turing & Dr. Ada Shannon, Fri,
https://podcast.do-not-panic.com/episodes/agentic-context-engineering-evolving-contexts-for-self-improving-llms/
22. AI Post Transformers: Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models — Hal Turing & Dr. Ada Shannon, Wed,
https://podcast.do-not-panic.com/episodes/agentic-context-engineering-evolving-contexts-for-self-improving-language-models/
23. AI Post Transformers: Experiential Reinforcement Learning: Internalizing Reflection for Better Policy Training — Hal Turing & Dr. Ada Shannon, Fri,
https://podcast.do-not-panic.com/episodes/experiential-reinforcement-learning-internalizing-reflection-for-better-policy-t/
24. AI Post Transformers: Metacognition and Skill Discovery in LLM Math Reasoning — Hal Turing & Dr. Ada Shannon, Mon,
https://podcast.do-not-panic.com/episodes/metacognition-and-skill-discovery-in-llm-math-reasoning/
25. AI Post Transformers: NeurIPS 2025: SeRL: Self-Play Reinforcement Learning for Large Language Models with Limited Data — Hal Turing & Dr. Ada Shannon, Sat,
https://podcast.do-not-panic.com/episodes/neurips-2025-serl-self-play-reinforcement-learning-for-large-language-models-wit/
26. AI Post Transformers: MATTRL: Collaborative Test-Time Reinforcement Learning for Multi-Agent Reasoning — Hal Turing & Dr. Ada Shannon, Mon,
https://podcast.do-not-panic.com/episodes/mattrl-collaborative-test-time-reinforcement-learning-for-multi-agent-reasoning/
27. AI Post Transformers: A Framework for LLM Application Safety Evaluation — Hal Turing & Dr. Ada Shannon, Mon,
https://podcast.do-not-panic.com/episodes/a-framework-for-llm-application-safety-evaluation/

This episode examines Splitwise: Efficient Generative LLM Inference Using Phase Splitting, a 2024 systems paper from researchers at the University of Washington and Microsoft, and centers the discussion on a simple claim with large deployment consequences: prompt prefill and token decode are different enough that they should not necessarily run on the same hardware. The hosts walk through the basic mechanics of generative inference, explaining prefill as the parallel, compute-heavy stage that processes the prompt, and decode as the sequential, KV-cache-driven stage that generates tokens one by one. That distinction sets up the paper’s core argument that modern serving stacks are paying a penalty by treating inference as a uniform workload when its phases are constrained by very different resources. The conversation stays focused on why that split matters in practice. It unpacks phase heterogeneity in terms of throughput, latency, utilization, memory pressure, and power draw, and explains why decode can remain bottlenecked by memory bandwidth and capacity even on newer accelerators with far more raw FLOPs. From there, the episode explores Splitwise’s broader systems framing: if compute is scaling faster than memory, then assigning prefill to high-throughput hardware and decode to cheaper or lower-power machines may be a more realistic datacenter strategy than continuing to push everything through one homogeneous GPU fleet. The hosts also emphasize power-normalized evaluation as a more honest lens for operators than simple box-for-box performance comparisons. Along the way, the episode places Splitwise in public context alongside ORCA, PagedAttention, and SARATHI without losing its anchor. Those earlier systems are used to clarify what Splitwise does and does not claim: continuous batching, KV-cache-aware memory management, and batch reshaping all improve serving efficiency, but they do not eliminate the underlying asymmetry between prefill and decode. The result is a grounded discussion of phase splitting as a deployment decision rather than a purely algorithmic trick, with particular attention to where prefill-decode disaggregation looks compelling, where it depends on the realities of cluster design, and where the limits of PD disaggregation still leave open systems questions.

Interactive Visualization: Splitwise: Phase-Split LLM Inference
Sources:
1. Splitwise: Efficient generative LLM inference using phase splitting — Pratyush Patel, Esha Choukse, Chaojie Zhang, Aashaka Shah, Íñigo Goiri, Saeed Maleki, Ricardo Bianchini, 2023
http://arxiv.org/abs/2311.18677
2. DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving — Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, Hao Zhang, 2024
http://arxiv.org/abs/2401.09670
3. DualPath: Breaking the Storage Bandwidth Bottleneck in Agentic LLM Inference — Yongtong Wu, Shaoyuan Chen, Yinmin Zhong, Rilin Huang, Yixuan Tan, Wentao Zhang, Liyue Zhang, Shangyan Zhou, Yuxuan Liu, Shunfeng Zhou, Mingxing Zhang, Xin Jin, Panpan Huang, 2026
http://arxiv.org/abs/2602.21548
4. Not All Prefills Are Equal: PPD Disaggregation for Multi-turn LLM Serving — Zongze Li, Jingyu Liu, Zach Xu, Yineng Zhang, Tahseen Rabbani, Ce Zhang, 2026
http://arxiv.org/abs/2603.13358
5. Orca: A Distributed Serving System for Transformer-Based Generative Models — Gyeong-In Yu, Joo Seong Jeong, Gun-Woo Kim, Seungtae Kim, Byung-Gon Chun, 2022
https://scholar.google.com/scholar?q=Orca%3A+A+Distributed+Serving+System+for+Transformer-Based+Generative+Models
6. Efficient Memory Management for Large Language Model Serving with PagedAttention — Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, Ion Stoica, 2023
https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention
7. SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills — Animesh Agrawal, Aakash Panwar, Jaya Mohan, Nakul Kwatra, Bhaskar S. Gulavani, Ramachandran Ramjee, 2023
https://scholar.google.com/scholar?q=SARATHI%3A+Efficient+LLM+Inference+by+Piggybacking+Decodes+with+Chunked+Prefills
8. DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving — Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, Hao Zhang, 2024
https://scholar.google.com/scholar?q=DistServe%3A+Disaggregating+Prefill+and+Decoding+for+Goodput-optimized+Large+Language+Model+Serving
9. Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving — Ruoyu Qin, Zheming Li, Weiran He, Mingxing Zhang, Yongwei Wu, Weimin Zheng, Xinran Xu, 2024
https://scholar.google.com/scholar?q=Mooncake%3A+A+KVCache-centric+Disaggregated+Architecture+for+LLM+Serving
10. KVLink: Accelerating Large Language Models via Efficient KV Cache Reuse — Jingbo Yang, Bairu Hou, Wei Wei, Yujia Bao, Shiyu Chang, 2025
https://scholar.google.com/scholar?q=KVLink%3A+Accelerating+Large+Language+Models+via+Efficient+KV+Cache+Reuse
11. HyperRAG: Enhancing Quality-Efficiency Tradeoffs in Retrieval-Augmented Generation with Reranker KV-Cache Reuse — Yuwei An, Yihua Cheng, Seo Jin Park, Junchen Jiang, 2025
https://scholar.google.com/scholar?q=HyperRAG%3A+Enhancing+Quality-Efficiency+Tradeoffs+in+Retrieval-Augmented+Generation+with+Reranker+KV-Cache+Reuse
12. ProphetKV: User-Query-Driven Selective Recomputation for Efficient KV Cache Reuse in Retrieval-Augmented Generation — Shihao Wang, Jiahao Chen, Yanqi Pan, Hao Huang and colleagues, 2026
https://scholar.google.com/scholar?q=ProphetKV%3A+User-Query-Driven+Selective+Recomputation+for+Efficient+KV+Cache+Reuse+in+Retrieval-Augmented+Generation
13. Accelerating LLM Inference with Staged Speculative Decoding — Benjamin Spector, Chris Re, 2023
https://scholar.google.com/scholar?q=Accelerating+LLM+Inference+with+Staged+Speculative+Decoding
14. SpecExec: Massively Parallel Speculative Decoding for Interactive LLM Inference on Consumer Devices — Ruslan Svirschevski, Avner May, Zhuoming Chen, Beidi Chen, Zhihao Jia, Max Ryabinin, 2024
https://scholar.google.com/scholar?q=SpecExec%3A+Massively+Parallel+Speculative+Decoding+for+Interactive+LLM+Inference+on+Consumer+Devices
15. KVDirect: Distributed Disaggregated LLM Inference — Shiyang Chen, Rain Jiang, Dezhi Yu, Jinlai Xu, Mengyuan Chao and colleagues, 2025
https://scholar.google.com/scholar?q=KVDirect%3A+Distributed+Disaggregated+LLM+Inference
16. Arrow: Adaptive Scheduling Mechanisms for Disaggregated LLM Inference Architecture — Yu Wu, Tongxuan Liu, Yuting Zeng, Siyu Wu, Jun Xiong, Xianzhe Dong and colleagues, 2025
https://scholar.google.com/scholar?q=Arrow%3A+Adaptive+Scheduling+Mechanisms+for+Disaggregated+LLM+Inference+Architecture
17. WindServe: Efficient Phase-Disaggregated LLM Serving with Stream-based Dynamic Scheduling — Jingqi Feng, Yukai Huang, Rui Zhang, Sicheng Liang, Ming Yan, Jie Wu, 2025
https://scholar.google.com/scholar?q=WindServe%3A+Efficient+Phase-Disaggregated+LLM+Serving+with+Stream-based+Dynamic+Scheduling
18. AI Post Transformers: CacheSlide: Position-Aware KV Cache Reuse for Agent LLMs — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-16-cacheslide-position-aware-kv-cache-reuse-cd59c7.mp3
19. AI Post Transformers: SGLang: Efficient Language Model Program Execution — Hal Turing & Dr. Ada Shannon, Sun,
https://podcast.do-not-panic.com/episodes/sglang-efficient-language-model-program-execution/
20. AI Post Transformers: Episode: Speculative Speculative Decoding — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-25-speculative-speculative-decoding-1b7a10.mp3
21. AI Post Transformers: xLLM: Co-Locating Online and Offline LLM Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-16-xllm-co-locating-online-and-offline-llm-10bb81.mp3
22. AI Post Transformers: FAST26: Bidaw: Enhancing Key-Value Caching for Interactive LLM Serving via Bidirectional — Hal Turing & Dr. Ada Shannon, Wed,
https://podcast.do-not-panic.com/episodes/fast26-bidaw-enhancing-key-value-caching-for-interactive-llm-serving-via-bidirec/
23. AI Post Transformers: Lookahead Q-Cache for Consistent KV Eviction — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-25-lookahead-q-cache-for-consistent-kv-evic-d97b09.mp3
Interactive Visualization: Splitwise: Phase-Split LLM Inference

This episode explores TurboQuant, a method for compressing high-dimensional vectors online without learning a dataset-specific codebook first, aimed at settings like LLM KV-cache compression and approximate nearest neighbor search. It explains why vector quantization is a different problem from ordinary weight quantization, and why preserving inner products can matter just as much as minimizing reconstruction error for retrieval quality and attention behavior. The discussion focuses on the paper’s central idea that a random rotation can regularize vectors enough for simple scalar quantization to approach information-theoretic distortion limits, at least under the paper’s theoretical assumptions. Listeners would find it interesting because it connects rate-distortion theory to concrete systems bottlenecks in modern AI, while also critically examining where the paper’s theoretical strength outpaces its empirical validation.

Sources:
1. TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate — Amir Zandieh, Majid Daliri, Majid Hadian, Vahab Mirrokni, 2025
http://arxiv.org/abs/2504.19874
2. Product Quantization for Nearest Neighbor Search — Herve Jegou, Matthijs Douze, Cordelia Schmid, 2011
https://scholar.google.com/scholar?q=Product+Quantization+for+Nearest+Neighbor+Search
3. Quantization based Fast Inner Product Search — Ruiqi Guo, Sanjiv Kumar, Krzysztof Choromanski, David Simcha, 2016
https://scholar.google.com/scholar?q=Quantization+based+Fast+Inner+Product+Search
4. Norm-Explicit Quantization: Improving Vector Quantization for Maximum Inner Product Search — Xinyan Dai, Xiao Yan, Kelvin K. W. Ng, Jiu Liu, James Cheng, 2020
https://scholar.google.com/scholar?q=Norm-Explicit+Quantization%3A+Improving+Vector+Quantization+for+Maximum+Inner+Product+Search
5. Accelerating Large-Scale Inference with Anisotropic Vector Quantization — Ruiqi Guo, Philip Sun, Erik Lindgren, Quan Geng, David Simcha, Felix Chern, Sanjiv Kumar, 2020
https://scholar.google.com/scholar?q=Accelerating+Large-Scale+Inference+with+Anisotropic+Vector+Quantization
6. QJL: 1-bit Quantized JL Transform for KV Cache Quantization with Zero Overhead — Amir Zandieh, Majid Daliri, Iman Han, 2024
https://scholar.google.com/scholar?q=QJL%3A+1-bit+Quantized+JL+Transform+for+KV+Cache+Quantization+with+Zero+Overhead
7. PolarQuant: Quantizing KV Caches with Polar Transformation — Iman Han, Prannay Kacham, Amin Karbasi, Vahab Mirrokni, Amir Zandieh, 2025
https://scholar.google.com/scholar?q=PolarQuant%3A+Quantizing+KV+Caches+with+Polar+Transformation
8. Practical and Asymptotically Optimal Quantization of High-Dimensional Vectors in Euclidean Space for Approximate Nearest Neighbor Search — Jianqiao Gao, Yuxuan Gou, Yiming Xu, Yuting Yang, Cheng Long, Raymond Chi-Wing Wong, 2024
https://scholar.google.com/scholar?q=Practical+and+Asymptotically+Optimal+Quantization+of+High-Dimensional+Vectors+in+Euclidean+Space+for+Approximate+Nearest+Neighbor+Search
9. KIVI: A Tuning-Free Asymmetric 2-bit Quantization for KV Cache — Zefan Liu, Jiapeng Yuan, Hongyin Jin, Shanghang Zhong, Zhiyuan Xu, Vladimir Braverman, Beidi Chen, Xia Hu, 2024
https://scholar.google.com/scholar?q=KIVI%3A+A+Tuning-Free+Asymmetric+2-bit+Quantization+for+KV+Cache
10. KVSink: Understanding and Enhancing the Preservation of Attention Sinks in KV Cache Quantization for LLMs — Zunhai Su, Kehong Yuan, 2025
https://scholar.google.com/scholar?q=KVSink%3A+Understanding+and+Enhancing+the+Preservation+of+Attention+Sinks+in+KV+Cache+Quantization+for+LLMs
11. ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification — Yefei He, Luoming Zhang, Weijia Wu, Jing Liu, Hong Zhou, Bohan Zhuang, 2024
https://scholar.google.com/scholar?q=ZipCache%3A+Accurate+and+Efficient+KV+Cache+Quantization+with+Salient+Token+Identification
12. AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models — Zunhai Su, Wang Shen, Linge Li, Zhe Chen, Hanyu Wei, Huangqi Yu, Kehong Yuan, 2025
https://scholar.google.com/scholar?q=AKVQ-VL%3A+Attention-Aware+KV+Cache+Adaptive+2-Bit+Quantization+for+Vision-Language+Models
13. SpinQuant: LLM Quantization with Learned Rotations — Zechun Liu, Changsheng Zhao, Igor Fedorov, Bilge Soran, Dhruv Choudhary, Raghuraman Krishnamoorthi, Vikas Chandra, Yuandong Tian, Tijmen Blankevoort, 2024
https://scholar.google.com/scholar?q=SpinQuant%3A+LLM+Quantization+with+Learned+Rotations
14. Rotate, Clip, and Partition: Towards W2A4KV4 Quantization by Integrating Rotation and Learnable Non-uniform Quantizer — Euntae Choi, Sumin Song, Woosang Lim, Sungjoo Yoo, 2025
https://scholar.google.com/scholar?q=Rotate%2C+Clip%2C+and+Partition%3A+Towards+W2A4KV4+Quantization+by+Integrating+Rotation+and+Learnable+Non-uniform+Quantizer
15. Locally-Adaptive Quantization for Streaming Vector Search — Cecilia Aguerrebere, Mark Hildebrand, Ishwar Singh Bhati, Theodore Willke, Mariano Tepper, 2024
https://scholar.google.com/scholar?q=Locally-Adaptive+Quantization+for+Streaming+Vector+Search
16. Sampling Methods for Inner Product Sketching — Majid Daliri, Juliana Freire, Christopher Musco, Aecio Santos, Haoxiang Zhang, 2024
https://scholar.google.com/scholar?q=Sampling+Methods+for+Inner+Product+Sketching
17. AI Post Transformers: Memory Traffic Saturation in Transformer Decode — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-20-memory-traffic-saturation-in-transformer-cd4961.mp3
18. AI Post Transformers: LAQ for Smarter KV Cache Eviction — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-23-laq-for-smarter-kv-cache-eviction-3ea2b8.mp3
19. AI Post Transformers: Lookahead Q-Cache for Consistent KV Eviction — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-25-lookahead-q-cache-for-consistent-kv-evic-d97b09.mp3
20. AI Post Transformers: AWQ: On-Device LLM Compression and Acceleration — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/awq-on-device-llm-compression-and-acceleration/
21. AI Post Transformers: Sentence-BERT: Siamese Networks for Sentence Embeddings — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/sentence-bert-siamese-networks-for-sentence-embeddings/
22. AI Post Transformers: SolidAttention: Co-Designing Sparse Attention and SSD I/O — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-18-solidattention-co-designing-sparse-atten-5a8622.mp3
Interactive Visualization: Episode: TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate

This episode explores the HyperAgents paper and its central claim that an AI system can improve not only its task behavior but also the procedure it uses to generate future improvements. It explains recursive self-improvement in practical terms as an outer engineering loop over prompts, code, tools, memory, and evaluators, and contrasts that with standard deep learning, where the learning process itself stays fixed. The discussion focuses on why freezing the meta-agent creates a conceptual and practical bottleneck, how HyperAgents try to remove that ceiling by making the improver editable too, and why that could matter beyond coding in domains like reviewing or grading. A listener would find it interesting for its clear debate over whether this is a genuine step toward more general self-improving agents or simply a cleaner packaging of familiar external scaffolding and control mechanisms.

Sources:
1. Hyperagents — Jenny Zhang, Bingchen Zhao, Wannan Yang, Jakob Foerster, Jeff Clune, Minqi Jiang, Sam Devlin, Tatiana Shavrina, 2026
http://arxiv.org/abs/2603.19461
2. Gödel Machines: Fully Self-Referential Optimal Universal Self-Improvers — Jürgen Schmidhuber, 2003
https://scholar.google.com/scholar?q=G%C3%B6del+Machines%3A+Fully+Self-Referential+Optimal+Universal+Self-Improvers
3. Self-Taught Optimizer (STOP): Recursively Self-Improving Code Generation — Eric Zelikman, Eliana Lorch, Lester Mackey, Adam Tauman Kalai, 2023
https://scholar.google.com/scholar?q=Self-Taught+Optimizer+%28STOP%29%3A+Recursively+Self-Improving+Code+Generation
4. Darwin Gödel Machine: Open-Ended Evolution of Self-Improving Agents — Jenny Zhang, Shengran Hu, Cong Lu, Robert Lange, Jeff Clune, 2025
https://scholar.google.com/scholar?q=Darwin+G%C3%B6del+Machine%3A+Open-Ended+Evolution+of+Self-Improving+Agents
5. AlphaEvolve: A Coding Agent for Scientific and Algorithmic Discovery — Alexander Novikov, Ngân Vũ, Marvin Eisenberger and colleagues, 2025
https://scholar.google.com/scholar?q=AlphaEvolve%3A+A+Coding+Agent+for+Scientific+and+Algorithmic+Discovery
6. Self-Referential Meta Learning — Louis Kirsch, Jürgen Schmidhuber, 2022
https://scholar.google.com/scholar?q=Self-Referential+Meta+Learning
7. A Modern Self-Referential Weight Matrix That Learns to Modify Itself — Kazuki Irie, Imanol Schlag, Róbert Csordás, Jürgen Schmidhuber, 2022
https://scholar.google.com/scholar?q=A+Modern+Self-Referential+Weight+Matrix+That+Learns+to+Modify+Itself
8. Gödel Agent: A Self-Referential Agent Framework for Recursive Self-Improvement — Xunjian Yin, Xinyi Wang, Liangming Pan, Xiaojun Wan, William Yang Wang, 2025
https://scholar.google.com/scholar?q=G%C3%B6del+Agent%3A+A+Self-Referential+Agent+Framework+for+Recursive+Self-Improvement
9. Darwin Gödel Machine — Jenny Zhang et al., 2025
https://scholar.google.com/scholar?q=Darwin+G%C3%B6del+Machine
10. Self-Referential AI Systems — Louis Kirsch and Jürgen Schmidhuber, 2022
https://scholar.google.com/scholar?q=Self-Referential+AI+Systems
11. Recursive Self-Improvement Can Be Bounded or Self-Accelerating — Lu et al., 2023
https://scholar.google.com/scholar?q=Recursive+Self-Improvement+Can+Be+Bounded+or+Self-Accelerating
12. Polyglot — Gauthier, 2024
https://scholar.google.com/scholar?q=Polyglot
13. Paper Review Benchmark — Zhao et al., 2026
https://scholar.google.com/scholar?q=Paper+Review+Benchmark
14. Genesis — Genesis authors, 2024
https://scholar.google.com/scholar?q=Genesis
15. Learning How to Remember: A Meta-Cognitive Management Method for Structured and Transferable Agent Memory — approx. unknown from snippet, recent
https://scholar.google.com/scholar?q=Learning+How+to+Remember%3A+A+Meta-Cognitive+Management+Method+for+Structured+and+Transferable+Agent+Memory
16. Memory in the Age of AI Agents — approx. unknown from snippet, recent
https://scholar.google.com/scholar?q=Memory+in+the+Age+of+AI+Agents
17. Discovering Hierarchical Software Engineering Agents via Bandit Optimization — approx. unknown from snippet, recent
https://scholar.google.com/scholar?q=Discovering+Hierarchical+Software+Engineering+Agents+via+Bandit+Optimization
18. BOAD: Discovering Hierarchical Software Engineering Agents via Bandit Optimization — approx. unknown from snippet, recent
https://scholar.google.com/scholar?q=BOAD%3A+Discovering+Hierarchical+Software+Engineering+Agents+via+Bandit+Optimization
19. Automated Design of Agentic Systems — approx. unknown from snippet, recent
https://scholar.google.com/scholar?q=Automated+Design+of+Agentic+Systems
20. AI Post Transformers: Experiential Reinforcement Learning: Internalizing Reflection for Better Policy Training — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/experiential-reinforcement-learning-internalizing-reflection-for-better-policy-t/
21. AI Post Transformers: LLM Agents Reason About Code Without Running It — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-15-llm-agents-reason-about-code-without-run-2a1876.mp3
22. AI Post Transformers: Mem0: Scalable Long-Term Memory for AI Agents — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/mem0-scalable-long-term-memory-for-ai-agents/
Interactive Visualization: HyperAgents and Metacognitive Self-Improvement

This episode explores a systems paper on speeding up retrieval-augmented generation by reusing transformer KV cache state more intelligently, instead of recomputing long retrieved prompts from scratch on every request. It explains why RAG often improves grounding yet suffers from high time-to-first-token, especially when multiple retrieved chunks must be prefetched and encoded together. The discussion focuses on the paper’s central argument that naive chunk-level cache reuse breaks important cross-chunk interactions, and that the proposed FusionRAG Cache tries to preserve quality through offline chunk enrichment and selective online recomputation. A listener would find it interesting because it connects familiar RAG concepts to the real serving bottlenecks that determine whether enterprise assistants feel practical or painfully slow.

Sources:
1. From Prefix Cache to Fusion RAG Cache: Accelerating LLM Inference in Retrieval-Augmented Generation — Jiahao Wang, Weiyu Xie, Mingxing Zhang, Boxing Zhang, Jianwei Dong, Yuening Zhu, Chen Lin, Jinqi Tang, Yaochen Han, Zhiyuan Ai, Xianglin Chen, Yongwei Wu, Congfeng Jiang, 2026
http://arxiv.org/abs/2601.12904v1
2. REALM: Retrieval-Augmented Language Model Pre-Training — Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, Ming-Wei Chang, 2020
https://scholar.google.com/scholar?q=REALM%3A+Retrieval-Augmented+Language+Model+Pre-Training
3. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks — Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Kuttler, Mike Lewis, Wen-tau Yih, Tim Rocktaschel, Sebastian Riedel, Douwe Kiela, 2020
https://scholar.google.com/scholar?q=Retrieval-Augmented+Generation+for+Knowledge-Intensive+NLP+Tasks
4. Few-shot Learning with Retrieval Augmented Language Models — Gautier Izacard, Edouard Grave, Patrick Lewis, Benjamin Chintala, Tim Rocktaschel, Fabio Petroni, 2022
https://scholar.google.com/scholar?q=Few-shot+Learning+with+Retrieval+Augmented+Language+Models
5. Retrieval-Augmented Generation for Large Language Models: A Survey — Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, Meng Wang, Haofen Wang, 2023
https://scholar.google.com/scholar?q=Retrieval-Augmented+Generation+for+Large+Language+Models%3A+A+Survey
6. APE: Faster and Longer Context-Augmented Generation via Adaptive Parallel Encoding — Xinyu Yang, Tianqi Chen, Beidi Chen, 2025
https://scholar.google.com/scholar?q=APE%3A+Faster+and+Longer+Context-Augmented+Generation+via+Adaptive+Parallel+Encoding
7. CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowledge Fusion — Jiayi Yao, Hanchen Li, Yuhan Liu, Siddhant Ray, Yihua Cheng, Qizheng Zhang, Kuntai Du, Shan Lu, Junchen Jiang, 2024
https://scholar.google.com/scholar?q=CacheBlend%3A+Fast+Large+Language+Model+Serving+for+RAG+with+Cached+Knowledge+Fusion
8. Cache-Craft: Managing Chunk-Caches for Efficient Retrieval-Augmented Generation — Shubham Agarwal, Sai Narayan Sundaresan, Subrata Mitra, Deb Mahapatra, Tong Yu, Shiv Saini, 2025
https://scholar.google.com/scholar?q=Cache-Craft%3A+Managing+Chunk-Caches+for+Efficient+Retrieval-Augmented+Generation
9. From Prefix Cache to Fusion RAG Cache: Accelerating LLM Inference in Retrieval-Augmented Generation — Jiahao Wang, Weiyu Xie, Mingxing Zhang, Boxing Zhang, Jianwei Dong, Yuening Zhu, Chen Lin, Jinqi Tang, Yaochen Han, Zhiyuan Ai, Xianglin Chen, Yongwei Wu, Congfeng Jiang, 2026
https://scholar.google.com/scholar?q=From+Prefix+Cache+to+Fusion+RAG+Cache%3A+Accelerating+LLM+Inference+in+Retrieval-Augmented+Generation
10. HyperRAG: Enhancing Quality-Efficiency Tradeoffs in Retrieval-Augmented Generation with Reranker KV-Cache Reuse — authors not reliably recoverable from the browsed snippets, 2025
https://scholar.google.com/scholar?q=HyperRAG%3A+Enhancing+Quality-Efficiency+Tradeoffs+in+Retrieval-Augmented+Generation+with+Reranker+KV-Cache+Reuse
11. Kvlink: Accelerating Large Language Models via Efficient KV Cache Reuse — approx. recent systems/LLM inference authors, 2025/2026
https://scholar.google.com/scholar?q=Kvlink%3A+Accelerating+Large+Language+Models+via+Efficient+KV+Cache+Reuse
12. Sparse Attention across Multiple-context KV Cache — approx. recent long-context/RAG inference authors, 2025/2026
https://scholar.google.com/scholar?q=Sparse+Attention+across+Multiple-context+KV+Cache
13. CacheClip: Accelerating RAG with Effective KV Cache Reuse — approx. recent RAG systems authors, 2025/2026
https://scholar.google.com/scholar?q=CacheClip%3A+Accelerating+RAG+with+Effective+KV+Cache+Reuse
14. CHESS: Context-aware Hierarchical Efficient Semantic Selection for Long-Context LLM Inference — approx. recent long-context inference authors, 2025/2026
https://scholar.google.com/scholar?q=CHESS%3A+Context-aware+Hierarchical+Efficient+Semantic+Selection+for+Long-Context+LLM+Inference
15. BudgetMem: Learning Selective Memory Policies for Cost-Efficient Long-Context Processing in Language Models — approx. recent memory-policy learning authors, 2025/2026
https://scholar.google.com/scholar?q=BudgetMem%3A+Learning+Selective+Memory+Policies+for+Cost-Efficient+Long-Context+Processing+in+Language+Models
16. TeleRAG: Efficient Retrieval-Augmented Generation Inference with Lookahead Retrieval — approx. recent RAG systems authors, 2025/2026
https://scholar.google.com/scholar?q=TeleRAG%3A+Efficient+Retrieval-Augmented+Generation+Inference+with+Lookahead+Retrieval
17. Understanding and Optimizing Multi-Stage AI Inference Pipelines — approx. recent systems authors, 2025/2026
https://scholar.google.com/scholar?q=Understanding+and+Optimizing+Multi-Stage+AI+Inference+Pipelines
18. AI Post Transformers: Episode: LookaheadKV: Fast and Accurate KV Cache Eviction by Glimpsing into the Future without Generation — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-21-lookaheadkv-fast-and-accurate-kv-c9d436.mp3
19. AI Post Transformers: CacheSlide: Position-Aware KV Cache Reuse for Agent LLMs — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-16-cacheslide-position-aware-kv-cache-reuse-cd59c7.mp3
20. AI Post Transformers: SolidAttention: Co-Designing Sparse Attention and SSD I/O — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-18-solidattention-co-designing-sparse-atten-5a8622.mp3
Interactive Visualization: Episode: From Prefix Cache to Fusion RAG Cache: Accelerating LLM Inference in Retrieval-Augmented Generation

This episode explores how KV cache eviction shapes the speed and usability of long-context language models, focusing on the March 2026 paper LookaheadKV and the broader problem of managing transformer memory under tight GPU budgets. It explains why KV caches are essential for autoregressive decoding, why their linear growth becomes a major inference bottleneck, and how eviction policies differ from related approaches such as cache compression. The discussion highlights the paper’s central argument: future-aware eviction can outperform simple recency-based heuristics, but only if it avoids the heavy latency costs that make some draft-generation methods impractical. A listener would find it interesting for its clear systems-level view of transformer inference, especially the tradeoff between smarter cache decisions and time-to-first-token in real production settings.

Sources:
1. LookaheadKV: Fast and Accurate KV Cache Eviction by Glimpsing into the Future without Generation — Jinwoo Ahn, Ingyu Seong, Akhil Kedia, Junhan Kim, Hyemi Jang, Kangwook Lee, Yongkweon Jeon, 2026
http://arxiv.org/abs/2603.10899v1
2. SnapKV: LLM Knows What You Are Looking for before Generation — Yuhong Li, Yingbing Huang, Bowen Yang, Bharat Venkitesh, Acyr Locatelli, Hanchen Ye, Tianle Cai, Patrick Lewis, Deming Chen, 2024
https://scholar.google.com/scholar?q=SnapKV%3A+LLM+Knows+What+You+Are+Looking+for+before+Generation
3. Lookahead Q-Cache: Achieving More Consistent KV Cache Eviction via Pseudo Query — Yixuan Wang, Shiyu Ji, Yijun Liu, Yuzhuang Xu, Yang Xu, Qingfu Zhu, Wanxiang Che, 2025
https://scholar.google.com/scholar?q=Lookahead+Q-Cache%3A+Achieving+More+Consistent+KV+Cache+Eviction+via+Pseudo+Query
4. KVzip: Query-Agnostic KV Cache Compression with Context Reconstruction — Jang-Hyun Kim, Jinuk Kim, Sangwoo Kwon, Jae W. Lee, Sangdoo Yun, Hyun Oh Song, 2025
https://scholar.google.com/scholar?q=KVzip%3A+Query-Agnostic+KV+Cache+Compression+with+Context+Reconstruction
5. PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling — Zefan Cai, Yichi Zhang, Bofei Gao, Yuliang Liu, Tianyu Liu, Keming Lu, Wayne Xiong, Yue Dong, Baobao Chang, Junjie Hu, Wen Xiao, 2024
https://scholar.google.com/scholar?q=PyramidKV%3A+Dynamic+KV+Cache+Compression+based+on+Pyramidal+Information+Funneling
6. HyperRAG: Enhancing Quality-Efficiency Tradeoffs in Retrieval-Augmented Generation with Reranker KV-Cache Reuse — approx. recent RAG systems authors, 2025/2026
https://scholar.google.com/scholar?q=HyperRAG%3A+Enhancing+Quality-Efficiency+Tradeoffs+in+Retrieval-Augmented+Generation+with+Reranker+KV-Cache+Reuse
7. KVLink: Accelerating Large Language Models via Efficient KV Cache Reuse — approx. recent LLM inference authors, 2025/2026
https://scholar.google.com/scholar?q=KVLink%3A+Accelerating+Large+Language+Models+via+Efficient+KV+Cache+Reuse
8. ProphetKV: User-Query-Driven Selective Recomputation for Efficient KV Cache Reuse in Retrieval-Augmented Generation — approx. recent RAG/inference authors, 2025/2026
https://scholar.google.com/scholar?q=ProphetKV%3A+User-Query-Driven+Selective+Recomputation+for+Efficient+KV+Cache+Reuse+in+Retrieval-Augmented+Generation
9. LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference — approx. recent long-context inference authors, 2024/2025
https://scholar.google.com/scholar?q=LazyLLM%3A+Dynamic+Token+Pruning+for+Efficient+Long+Context+LLM+Inference
10. TokenButler: Token Importance Is Predictable — approx. recent token pruning authors, 2025
https://scholar.google.com/scholar?q=TokenButler%3A+Token+Importance+Is+Predictable
11. LongHeads: Multi-Head Attention Is Secretly a Long Context Processor — approx. recent mechanistic interpretability / long-context authors, 2025
https://scholar.google.com/scholar?q=LongHeads%3A+Multi-Head+Attention+Is+Secretly+a+Long+Context+Processor
12. How Transformers Implement Induction Heads: Approximation and Optimization Analysis — approx. mechanistic interpretability authors, 2024/2025
https://scholar.google.com/scholar?q=How+Transformers+Implement+Induction+Heads%3A+Approximation+and+Optimization+Analysis
13. AI Post Transformers: LAQ for Smarter KV Cache Eviction — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-23-laq-for-smarter-kv-cache-eviction-3ea2b8.mp3
14. AI Post Transformers: Lookahead Q-Cache for Consistent KV Eviction — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-25-lookahead-q-cache-for-consistent-kv-evic-d97b09.mp3
15. AI Post Transformers: Memory Traffic Saturation in Transformer Decode — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-20-memory-traffic-saturation-in-transformer-cd4961.mp3
16. AI Post Transformers: SolidAttention: Co-Designing Sparse Attention and SSD I/O — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-18-solidattention-co-designing-sparse-atten-5a8622.mp3
Interactive Visualization: Episode: LookaheadKV: Fast and Accurate KV Cache Eviction by Glimpsing into the Future without Generation

In this episode, the hosts examine LeWorldModel, a 2026 paper from researchers across Mila, Universite de Montreal, NYU, Samsung SAIL, and Brown that asks whether a joint-embedding predictive architecture can finally be trained end to end from raw pixels without the usual stack of stabilizers. The discussion situates the work in the broader world-model lineage from Ha and Schmidhuber through Dreamer, and in the JEPA program associated with Yann LeCun’s predictive-learning agenda. The paper’s core claim is unusually narrow and concrete: a small model, around 15 million parameters, can learn action-conditioned dynamics directly from images using just next-embedding prediction plus a Gaussian latent regularizer called SIGReg, avoiding EMA teachers, pretrained encoders, reconstruction losses, and other auxiliary machinery that many related systems rely on. The conversation focuses on why that claim matters. The hosts explain that JEPA-style methods are attractive because they predict semantic embeddings rather than reconstructing every pixel, but they have been plagued by representation collapse and fragile training recipes. Most of the technical attention therefore goes to how LeWorldModel tries to keep the latent space informative while staying simple enough to train jointly on a single GPU in a few hours. They walk through the paper’s framing around offline control and latent-space planning, where forecasting compact future states can make imagined rollouts cheap. They also discuss the project-page claim that LeWorldModel can plan up to roughly 48 times faster than DINO-WM because each frame is compressed to a single 192-dimensional token, while noting that this is part of the system’s pitch and should be separated from broader claims about downstream capability. The episode also digs into the evidence and the limits. The hosts cover the benchmark results across Two-Room, Reacher, Push-T, and OGBench-Cube, where LeWorldModel appears competitive overall, broadly stronger than PLDM, and better than DINO-WM on Push-T and Reacher, while DINO-WM still looks stronger on the more visually complex OGBench-Cube setting, likely because richer pretrained visual priors still help there. They also discuss the paper and project-page attempts to show “physical understanding” through latent probing, decoded latent visualizations, and surprise-style tests for implausible events. Throughout, the description stays skeptical about the gap between paper evidence and project-page marketing: the system looks like a real simplification of JEPA world modeling, but not yet a final verdict that minimal end-to-end predictive learning has solved robust visual control.

Sources:
1. LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels — Lucas Maes, Quentin Le Lidec, Damien Scieur, Yann LeCun, Randall Balestriero, 2026
http://arxiv.org/abs/2603.19312v1
2. World Models — David Ha, Jürgen Schmidhuber, 2018
https://scholar.google.com/scholar?q=World+Models
3. Learning Latent Dynamics for Planning from Pixels — Danijar Hafner, Timothy Lillicrap, Ian Fischer, Ruben Villegas, David Ha, Honglak Lee, James Davidson, 2018
https://scholar.google.com/scholar?q=Learning+Latent+Dynamics+for+Planning+from+Pixels
4. Dream to Control: Learning Behaviors by Latent Imagination — Danijar Hafner, Timothy Lillicrap, Jimmy Ba, Mohammad Norouzi, 2019
https://scholar.google.com/scholar?q=Dream+to+Control%3A+Learning+Behaviors+by+Latent+Imagination
5. TD-MPC2: Scalable, Robust World Models for Continuous Control — Nicklas Hansen, Hao Su, Xiaolong Wang, 2024
https://scholar.google.com/scholar?q=TD-MPC2%3A+Scalable%2C+Robust+World+Models+for+Continuous+Control
6. A Path Towards Autonomous Machine Intelligence — Yann LeCun, 2022
https://scholar.google.com/scholar?q=A+Path+Towards+Autonomous+Machine+Intelligence
7. Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture — Mahmoud Assran, Quentin Duval, Ishan Misra, Piotr Bojanowski, Pascal Vincent, Michael Rabbat, Yann LeCun, Nicolas Ballas, 2023
https://scholar.google.com/scholar?q=Self-Supervised+Learning+from+Images+with+a+Joint-Embedding+Predictive+Architecture
8. Revisiting Feature Prediction for Learning Visual Representations from Video — Adrien Bardes, Quentin Garrido, Jean Ponce, Xinlei Chen, Michael Rabbat, Yann LeCun, Mahmoud Assran, Nicolas Ballas, 2024
https://scholar.google.com/scholar?q=Revisiting+Feature+Prediction+for+Learning+Visual+Representations+from+Video
9. LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels — Lucas Maes, Quentin Le Lidec, Damien Scieur, Yann LeCun, Randall Balestriero, 2026
https://scholar.google.com/scholar?q=LeWorldModel%3A+Stable+End-to-End+Joint-Embedding+Predictive+Architecture+from+Pixels
10. Bootstrap Your Own Latent: A New Approach to Self-Supervised Learning — Jean-Bastien Grill, Florian Strub, Florent Altche, Corentin Tallec and others, 2020
https://scholar.google.com/scholar?q=Bootstrap+Your+Own+Latent%3A+A+New+Approach+to+Self-Supervised+Learning
11. Exploring Simple Siamese Representation Learning — Xinlei Chen, Kaiming He, 2021
https://scholar.google.com/scholar?q=Exploring+Simple+Siamese+Representation+Learning
12. VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised Learning — Adrien Bardes, Jean Ponce, Yann LeCun, 2021
https://scholar.google.com/scholar?q=VICReg%3A+Variance-Invariance-Covariance+Regularization+for+Self-Supervised+Learning
13. Learning from Reward-Free Offline Data: A Case for Planning with Latent Dynamics Models — Vlad Sobal, Wancong Zhang, Kyunghyun Cho, Randall Balestriero, Tim G. J. Rudner, Yann LeCun, 2025
https://scholar.google.com/scholar?q=Learning+from+Reward-Free+Offline+Data%3A+A+Case+for+Planning+with+Latent+Dynamics+Models
14. Human-level Control through Deep Reinforcement Learning — Volodymyr Mnih, Koray Kavukcuoglu, David Silver and others, 2015
https://scholar.google.com/scholar?q=Human-level+Control+through+Deep+Reinforcement+Learning
15. Mastering Diverse Domains through World Models — Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, Timothy Lillicrap, 2023
https://scholar.google.com/scholar?q=Mastering+Diverse+Domains+through+World+Models
16. TD-MPC: Learning to Plan in Latent Space for Visual Control — Nicklas Hansen, Xiaolong Wang, Hao Su, 2022
https://scholar.google.com/scholar?q=TD-MPC%3A+Learning+to+Plan+in+Latent+Space+for+Visual+Control
17. PlaNet: Learning Latent Dynamics for Planning from Pixels — Danijar Hafner, Timothy Lillicrap, Jimmy Ba, Mohammad Norouzi, 2019
https://scholar.google.com/scholar?q=PlaNet%3A+Learning+Latent+Dynamics+for+Planning+from+Pixels
18. Rethinking JEPA: Compute-Efficient Video Self-Supervised Learning with Frozen Teachers — approx. unknown from snippet, 2025-2026
https://scholar.google.com/scholar?q=Rethinking+JEPA%3A+Compute-Efficient+Video+Self-Supervised+Learning+with+Frozen+Teachers
19. Efficient reinforcement learning through adaptively pretrained visual encoder — approx. unknown from snippet, 2023-2026
https://scholar.google.com/scholar?q=Efficient+reinforcement+learning+through+adaptively+pretrained+visual+encoder
20. Adaworld: Learning adaptable world models with latent actions — approx. unknown from snippet, 2023-2026
https://scholar.google.com/scholar?q=Adaworld%3A+Learning+adaptable+world+models+with+latent+actions
21. AI Post Transformers: Unified Latents (UL): How to train your latents — Hal Turing & Dr. Ada Shannon, Sat,
https://podcast.do-not-panic.com/episodes/unified-latents-ul-how-to-train-your-latents/
22. AI Post Transformers: Contrastive Behavioral Similarity Embeddings for Generalization in Reinforcement Learning — Hal Turing & Dr. Ada Shannon, Fri,
https://podcast.do-not-panic.com/episodes/contrastive-behavioral-similarity-embeddings-for-generalization-in-reinforcement/
Interactive Visualization: LeWorldModel: Stable Joint-Embedding World Models from Pixels

This episode explores TurboQuant, a method for online vector quantization that aims to compress high-dimensional embeddings and transformer KV caches without any pretrained codebook or calibration pass. It explains how the paper connects classical rate-distortion theory to practical ML systems, contrasting mean-squared reconstruction error with inner-product preservation for tasks like retrieval and attention. The discussion highlights TurboQuant’s core idea of using random rotations and quantized Johnson-Lindenstrauss style sketches to make data-oblivious compression theoretically strong while still relevant to modern workloads. A listener would find it interesting because it probes whether a single, plug-and-play quantization scheme can approach information-theoretic limits while addressing real memory and bandwidth bottlenecks in large-scale AI systems.

Sources:
1. TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate — Amir Zandieh, Majid Daliri, Majid Hadian, Vahab Mirrokni, 2025
http://arxiv.org/abs/2504.19874
2. Vector Quantization — Robert M. Gray, 1984
https://scholar.google.com/scholar?q=Vector+Quantization
3. Product Quantization for Nearest Neighbor Search — Herve Jegou, Matthijs Douze, Cordelia Schmid, 2011
https://scholar.google.com/scholar?q=Product+Quantization+for+Nearest+Neighbor+Search
4. Optimized Product Quantization — Tiezheng Ge, Kaiming He, Qifa Ke, Jian Sun, 2013
https://scholar.google.com/scholar?q=Optimized+Product+Quantization
5. Additive Quantization for Extreme Vector Compression — Artem Babenko, Victor Lempitsky, 2014
https://scholar.google.com/scholar?q=Additive+Quantization+for+Extreme+Vector+Compression
6. Online Product Quantization — Donna Xu, Ivor W. Tsang, Ying Zhang, 2018
https://scholar.google.com/scholar?q=Online+Product+Quantization
7. Online Optimized Product Quantization — Chao Liu, Defu Lian, Min Nie, Huabin Xia, 2020
https://scholar.google.com/scholar?q=Online+Optimized+Product+Quantization
8. Locally-Adaptive Quantization for Streaming Vector Search — Cecilia Aguerrebere, Mark Hildebrand, Ishwar Singh Bhati, Theodore L. Willke, Mariano Tepper, 2024
https://scholar.google.com/scholar?q=Locally-Adaptive+Quantization+for+Streaming+Vector+Search
9. Real-time Indexing for Large-scale Recommendation by Streaming Vector Quantization Retriever — Xingyan Bin, Jianfei Cui, Wujie Yan, Zhichen Zhao, Xintian Han, Chongyang Yan, Feng Zhang, Xun Zhou, Qi Wu, Zuotao Liu, 2025
https://scholar.google.com/scholar?q=Real-time+Indexing+for+Large-scale+Recommendation+by+Streaming+Vector+Quantization+Retriever
10. Coding Theorems for a Discrete Source with a Fidelity Criterion — Claude E. Shannon, 1959
https://scholar.google.com/scholar?q=Coding+Theorems+for+a+Discrete+Source+with+a+Fidelity+Criterion
11. Rate Distortion Theory: A Mathematical Basis for Data Compression — Thomas Berger, 1971
https://scholar.google.com/scholar?q=Rate+Distortion+Theory%3A+A+Mathematical+Basis+for+Data+Compression
12. The Information Bottleneck Method — Naftali Tishby, Fernando C. Pereira, William Bialek, 1999
https://scholar.google.com/scholar?q=The+Information+Bottleneck+Method
13. End-to-end Optimized Image Compression — Johannes Balle, Valero Laparra, Eero P. Simoncelli, 2017
https://scholar.google.com/scholar?q=End-to-end+Optimized+Image+Compression
14. Similarity Estimation Techniques from Rounding Algorithms — Moses S. Charikar, 2002
https://scholar.google.com/scholar?q=Similarity+Estimation+Techniques+from+Rounding+Algorithms
15. A Quantized Johnson-Lindenstrauss Lemma: The Finding of Buffon's Needle — Laurent Jacques, 2015
https://scholar.google.com/scholar?q=A+Quantized+Johnson-Lindenstrauss+Lemma%3A+The+Finding+of+Buffon%27s+Needle
16. Quantized Random Projections and Non-Linear Estimation of Cosine Similarity — Ping Li, Michael Mitzenmacher, Martin Slawski, 2016
https://scholar.google.com/scholar?q=Quantized+Random+Projections+and+Non-Linear+Estimation+of+Cosine+Similarity
17. QJL: 1-Bit Quantized JL Transform for KV Cache Quantization with Zero Overhead — Amir Zandieh, Majid Daliri, Insu Han, 2025
https://scholar.google.com/scholar?q=QJL%3A+1-Bit+Quantized+JL+Transform+for+KV+Cache+Quantization+with+Zero+Overhead
18. PolarQuant: Quantizing KV Caches with Polar Transformation — Iman Han, Praveen Kacham, Amin Karbasi, Vahab Mirrokni, and Amir Zandieh, 2025
https://scholar.google.com/scholar?q=PolarQuant%3A+Quantizing+KV+Caches+with+Polar+Transformation
19. Practical and Asymptotically Optimal Quantization of High-Dimensional Vectors in Euclidean Space for Approximate Nearest Neighbor Search — Jianqiu Gao, Yuxuan Gou, Yifan Xu, Yifan Yang, Cheng Long, and Raymond Chi-Wing Wong, 2024
https://scholar.google.com/scholar?q=Practical+and+Asymptotically+Optimal+Quantization+of+High-Dimensional+Vectors+in+Euclidean+Space+for+Approximate+Nearest+Neighbor+Search
20. Accelerating Large-Scale Inference with Anisotropic Vector Quantization — Ruoming Guo, Philip Sun, Erik Lindgren, Quan Geng, David Simcha, Felix Chern, and Sanjiv Kumar, 2020
https://scholar.google.com/scholar?q=Accelerating+Large-Scale+Inference+with+Anisotropic+Vector+Quantization
21. KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache — Zihao Liu, Jun Yuan, Haotian Jin, Sheng Zhong, Ziqi Xu, Vladimir Braverman, Beidi Chen, and Xia Hu, 2024
https://scholar.google.com/scholar?q=KIVI%3A+A+Tuning-Free+Asymmetric+2bit+Quantization+for+KV+Cache
22. HyperRAG: Enhancing Quality-Efficiency Tradeoffs in Retrieval-Augmented Generation with Reranker KV-Cache Reuse — approx. recent RAG / LLM systems authors, 2025
https://scholar.google.com/scholar?q=HyperRAG%3A+Enhancing+Quality-Efficiency+Tradeoffs+in+Retrieval-Augmented+Generation+with+Reranker+KV-Cache+Reuse
23. KVLink: Accelerating Large Language Models via Efficient KV Cache Reuse — approx. recent LLM systems authors, 2025
https://scholar.google.com/scholar?q=KVLink%3A+Accelerating+Large+Language+Models+via+Efficient+KV+Cache+Reuse
24. Cache-Craft: Managing Chunk-Caches for Efficient Retrieval-Augmented Generation — approx. recent RAG systems authors, 2025
https://scholar.google.com/scholar?q=Cache-Craft%3A+Managing+Chunk-Caches+for+Efficient+Retrieval-Augmented+Generation
25. Weighted Minwise Hashing Beats Linear Sketching for Inner Product Estimation — approx. sketching / similarity estimation authors, recent
https://scholar.google.com/scholar?q=Weighted+Minwise+Hashing+Beats+Linear+Sketching+for+Inner+Product+Estimation
26. AI Post Transformers: Memory Traffic Saturation in Transformer Decode — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-20-memory-traffic-saturation-in-transformer-cd4961.mp3
27. AI Post Transformers: Quantizing Diffusion LLMs: A Systematic Study — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/quantizing-diffusion-llms-a-systematic-study/
28. AI Post Transformers: Limitations of Embedding-Based Retrieval — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/limitations-of-embedding-based-retrieval/

This episode examines Lookahead Q-Cache as a very specific kind of inference optimization: a decode-stage KV-cache eviction method for long-context serving. The discussion explains the paper’s core claim that prefill-time attention is a weak proxy for what will matter once generation actually begins, because decode-time queries are conditioned on the answer the model is actively writing rather than on the prompt alone. That is the real novelty here. Static selection methods such as SnapKV and related heavy-hitter or cumulative-attention schemes mostly infer importance from prompt-side attention patterns, often using a suffix window as a stand-in for future need. Lookahead Q-Cache instead uses a pseudo-query to approximate upcoming decode queries, making eviction more dynamic and more aligned to generation. The hosts are explicit that this is mostly a decode-only idea, not a general cure for transformer inference cost, and they keep returning to that point so the scope is not overstated. The conversation places the paper inside the broader acceleration landscape rather than treating it as a standalone breakthrough. Speculative decoding, Medusa-style multi-head prediction, and layered drafting ideas such as inference blending or Matryoshka-like speculative schemes all attack a different bottleneck: they try to reduce the cost of producing future tokens by drafting and verifying them more efficiently. Lookahead Q-Cache attacks the memory and attention burden of carrying long prefixes during decode. Those are not the same problem, which means they are not simple substitutes and can in principle be complementary in one serving stack. The episode also contrasts this test-time cache-management line with architecture-level efficiency work such as grouped-query attention, Nemotron 3 style system-model co-design, and Kimi-like efficient long-context efforts, where the gains often come from changing the model or attention structure rather than making smarter runtime eviction decisions. The tone stays skeptical about deployment significance. The hosts ask the hard scaling question directly: does smarter KV eviction materially change long-context serving economics, or does it mainly deliver narrower decode wins inside a larger bottleneck stack that still includes prefill cost, bandwidth pressure, scheduler behavior, batching constraints, quantization tradeoffs, and model architecture limits? They argue that benchmark improvements in eviction consistency are interesting, but the real bar is whether operators would trust aggressive dynamic cache pruning in production compared with more predictable approaches like GQA, FlashAttention, quantization, or speculative decode pipelines already discussed elsewhere on the podcast. The result is a grounded episode about what is genuinely new in Lookahead Q-Cache, where it fits, and why decode-specific cache tricks should not be confused with a full solution to long-context serving.

Sources:
1. Lookahead Q-Cache: Achieving More Consistent KV Cache Eviction via Pseudo Query — Yixuan Wang, Shiyu Ji, Yijun Liu, Yuzhuang Xu, Yang Xu, Qingfu Zhu, Wanxiang Che, 2025
http://arxiv.org/abs/2505.20334
2. SnapKV: LLM Knows What You are Looking for Before Generation — Zhenyu Li et al., 2024
https://scholar.google.com/scholar?q=SnapKV%3A+LLM+Knows+What+You+are+Looking+for+Before+Generation
3. H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models — Zhenyu Zhang et al., 2023
https://scholar.google.com/scholar?q=H2O%3A+Heavy-Hitter+Oracle+for+Efficient+Generative+Inference+of+Large+Language+Models
4. Scissorhands: Exploiting the Persistence of Importance Hypothesis for LLM KV Cache Compression at Test Time — Zhenyu Liu et al., 2023
https://scholar.google.com/scholar?q=Scissorhands%3A+Exploiting+the+Persistence+of+Importance+Hypothesis+for+LLM+KV+Cache+Compression+at+Test+Time
5. Fast and Accurate Transformer Decoding via Dynamic Compression of KV Cache — likely Tang et al., 2024
https://scholar.google.com/scholar?q=Fast+and+Accurate+Transformer+Decoding+via+Dynamic+Compression+of+KV+Cache
6. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness — Tri Dao et al., 2022
https://scholar.google.com/scholar?q=FlashAttention%3A+Fast+and+Memory-Efficient+Exact+Attention+with+IO-Awareness
7. FlashAttention-2 or subsequent FlashAttention work — Tri Dao, 2024
https://scholar.google.com/scholar?q=FlashAttention-2+or+subsequent+FlashAttention+work
8. RazorAttention: Efficient KV Cache Compression through Retrieval Heads — not recovered from snippet, 2024-2025
https://scholar.google.com/scholar?q=RazorAttention%3A+Efficient+KV+Cache+Compression+through+Retrieval+Heads
9. FastKV: KV Cache Compression for Fast Long-Context Processing with Token-Selective Propagation — not recovered from snippet, 2024-2025
https://scholar.google.com/scholar?q=FastKV%3A+KV+Cache+Compression+for+Fast+Long-Context+Processing+with+Token-Selective+Propagation
10. Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity — not recovered from snippet, 2024-2025
https://scholar.google.com/scholar?q=Compressing+KV+Cache+for+Long-Context+LLM+Inference+with+Inter-Layer+Attention+Similarity
11. KVLink: Accelerating Large Language Models via Efficient KV Cache Reuse — not recovered from snippet, 2024-2025
https://scholar.google.com/scholar?q=KVLink%3A+Accelerating+Large+Language+Models+via+Efficient+KV+Cache+Reuse
12. DepCache: A KV Cache Management Framework for GraphRAG with Dependency Attention — not recovered from snippet, 2024-2025
https://scholar.google.com/scholar?q=DepCache%3A+A+KV+Cache+Management+Framework+for+GraphRAG+with+Dependency+Attention
13. End-to-End Acceleration of Generative Models with Runtime Regularized KV Cache Management — not recovered from snippet, 2024-2025
https://scholar.google.com/scholar?q=End-to-End+Acceleration+of+Generative+Models+with+Runtime+Regularized+KV+Cache+Management
14. AI Post Transformers: LAQ for Smarter KV Cache Eviction — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-23-laq-for-smarter-kv-cache-eviction-3ea2b8.mp3
15. AI Post Transformers: LookaheadKV: Fast and Accurate KV Cache Eviction by Glimpsing into the Future without Generation — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-21-lookaheadkv-fast-and-accurate-kv-c9d436.mp3
16. AI Post Transformers: Quest: Query-Aware Sparsity for Efficient LLM Inference — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/quest-query-aware-sparsity-for-efficient-llm-inference/
17. AI Post Transformers: Hyper-Scaling LLM Inference with KV Cache Compression — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/hyper-scaling-llm-inference-with-kv-cache-compression/
18. AI Post Transformers: Memory Traffic Saturation in Transformer Decode — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-20-memory-traffic-saturation-in-transformer-cd4961.mp3
Interactive Visualization: Lookahead Q-Cache for Consistent KV Eviction

This episode explores Speculative Speculative Decoding, a technique for reducing LLM inference latency by overlapping drafting and verification more aggressively than standard speculative decoding. It explains how the method predicts likely verification outcomes in advance so the draft model can prepare multiple next-step continuations while the target model is still checking the current block, with the target model still preserving the final output. The discussion focuses on the Saguaro algorithm, the distinction between true parallel generation and latency-hiding around autoregressive dependencies, and the practical tradeoff between useful overlap and wasted draft-side computation. A listener would find it interesting for its clear look at where modern inference systems still lose time and how smarter scheduling, rather than changing model semantics, can unlock additional speed.

Sources:
1. Speculative Speculative Decoding — Tanishq Kumar, Tri Dao, Avner May, 2026
http://arxiv.org/abs/2603.03251
2. Fast Inference from Transformers via Speculative Decoding — Yaniv Leviathan, Matan Kalman, Yossi Matias, 2023
https://scholar.google.com/scholar?q=Fast+Inference+from+Transformers+via+Speculative+Decoding
3. EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty — Yuhui Li, Fangyun Wei, Chao Zhang, Hongyang Zhang, 2024
https://scholar.google.com/scholar?q=EAGLE%3A+Speculative+Sampling+Requires+Rethinking+Feature+Uncertainty
4. SpecInfer: Accelerating Generative LLM Serving with Speculative Inference and Token Tree Verification — Xupeng Miao, Gabriele Oliaro, Zhihao Zhang, Xinhao Cheng, Zeyu Wang, Rae Ying Yee Wong, Zhuoming Chen, Daiyaan Arfeen, Reyna Abhyankar, Zhihao Jia, 2024
https://scholar.google.com/scholar?q=SpecInfer%3A+Accelerating+Generative+LLM+Serving+with+Speculative+Inference+and+Token+Tree+Verification
5. PEARL: Parallel Speculative Decoding with Adaptive Draft Length — Tianyu Liu, Yun Li, Qitan Lv, Kai Liu, Jianchen Zhu, Winston Hu, Xiao Sun, 2025
https://scholar.google.com/scholar?q=PEARL%3A+Parallel+Speculative+Decoding+with+Adaptive+Draft+Length
6. AMUSD: Asynchronous Multi-Device Speculative Decoding for LLM Acceleration — Bradley McDanel, 2025
https://scholar.google.com/scholar?q=AMUSD%3A+Asynchronous+Multi-Device+Speculative+Decoding+for+LLM+Acceleration
7. Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding — Jun Zhang et al., 2023
https://scholar.google.com/scholar?q=Draft+%26+Verify%3A+Lossless+Large+Language+Model+Acceleration+via+Self-Speculative+Decoding
8. SWIFT: On-the-Fly Self-Speculative Decoding for LLM Inference Acceleration — Heming Xia et al., 2024
https://scholar.google.com/scholar?q=SWIFT%3A+On-the-Fly+Self-Speculative+Decoding+for+LLM+Inference+Acceleration
9. SpecDec++: Boosting Speculative Decoding via Adaptive Candidate Lengths — Kaixuan Huang, Xudong Guo, Mengdi Wang, 2024
https://scholar.google.com/scholar?q=SpecDec%2B%2B%3A+Boosting+Speculative+Decoding+via+Adaptive+Candidate+Lengths
10. Accelerating Transformer Inference for Translation via Parallel Decoding — Andrea Santilli et al., 2023
https://scholar.google.com/scholar?q=Accelerating+Transformer+Inference+for+Translation+via+Parallel+Decoding
11. Self-Speculative Decoding for LLM-based ASR with CTC Encoder Drafts — George Saon et al., 2026
https://scholar.google.com/scholar?q=Self-Speculative+Decoding+for+LLM-based+ASR+with+CTC+Encoder+Drafts
12. AI Post Transformers: Apple's Speculative Streaming: Fast LLM Inference without Auxiliary Models — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/apples-speculative-streaming-fast-llm-inference-without-auxiliary-models/
13. AI Post Transformers: Adaptive Control for Batched Speculative Decoding in LLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/adaptive-control-for-batched-speculative-decoding-in-llm-serving/
14. AI Post Transformers: FastGRPO: Concurrency-Aware Speculative Decoding for Policy Optimization — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/fastgrpo-concurrency-aware-speculative-decoding-for-policy-optimization/
Interactive Visualization: Episode: Speculative Speculative Decoding

This episode explores how speculative decoding can itself be pipelined, using the March 2026 paper "Speculative Speculative Decoding" to examine whether drafting and verification in LLM inference can be overlapped instead of run in a stop-and-go loop. It explains the bottleneck of standard autoregressive generation, reviews classic speculative decoding as introduced by Leviathan et al., and then focuses on the paper’s key idea: predicting likely verification outcomes so the next draft is ready before the verifier finishes. The discussion frames this as a scheduling and systems optimization problem rather than a new model architecture, connecting it to related work such as Lookahead Decoding, Medusa, and EAGLE. Listeners would find it interesting because it shows how careful inference-time execution design can deliver major practical speedups, including roughly a 30 percent average gain over strong speculative decoding baselines in the paper’s optimized Saguaro system.

Sources:
1. Speculative Speculative Decoding — Tanishq Kumar, Tri Dao, Avner May, 2026
http://arxiv.org/abs/2603.03251
2. Fast Inference from Transformers via Speculative Decoding — Yaniv Leviathan, Matan Kalman, Yossi Matias, 2023
https://scholar.google.com/scholar?q=Fast+Inference+from+Transformers+via+Speculative+Decoding
3. Break the Sequential Dependency of LLM Inference Using Lookahead Decoding — Yichao Fu, Peter Bailis, Ion Stoica, Hao Zhang, 2024
https://scholar.google.com/scholar?q=Break+the+Sequential+Dependency+of+LLM+Inference+Using+Lookahead+Decoding
4. Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads — Tianle Cai, Yuhong Li, Zhengyang Geng, Hongwu Peng, Jason D. Lee, Deming Chen, Tri Dao, 2024
https://scholar.google.com/scholar?q=Medusa%3A+Simple+LLM+Inference+Acceleration+Framework+with+Multiple+Decoding+Heads
5. Speculative Speculative Decoding — Tanishq Kumar, Tri Dao, Avner May, 2026
https://scholar.google.com/scholar?q=Speculative+Speculative+Decoding
6. EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty — Yuhui Li, Fangyun Wei, Chao Zhang, Hongyang Zhang, 2024
https://scholar.google.com/scholar?q=EAGLE%3A+Speculative+Sampling+Requires+Rethinking+Feature+Uncertainty
7. Hydra: Sequentially-Dependent Draft Heads for Medusa Decoding — Zachary Ankner, Rishab Parthasarathy, Aniruddha Nrusimha, Christopher Rinard, Jonathan Ragan-Kelley, William Brandon, 2024
https://scholar.google.com/scholar?q=Hydra%3A+Sequentially-Dependent+Draft+Heads+for+Medusa+Decoding
8. Accelerating Production LLMs with Combined Token/Embedding Speculators — Davis Wertheimer, Joshua Rosenkranz, Thomas Parnell, Sahil Suneja, Pavithra Ranganathan, Raghu Ganti, Mudhakar Srivatsa, 2024
https://scholar.google.com/scholar?q=Accelerating+Production+LLMs+with+Combined+Token%2FEmbedding+Speculators
9. Dynamic Speculation Lookahead Accelerates Speculative Decoding of Large Language Models — Jonathan Mamou, Oren Pereg, Daniel Korat, Moshe Berchansky, Nadav Timor, Moshe Wasserblat, Roy Schwartz, 2024
https://scholar.google.com/scholar?q=Dynamic+Speculation+Lookahead+Accelerates+Speculative+Decoding+of+Large+Language+Models
10. Accelerating Large Language Model Decoding with Speculative Sampling — Charlie Chen, Sebastian Borgeaud, Geoffrey Irving, Jean-Baptiste Lespiau, Laurent Sifre, John Jumper, 2023
https://scholar.google.com/scholar?q=Accelerating+Large+Language+Model+Decoding+with+Speculative+Sampling
11. SpecInfer: Accelerating Large Language Model Serving with Tree-Based Speculative Inference and Verification — Xupeng Miao, Gabriele Oliaro, Zhihao Zhang, Xinhao Cheng, Zeyu Wang, Zhengxin Zhang, Rae Ying Yee Wong, Alan Zhu, Lijie Yang, Xiaoxiang Shi, Chunan Shi, Zhuoming Chen, Daiyaan Arfeen, Reyna Abhyankar, Zhihao Jia, 2024
https://scholar.google.com/scholar?q=SpecInfer%3A+Accelerating+Large+Language+Model+Serving+with+Tree-Based+Speculative+Inference+and+Verification
12. EAGLE-3: Scaling Up Inference Acceleration of Large Language Models via Training-Time Test — Yuhui Li, Fangyun Wei, Chao Zhang, Hongyang Zhang, 2025
https://scholar.google.com/scholar?q=EAGLE-3%3A+Scaling+Up+Inference+Acceleration+of+Large+Language+Models+via+Training-Time+Test
13. AMUSD: Asynchronous Multi-Device Speculative Decoding for LLM Acceleration — Bradley McDanel, 2025
https://scholar.google.com/scholar?q=AMUSD%3A+Asynchronous+Multi-Device+Speculative+Decoding+for+LLM+Acceleration
14. DistServe: Disaggregating Prefill and Decoding for Goodput-Optimized Large Language Model Serving — Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, Hao Zhang, 2024
https://scholar.google.com/scholar?q=DistServe%3A+Disaggregating+Prefill+and+Decoding+for+Goodput-Optimized+Large+Language+Model+Serving
15. Layerskip: Enabling early exit inference and self-speculative decoding — not verified from snippet, 2024/2025
https://scholar.google.com/scholar?q=Layerskip%3A+Enabling+early+exit+inference+and+self-speculative+decoding
16. Self-Speculative Decoding Accelerates Lossless Inference in Any-Order and Any-Subset Autoregressive Models — not verified from snippet, 2025
https://scholar.google.com/scholar?q=Self-Speculative+Decoding+Accelerates+Lossless+Inference+in+Any-Order+and+Any-Subset+Autoregressive+Models
17. Speculative Verification: Exploiting Information Gain to Refine Speculative Decoding — not verified from snippet, 2025
https://scholar.google.com/scholar?q=Speculative+Verification%3A+Exploiting+Information+Gain+to+Refine+Speculative+Decoding
18. Parallelspec: Parallel drafter for efficient speculative decoding — not verified from snippet, 2025
https://scholar.google.com/scholar?q=Parallelspec%3A+Parallel+drafter+for+efficient+speculative+decoding
19. Opt-tree: Speculative decoding with adaptive draft tree structure — not verified from snippet, 2024/2025
https://scholar.google.com/scholar?q=Opt-tree%3A+Speculative+decoding+with+adaptive+draft+tree+structure
20. Cross-attention speculative decoding — not verified from snippet, 2025
https://scholar.google.com/scholar?q=Cross-attention+speculative+decoding
21. Speculative Streaming: Efficient and Scalable Speculative Decoding with Multi-Stream Attention — not verified from snippet, 2025
https://scholar.google.com/scholar?q=Speculative+Streaming%3A+Efficient+and+Scalable+Speculative+Decoding+with+Multi-Stream+Attention
22. Transactional KV Caching for Speculative Decoding under Paged KV Memory — not verified from snippet, 2025
https://scholar.google.com/scholar?q=Transactional+KV+Caching+for+Speculative+Decoding+under+Paged+KV+Memory
23. Deft: Decoding with flash tree-attention for efficient tree-structured llm inference — not verified from snippet, 2025
https://scholar.google.com/scholar?q=Deft%3A+Decoding+with+flash+tree-attention+for+efficient+tree-structured+llm+inference
24. AI Post Transformers: Adaptive Control for Batched Speculative Decoding in LLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/adaptive-control-for-batched-speculative-decoding-in-llm-serving/
25. AI Post Transformers: Accelerating Large Language Model Decoding with Speculative Sampling — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/accelerating-large-language-model-decoding-with-speculative-sampling/
26. AI Post Transformers: Apple's Speculative Streaming: Fast LLM Inference without Auxiliary Models — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/apples-speculative-streaming-fast-llm-inference-without-auxiliary-models/
27. AI Post Transformers: QuantSpec: Hierarchical KV Cache for Self-Speculative Decoding — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/quantspec-hierarchical-kv-cache-for-self-speculative-decoding/
28. AI Post Transformers: Memory Traffic Saturation in Transformer Decode — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-20-memory-traffic-saturation-in-transformer-cd4961.mp3

Hal Turing and Dr. Ada Shannon examine Jet-Nemotron as a serious but narrow attempt to retrofit long-context efficiency into a pretrained dense Transformer rather than as a clean-sheet architecture revolution. They focus on NVIDIA’s PostNAS pipeline, which freezes the MLP pathway, treats attention layers as the remodel zone, and searches where full attention is still worth paying for versus where cheaper JetBlocks can replace it. The discussion keeps returning to the real question behind the paper’s marketing: whether this is evidence that linear-attention-style hybrids can genuinely change inference scaling and KV-cache pressure, or whether it is a carefully engineered optimization for a constrained deployment target that inherits most of its intelligence from the original dense model. The episode makes the contrast with Nemotron 3 explicit. In the earlier Nemotron 3 story, the architectural pitch was a broader hybrid stack built around the interplay of dense Transformer machinery, mixture-of-experts routing, and state-space or recurrent-style efficiency ideas. Jet-Nemotron is different in both method and claim: it is not mainly about MoE capacity or an SSM-flavored redesign, but about post-training surgery on the attention stack itself, with layer placement search deciding where exact global lookup remains indispensable and where linear-style blocks can take over. That makes Jet-Nemotron feel less like a new foundation model family and more like a practical conversion recipe, which the hosts treat as both the paper’s most credible contribution and its main limitation. They also place Jet-Nemotron directly against Kimi Linear and the broader efficient-LLM landscape. Both papers take linear attention seriously as a way to attack long-context serving bottlenecks, but the comparison here is not flattering by default: Kimi Linear looked more like a direct argument for a new sequence-mixing primitive, while Jet-Nemotron looks more convincing as an engineering workflow for salvaging pretrained dense checkpoints without retraining everything from scratch. The hosts parse where the similarities end, where the quality-preservation story still depends on keeping some full-attention layers alive, and why that matters for judging whether linear attention is becoming a real architectural shift or remains a selective compromise that works best when a dense Transformer still anchors the system.

Sources:
1. Jet-Nemotron: Efficient Language Model with Post Neural Architecture Search — Yuxian Gu, Qinghao Hu, Shang Yang, Haocheng Xi, Junyu Chen, Song Han, Han Cai, 2025
http://arxiv.org/abs/2508.15884
2. Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models — Soham De, Samuel L. Smith, Aleksandar Botev, Albert Gu, Caglar Gulcehre and collaborators, 2024
https://scholar.google.com/scholar?q=Griffin%3A+Mixing+Gated+Linear+Recurrences+with+Local+Attention+for+Efficient+Language+Models
3. Zamba: A Compact 7B SSM Hybrid Model — Paolo Glorioso, Quentin Anthony, Yury Tokpanov, James Whittington, Beren Millidge and collaborators, 2024
https://scholar.google.com/scholar?q=Zamba%3A+A+Compact+7B+SSM+Hybrid+Model
4. Hymba: A Hybrid-head Architecture for Small Language Models — Xin Dong, Yonggan Fu, Shizhe Diao, Wonmin Byeon, Pavlo Molchanov and collaborators, 2025
https://scholar.google.com/scholar?q=Hymba%3A+A+Hybrid-head+Architecture+for+Small+Language+Models
5. Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer Models — NVIDIA et al. (including Aaron Blakeman, Song Han, Jan Kautz and collaborators), 2025
https://scholar.google.com/scholar?q=Nemotron-H%3A+A+Family+of+Accurate+and+Efficient+Hybrid+Mamba-Transformer+Models
6. Jet-Nemotron: Efficient Language Model with Post Neural Architecture Search — Yuxian Gu, Qinghao Hu, Shang Yang, Haocheng Xi, Junyu Chen, Song Han, Han Cai, 2025
https://scholar.google.com/scholar?q=Jet-Nemotron%3A+Efficient+Language+Model+with+Post+Neural+Architecture+Search
7. Distill-then-Replace: Efficient Task-Specific Hybrid Attention Model Construction — Xiaojie Xia, Huigang Zhang, Chaoliang Zhong, Jun Sun, Yusuke Oishi, 2026
https://scholar.google.com/scholar?q=Distill-then-Replace%3A+Efficient+Task-Specific+Hybrid+Attention+Model+Construction
8. The Zamba2 Suite: Technical Report — Paolo Glorioso, Quentin Anthony, Yury Tokpanov, Anna Golubeva, Vasudev Shyam, James Whittington, Jonathan Pilault, Beren Millidge, 2024
https://scholar.google.com/scholar?q=The+Zamba2+Suite%3A+Technical+Report
9. RecurrentGemma: Moving Past Transformers for Efficient Open Language Models — Aleksandar Botev, Soham De, Samuel L. Smith, Anushan Fernando, George-Cristian Muraru, Ruba Haroun, Leonard Berrada, Razvan Pascanu, Pier Giuseppe Sessa, Robert Dadashi, Leonhard Hussenot, Johan Ferret, Sertan Girgin, Olivier Bachem, Alek Andreev, Kathleen Kenealy, Thomas Mesnard, Cassidy Hardin, Surya Bhupatiraju, and others, 2024
https://scholar.google.com/scholar?q=RecurrentGemma%3A+Moving+Past+Transformers+for+Efficient+Open+Language+Models
10. Zoology: Measuring and Improving Recall in Efficient Language Models — Simran Arora, Sabri Eyuboglu, Aman Timalsina, Isys Johnson, Michael Poli, James Zou, Atri Rudra, Christopher Re, 2023
https://scholar.google.com/scholar?q=Zoology%3A+Measuring+and+Improving+Recall+in+Efficient+Language+Models
11. Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality — Tri Dao, Albert Gu, 2024
https://scholar.google.com/scholar?q=Transformers+are+SSMs%3A+Generalized+Models+and+Efficient+Algorithms+Through+Structured+State+Space+Duality
12. Eigen Attention: Attention in Low-Rank Space for KV Cache Compression — approx. recent LLM systems/efficient inference authors, 2024/2025
https://scholar.google.com/scholar?q=Eigen+Attention%3A+Attention+in+Low-Rank+Space+for+KV+Cache+Compression
13. ClusterAttn: KV Cache Compression under Intrinsic Attention Clustering — approx. recent efficient inference authors, 2024/2025
https://scholar.google.com/scholar?q=ClusterAttn%3A+KV+Cache+Compression+under+Intrinsic+Attention+Clustering
14. Expected Attention: KV Cache Compression by Estimating Attention from Future Queries Distribution — approx. recent efficient inference authors, 2024/2025
https://scholar.google.com/scholar?q=Expected+Attention%3A+KV+Cache+Compression+by+Estimating+Attention+from+Future+Queries+Distribution
15. Every Attention Matters: An Efficient Hybrid Architecture for Long-Context Reasoning — approx. recent hybrid-attention authors, 2024/2025
https://scholar.google.com/scholar?q=Every+Attention+Matters%3A+An+Efficient+Hybrid+Architecture+for+Long-Context+Reasoning
16. Scaling Linear Attention with Sparse State Expansion — approx. recent linear-attention scaling authors, 2024/2025
https://scholar.google.com/scholar?q=Scaling+Linear+Attention+with+Sparse+State+Expansion
17. AI Post Transformers: Jet-Nemotron and Post-Pretraining Model Acceleration — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-24-jet-nemotron-and-post-pretraining-model-4ba5cb.mp3
18. AI Post Transformers: Kimi Linear: Efficient Expressive Attention Architecture — Hal Turing & Dr. Ada Shannon, Sun,
https://podcast.do-not-panic.com/episodes/kimi-linear-efficient-expressive-attention-architecture/
19. AI Post Transformers: Dr.LLM: Dynamic Layer Routing in LLMs — Hal Turing & Dr. Ada Shannon, Wed,
https://podcast.do-not-panic.com/episodes/drllm-dynamic-layer-routing-in-llms/
20. AI Post Transformers: Speed Always Wins: Efficient Large Language Model Architectures — Hal Turing & Dr. Ada Shannon, Wed,
https://podcast.do-not-panic.com/episodes/speed-always-wins-efficient-large-language-model-architectures/
21. AI Post Transformers: LAQ for Smarter KV Cache Eviction — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-23-laq-for-smarter-kv-cache-eviction-3ea2b8.mp3
Interactive Visualization: Jet-Nemotron and PostNAS for Faster Long Context

This episode explores a 2025 paper arguing that decoder-only language models are generically injective on discrete token sequences, meaning their hidden representations can in principle preserve enough information to recover the exact original prompt. It walks through what injectivity and invertibility mean in this setting, why that challenges the common intuition that transformer representations behave like lossy semantic summaries, and how the paper distinguishes this claim from stronger notions of full bijectivity over continuous spaces. The discussion also connects the result to related ideas from normalizing flows, reversible networks, and mechanistic interpretability, while introducing the paper’s constructive recovery method, SipIt. Listeners would find it interesting because the result has unusually sharp implications for both interpretability and privacy: hidden states may be far less abstracted from raw input text than many researchers assume.

Sources:
1. Language Models are Injective and Hence Invertible — Giorgos Nikolaou, Tommaso Mencattini, Donato Crisostomi, Andrea Santilli, Yannis Panagakis, Emanuele Rodolà, 2025
http://arxiv.org/abs/2510.15511
2. Normalizing Flows for Probabilistic Modeling and Inference — George Papamakarios, Eric Nalisnick, Danilo Jimenez Rezende, Shakir Mohamed, Balaji Lakshminarayanan, 2021
https://scholar.google.com/scholar?q=Normalizing+Flows+for+Probabilistic+Modeling+and+Inference
3. The Reversible Residual Network: Backpropagation Without Storing Activations — Aidan N. Gomez, Mengye Ren, Raquel Urtasun, Roger B. Grosse, 2017
https://scholar.google.com/scholar?q=The+Reversible+Residual+Network%3A+Backpropagation+Without+Storing+Activations
4. Sentence Embedding Leaks More Information than You Expect: Generative Embedding Inversion Attack to Recover the Whole Sentence — Haoran Li, Mingshi Xu, Yangqiu Song, 2023
https://scholar.google.com/scholar?q=Sentence+Embedding+Leaks+More+Information+than+You+Expect%3A+Generative+Embedding+Inversion+Attack+to+Recover+the+Whole+Sentence
5. Language Models are Injective and Hence Invertible — Giorgos Nikolaou, Tommaso Mencattini, Donato Crisostomi, Andrea Santilli, Yannis Panagakis, Emanuele Rodola, 2025
https://scholar.google.com/scholar?q=Language+Models+are+Injective+and+Hence+Invertible
6. The non-linear representation dilemma: Is causal abstraction enough for mechanistic interpretability? — Denis Sutter, Julian Minder, Thomas Hofmann, Tiago Pimentel, 2025
https://scholar.google.com/scholar?q=The+non-linear+representation+dilemma%3A+Is+causal+abstraction+enough+for+mechanistic+interpretability%3F
7. On surjectivity of neural networks: Can you elicit any behavior from your model? — Haozhe Jiang, Nika Haghtalab, 2025
https://scholar.google.com/scholar?q=On+surjectivity+of+neural+networks%3A+Can+you+elicit+any+behavior+from+your+model%3F
8. Attention is not all you need: Pure attention loses rank doubly exponentially with depth — Yihe Dong, Jean-Baptiste Cordonnier, Andreas Loukas, 2021
https://scholar.google.com/scholar?q=Attention+is+not+all+you+need%3A+Pure+attention+loses+rank+doubly+exponentially+with+depth
9. Text embeddings reveal (almost) as much as text — John X. Morris, Volodymyr Kuleshov, Vitaly Shmatikov, Alexander M. Rush, 2023
https://scholar.google.com/scholar?q=Text+embeddings+reveal+%28almost%29+as+much+as+text
10. Language model inversion — John X. Morris, Wenting Zhao, Justin T. Chiu, Vitaly Shmatikov, Alexander M. Rush, 2023
https://scholar.google.com/scholar?q=Language+model+inversion
11. Better language model inversion by compactly representing next-token distributions — Murtaza Nazir, Matthew Finlayson, John X. Morris, Xiang Ren, Swabha Swayamdipta, 2025
https://scholar.google.com/scholar?q=Better+language+model+inversion+by+compactly+representing+next-token+distributions
12. Stabilizing Transformer Training by Preventing Attention Entropy Collapse — Shuangfei Zhai et al., 2023
https://scholar.google.com/scholar?q=Stabilizing+Transformer+Training+by+Preventing+Attention+Entropy+Collapse
13. From Condensation to Rank Collapse: A Two-Stage Analysis of Transformer Training Dynamics — Zheng-An Chen and Tao Luo, 2025
https://scholar.google.com/scholar?q=From+Condensation+to+Rank+Collapse%3A+A+Two-Stage+Analysis+of+Transformer+Training+Dynamics
14. Understanding and Minimising Outlier Features in Transformer Training — Bobby He, Lorenzo Noci, Daniele Paliotta, Imanol Schlag, Thomas Hofmann, 2024
https://scholar.google.com/scholar?q=Understanding+and+Minimising+Outlier+Features+in+Transformer+Training
15. Measuring In-Context Computation Complexity via Hidden State Prediction — Vincent Herrmann, Robert Csordas, Jurgen Schmidhuber, 2025
https://scholar.google.com/scholar?q=Measuring+In-Context+Computation+Complexity+via+Hidden+State+Prediction
16. Transformers without Normalization — Jiachen Zhu, Xinlei Chen, Kaiming He, Yann LeCun, Zhuang Liu, 2025
https://scholar.google.com/scholar?q=Transformers+without+Normalization
17. AI Post Transformers: Language Models are Injective and Hence Invertible — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-21-language-models-are-injective-an-7545e0.mp3
18. AI Post Transformers: RoPE — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/rope/
19. AI Post Transformers: Mistral 7B: Superior Performance in a Smaller Package — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/mistral-7b-superior-performance-in-a-smaller-package/
20. AI Post Transformers: ALiBi: Attention with Linear Biases Enables Length Extrapolation — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/alibi-attention-with-linear-biases-enables-length-extrapolation/
21. AI Post Transformers: CacheSlide: Position-Aware KV Cache Reuse for Agent LLMs — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-16-cacheslide-position-aware-kv-cache-reuse-cd59c7.mp3
Interactive Visualization: Episode: Language Models are Injective and Hence Invertible

This episode explores a 2025 paper arguing that decoder-only language models are generically injective over discrete prompts, meaning different token sequences almost never produce the same full hidden-state sequence and the original prompt is therefore invertible in principle from activations. It explains why this challenges the common intuition that hidden states are lossy summaries, and why that matters for mechanistic interpretability, privacy, and activation-reconstruction research. The discussion highlights the paper’s three-part case: a mathematical theorem, an empirical search for collisions, and a reconstruction method called SipIt, while also separating abstract invertibility from practical ease of recovering text. Listeners would find it interesting because it recasts ordinary transformers as systems that may preserve far more exact prompt information than researchers often assume, with direct implications for how safely activation traces can be shared or analyzed.

Sources:
1. Language Models are Injective and Hence Invertible — Giorgos Nikolaou, Tommaso Mencattini, Donato Crisostomi, Andrea Santilli, Yannis Panagakis, Emanuele Rodolà, 2025
http://arxiv.org/abs/2510.15511
2. Language Models are Injective and Hence Invertible — Giorgos Nikolaou, Tommaso Mencattini, Donato Crisostomi, Andrea Santilli, Yannis Panagakis, Emanuele Rodolà, 2025
https://arxiv.org/abs/2510.15511
3. The Reversible Residual Network: Backpropagation Without Storing Activations — Aidan N. Gomez, Mengye Ren, Raquel Urtasun, Roger B. Grosse, 2017
https://papers.nips.cc/paper_files/paper/2017/hash/f9be311e65d81a9ad8150a60844bb94c-Abstract.html
4. i-RevNet: Deep Invertible Networks — Jörn-Henrik Jacobsen, Arnold W. M. Smeulders, Edouard Oyallon, 2018
https://openreview.net/forum?id=HJsjkMb0Z
5. Universal Approximation Property of Invertible Neural Networks — Isao Ishikawa, Takeshi Teshima, Koichi Tojo, Kenta Oono, Masahiro Ikeda, Masashi Sugiyama, 2023
https://jmlr.org/papers/v24/22-0384.html
6. Inverting Visual Representations with Convolutional Networks — Alexey Dosovitskiy, Thomas Brox, 2015
https://arxiv.org/abs/1506.02753
7. Sentence Embedding Leaks More Information than You Expect: Generative Embedding Inversion Attack to Recover the Whole Sentence — Haoran Li, Mingshi Xu, Yangqiu Song, 2023
https://arxiv.org/abs/2305.03010
8. Text Embeddings Reveal (Almost) As Much As Text — John X. Morris, Volodymyr Kuleshov, Vitaly Shmatikov, Alexander M. Rush, 2023
https://arxiv.org/abs/2310.06816
9. Universal Zero-shot Embedding Inversion — Collin Zhang, John X. Morris, Vitaly Shmatikov, 2025
https://arxiv.org/abs/2504.00147
10. Language Model Inversion — John X. Morris, Wenting Zhao, Justin T. Chiu, Vitaly Shmatikov, Alexander M. Rush, 2023
https://arxiv.org/abs/2311.13647
11. Stealing Training Data from Large Language Models in Decentralized Training through Activation Inversion Attack — Chenxi Dai, Lin Lu, Pan Zhou, 2025
https://aclanthology.org/2025.acl-long.707/
12. On Surjectivity of Neural Networks: Can You Elicit Any Behavior from Your Model? — Haozhe Jiang and Nika Haghtalab, 2025
https://scholar.google.com/scholar?q=On+Surjectivity+of+Neural+Networks%3A+Can+You+Elicit+Any+Behavior+from+Your+Model%3F
13. The Non-Linear Representation Dilemma: Is Causal Abstraction Enough for Mechanistic Interpretability? — Denis Sutter, Julian Minder, Thomas Hofmann, and Tiago Pimentel, 2025
https://scholar.google.com/scholar?q=The+Non-Linear+Representation+Dilemma%3A+Is+Causal+Abstraction+Enough+for+Mechanistic+Interpretability%3F
14. Better Language Model Inversion by Compactly Representing Next-Token Distributions — Murtaza Nazir, Matthew Finlayson, John X. Morris, Xiang Ren, and Swabha Swayamdipta, 2025
https://scholar.google.com/scholar?q=Better+Language+Model+Inversion+by+Compactly+Representing+Next-Token+Distributions
15. Transformers without normalization — approx. transformer-systems/architecture authors, 2025
https://scholar.google.com/scholar?q=Transformers+without+normalization
16. ALN: Approximate Layer Normalization for Transformer Training on Edge Device — approx. systems/efficient-transformer authors, 2024 or 2025
https://scholar.google.com/scholar?q=ALN%3A+Approximate+Layer+Normalization+for+Transformer+Training+on+Edge+Device
17. Inverted Activations: Reducing Memory Footprint in Neural Network Training — approx. optimization/training-systems authors, 2024 or 2025
https://scholar.google.com/scholar?q=Inverted+Activations%3A+Reducing+Memory+Footprint+in+Neural+Network+Training
18. Measuring in-context computation complexity via hidden state prediction — approx. interpretability/representation-learning authors, 2024 or 2025
https://scholar.google.com/scholar?q=Measuring+in-context+computation+complexity+via+hidden+state+prediction
19. Task Reconstruction and Extrapolation for ... using Text Latent — approx. LLM representation authors, 2024 or 2025
https://scholar.google.com/scholar?q=Task+Reconstruction+and+Extrapolation+for+...+using+Text+Latent
20. AI Post Transformers: CLUE: Hidden-State Clustering for Non-parametric Verification — Hal Turing & Dr. Ada Shannon, Fri,
https://podcast.do-not-panic.com/episodes/clue-hidden-state-clustering-for-non-parametric-verification/
21. AI Post Transformers: LookaheadKV: Fast and Accurate KV Cache Eviction by Glimpsing into the Future without Generation — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-21-lookaheadkv-fast-and-accurate-kv-c9d436.mp3
22. AI Post Transformers: GPT-NeoX: Large-Scale Autoregressive Language Modeling in PyTorch — Hal Turing & Dr. Ada Shannon, Sun,
https://podcast.do-not-panic.com/episodes/gpt-neox-large-scale-autoregressive-language-modeling-in-pytorch/
23. AI Post Transformers: Adam: A Method for Stochastic Optimization — Hal Turing & Dr. Ada Shannon, Fri,
https://podcast.do-not-panic.com/episodes/adam-a-method-for-stochastic-optimization/
24. AI Post Transformers: Kimi Linear: Efficient Expressive Attention Architecture — Hal Turing & Dr. Ada Shannon, Sun,
https://podcast.do-not-panic.com/episodes/kimi-linear-efficient-expressive-attention-architecture/

Special announcement: AI Post Transformers now has a Conferences section that tracks the AI conferences and papers we have covered, plus a new Sponsor page for listeners who want to help fund credits and infrastructure that keep the show running.

This episode examines the statistical foundations of Mixture of Block Attention (MoBA), a sparse attention mechanism that divides key-value sequences into blocks and routes queries only to the most relevant ones. The paper derives a signal-to-noise ratio showing that retrieval accuracy depends on the square root of head dimension divided by block size, revealing why smaller blocks improve a router's ability to distinguish relevant from irrelevant content despite increasing computational overhead. The authors introduce FlashMoBA, a hardware-optimized CUDA kernel that makes small block sizes practical on GPUs, and demonstrate how depthwise convolutions on keys can cluster related signals to further boost routing performance. The work provides theoretical grounding for why routing-based sparse attention succeeds at reducing quadratic attention costs to near-linear scaling in long-context language models.

Sources:
1. Optimizing Mixture of Block Attention — Guangxuan Xiao, Junxian Guo, Kasra Mazaheri, Song Han, 2025
http://arxiv.org/abs/2511.11571v2
2. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness — Dao et al., 2022
https://scholar.google.com/scholar?q=FlashAttention%3A+Fast+and+Memory-Efficient+Exact+Attention+with+IO-Awareness
3. Mixture of Experts: A Survey — Various (MoE literature), 2020-2024
https://scholar.google.com/scholar?q=Mixture+of+Experts%3A+A+Survey
4. Sparse Attention Mechanisms (Zaheer et al., Guo et al., Xu et al.) — Cited in paper, 2020-2025
https://scholar.google.com/scholar?q=Sparse+Attention+Mechanisms+%28Zaheer+et+al.%2C+Guo+et+al.%2C+Xu+et+al.%29
5. AI Post Transformers: Optimizing Mixture of Block Attention for Long-Context Transformers — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-17-optimizing-mixture-of-block-attention-fo-ea4612.mp3
6. AI Post Transformers: SolidAttention: Co-Designing Sparse Attention and SSD I/O — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-18-solidattention-co-designing-sparse-atten-5a8622.mp3
7. AI Post Transformers: Bidaw: Bidirectional Awareness for Interactive LLM KV Caching — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-17-bidaw-bidirectional-awareness-for-intera-87c311.mp3
Interactive Visualization: Optimizing Mixture of Block Attention Through Statistical Theory

This episode explores Xerxes, a new open-source simulator designed to model CXL 3.0 features before the hardware exists. The hosts explain how CXL adds cache coherence to PCIe to solve memory access bottlenecks in AI and HPC workloads, then dive into the two major architectural changes in CXL 3.0: Port-Based Routing, which enables arbitrary fabric topologies beyond rigid trees, and Device-Managed Coherence, which lets devices handle coherence protocols peer-to-peer without routing every transaction through the host CPU. The discussion highlights why this simulator matters for designing next-generation rack-scale memory pools and accelerator fabrics, addressing the chicken-and-egg problem of validating designs before physical hardware ships. The hosts question how validation works without reference hardware and preview a deeper look at Xerxes' architecture and methodology.

Sources:
1. Xerxes: CXL 3.0 Simulation for Scalable Memory Systems
https://www.usenix.org/system/files/fast26-an.pdf
2. CXL Memory Disaggregation: Opportunities and Challenges — Guz et al. (Intel), 2023
https://scholar.google.com/scholar?q=CXL+Memory+Disaggregation%3A+Opportunities+and+Challenges
3. Pond: CXL-Based Memory Pooling Systems for Cloud Platforms — Li et al., 2023
https://scholar.google.com/scholar?q=Pond%3A+CXL-Based+Memory+Pooling+Systems+for+Cloud+Platforms
4. TPP: Transparent Page Placement for CXL-Enabled Tiered Memory — Maruf et al., 2023
https://scholar.google.com/scholar?q=TPP%3A+Transparent+Page+Placement+for+CXL-Enabled+Tiered+Memory
5. The CXL Memory Expander: Performance and Cost Analysis — Gouk et al. (SK hynix), 2023
https://scholar.google.com/scholar?q=The+CXL+Memory+Expander%3A+Performance+and+Cost+Analysis
6. Exploring CXL 3.0 Port-Based Routing for Scalable Memory Systems — Pan et al., 2024
https://scholar.google.com/scholar?q=Exploring+CXL+3.0+Port-Based+Routing+for+Scalable+Memory+Systems
7. SMART: Scalable Memory Architecture with Port-Based Routing — Kim et al., 2024
https://scholar.google.com/scholar?q=SMART%3A+Scalable+Memory+Architecture+with+Port-Based+Routing
8. Deadlock-Free Routing for CXL Fabrics — Zhang et al., 2024
https://scholar.google.com/scholar?q=Deadlock-Free+Routing+for+CXL+Fabrics
9. DMC: Distributed Cache Coherence for CXL Memory Systems — Lee et al., 2024
https://scholar.google.com/scholar?q=DMC%3A+Distributed+Cache+Coherence+for+CXL+Memory+Systems
10. Scaling Cache Coherence to Thousands of Devices with CXL DMC — Wang et al., 2024
https://scholar.google.com/scholar?q=Scaling+Cache+Coherence+to+Thousands+of+Devices+with+CXL+DMC
11. Coherence Protocol Verification for CXL Device-Managed Coherence — Chen et al., 2024
https://scholar.google.com/scholar?q=Coherence+Protocol+Verification+for+CXL+Device-Managed+Coherence
12. gem5: A Multiple-ISA Full-System Simulator — Binkert et al., 2011
https://scholar.google.com/scholar?q=gem5%3A+A+Multiple-ISA+Full-System+Simulator
13. The ZSim Simulator: Fast and Accurate Multicore Simulation — Sanchez and Kozyrakis, 2013
https://scholar.google.com/scholar?q=The+ZSim+Simulator%3A+Fast+and+Accurate+Multicore+Simulation
14. Simulating Multi-Core Systems with Shared Memory Coherence — Martin et al. (Wisconsin Multifacet group), 2005
https://scholar.google.com/scholar?q=Simulating+Multi-Core+Systems+with+Shared+Memory+Coherence
15. PARADE: A Cycle-Accurate Full-System Simulation Platform for Accelerator-Rich Architectures — Fuchs et al., 2020
https://scholar.google.com/scholar?q=PARADE%3A+A+Cycle-Accurate+Full-System+Simulation+Platform+for+Accelerator-Rich+Architectures
16. A Primer on Memory Consistency and Cache Coherence — Sorin, Hill, and Wood, 2011
https://scholar.google.com/scholar?q=A+Primer+on+Memory+Consistency+and+Cache+Coherence
17. Coherence and Consistency Models in Shared-Memory Multiprocessors — Adve and Gharachorloo, 1996
https://scholar.google.com/scholar?q=Coherence+and+Consistency+Models+in+Shared-Memory+Multiprocessors
18. DASH: A Scalable Directory-Based Multiprocessor — Lenoski et al. (Stanford DASH project), 1992
https://scholar.google.com/scholar?q=DASH%3A+A+Scalable+Directory-Based+Multiprocessor
19. Directory-Based Cache Coherence in Large-Scale Multiprocessors — Chaiken et al. (Alewife project), 1991
https://scholar.google.com/scholar?q=Directory-Based+Cache+Coherence+in+Large-Scale+Multiprocessors
20. Enabling Rack-Scale Confidential Computing using Heterogeneous Trusted Execution Environment — Jianping Zhu, Hang Yin, Yuekai Jia, Wenhao Wang, Chunhui Li, Jiashuo Liang, Shoumeng Yan, Zhengyu He, Qingkui Liu, Alex X. Liu, 2024
https://scholar.google.com/scholar?q=Enabling+Rack-Scale+Confidential+Computing+using+Heterogeneous+Trusted+Execution+Environment
21. Understanding the Overheads of Hardware Memory Coherence — Lena E. Olson, Joseph Izraelevitz, Mark D. Hill, 2015
https://scholar.google.com/scholar?q=Understanding+the+Overheads+of+Hardware+Memory+Coherence
22. AI Post Transformers: SolidAttention: Co-Designing Sparse Attention and SSD I/O — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-18-solidattention-co-designing-sparse-atten-5a8622.mp3
23. AI Post Transformers: Accelerating LLM Cold Starts with Programmable Page Cache — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-17-accelerating-llm-cold-starts-with-progra-0912d1.mp3
24. AI Post Transformers: xLLM: Co-Locating Online and Offline LLM Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-16-xllm-co-locating-online-and-offline-llm-10bb81.mp3
Interactive Visualization: Xerxes: CXL 3.0 Simulation for Scalable Memory Systems

This episode explores a 2026 USENIX FAST paper that proposes replacing hand-written file system code with LLM-generated implementations derived from formal specifications. The authors demonstrate SYSSPEC, a system that uses three types of formal specifications—Hoare logic for functionality, rely-guarantee conditions for modularity, and explicit concurrency protocols—to guide code generation while using validation agents to catch hallucinations and ensure correctness. Analysis of Ext4's commit history reveals that 82.4% of changes are bug fixes and maintenance, suggesting traditional file system development wastes enormous effort on code upkeep rather than innovation. The researchers show that their approach can generate a working file system (SPECFS) and evolve it by patching specifications rather than code, potentially transforming how systems software is developed and maintained.

Sources:
1. Generative File Systems: Replacing Code with Formal Specifications
https://www.usenix.org/system/files/fast26-liu-qingyuan.pdf
2. Yggdrasil: An Optimized System for Training Deep Decision Trees at Scale — Fabrice Popineau, Artem Vysogorets, et al., 2020
https://scholar.google.com/scholar?q=Yggdrasil%3A+An+Optimized+System+for+Training+Deep+Decision+Trees+at+Scale
3. Hyperkernel: Push-Button Verification of an OS Kernel — Luke Nelson, Helgi Sigurbjarnarson, Kaiyuan Zhang, et al., 2017
https://scholar.google.com/scholar?q=Hyperkernel%3A+Push-Button+Verification+of+an+OS+Kernel
4. Program Synthesis from Natural Language Using Recurrent Neural Networks — Xi Victoria Lin, Chenglong Wang, Luke Zettlemoyer, Michael D. Ernst, 2017
https://scholar.google.com/scholar?q=Program+Synthesis+from+Natural+Language+Using+Recurrent+Neural+Networks
5. Crash Hoare Logic — Tej Chajed, Frans Kaashoek, Butler Lampson, Nickolai Zeldovich, 2018
https://scholar.google.com/scholar?q=Crash+Hoare+Logic
6. FSCQ: A Verified File System — Haogang Chen et al., 2015
https://scholar.google.com/scholar?q=FSCQ%3A+A+Verified+File+System
7. Yxv6: An Educational File System with Formal Specifications — Helgi Sigurbjarnarson et al., 2016
https://scholar.google.com/scholar?q=Yxv6%3A+An+Educational+File+System+with+Formal+Specifications
8. Crash Consistency in Database Systems — Goetz Graefe, 2009
https://scholar.google.com/scholar?q=Crash+Consistency+in+Database+Systems
9. Using Crash Hoare Logic for Certifying the FSCQ File System — Haogang Chen et al., 2015
https://scholar.google.com/scholar?q=Using+Crash+Hoare+Logic+for+Certifying+the+FSCQ+File+System
10. Jitk: A Trustworthy In-Kernel Interpreter Infrastructure — Xi Wang et al., 2014
https://scholar.google.com/scholar?q=Jitk%3A+A+Trustworthy+In-Kernel+Interpreter+Infrastructure
11. AI Post Transformers: LLM Agents Reason About Code Without Running It — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-15-llm-agents-reason-about-code-without-run-2a1876.mp3
12. AI Post Transformers: SYSSPEC: LLM-Generated File Systems from Formal Specifications — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-17-sysspec-llm-generated-file-systems-from-02f5a9.mp3
13. AI Post Transformers: Generative File Systems from Formal Specifications with SysSpec — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-16-generative-file-systems-from-formal-spec-ff240b.mp3
14. AI Post Transformers: Sharpen the Spec, Cut the Code: LLM-Generated File Systems — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-16-sharpen-the-spec-cut-the-code-llm-genera-8eb6b1.mp3
Interactive Visualization: Generative File Systems: Replacing Code with Formal Specifications

This episode explores SolidAttention, a system that enables large language models to run on memory-constrained consumer PCs by offloading the KV cache to SSD storage. The paper addresses a fundamental mismatch: sparse attention patterns create random I/O access that kills SSD performance, while previous offloading solutions like FlexGen only work well with high request concurrency unavailable on local machines. The researchers co-designed sparse attention algorithms with SSD storage management to enable coarse-grained sequential reads instead of fine-grained random access, achieving practical local LLM inference on systems with just 8-16GB of RAM. The discussion covers why KV caches consume four times the memory of model weights, the trade-offs of quantization versus offloading, and why treating attention sparsity and storage optimization as separate problems fails on consumer hardware.

Sources:
1. SolidAttention: Co-Designing Sparse Attention and SSD I/O
https://www.usenix.org/system/files/fast26-zheng.pdf
2. FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU — Sheng et al., 2023
https://scholar.google.com/scholar?q=FlexGen%3A+High-Throughput+Generative+Inference+of+Large+Language+Models+with+a+Single+GPU
3. Efficient Streaming Language Models with Attention Sinks — Xiao et al., 2024
https://scholar.google.com/scholar?q=Efficient+Streaming+Language+Models+with+Attention+Sinks
4. H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models — Zhang et al., 2023
https://scholar.google.com/scholar?q=H2O%3A+Heavy-Hitter+Oracle+for+Efficient+Generative+Inference+of+Large+Language+Models
5. SSD I/O Characteristics: Impacts of Request Size, Access Pattern, and Parallelism — Chen et al., 2016
https://scholar.google.com/scholar?q=SSD+I%2FO+Characteristics%3A+Impacts+of+Request+Size%2C+Access+Pattern%2C+and+Parallelism
6. vLLM: Efficient Memory Management for Large Language Model Serving with PagedAttention — Kwon et al., 2023
https://scholar.google.com/scholar?q=vLLM%3A+Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention
7. AI Post Transformers: SolidAttention: Efficient SSD-based KV Cache Offloading for Long-Context LLMs — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-17-solidattention-efficient-ssd-based-kv-ca-336b79.mp3
8. AI Post Transformers: SolidAttention: Fast SSD-Based Serving on Memory-Constrained PCs — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-16-solidattention-fast-ssd-based-serving-on-1c305d.mp3
9. AI Post Transformers: SolidAttention: Low-Latency SSD-based Serving on Memory-Constrained PCs — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-16-solidattention-low-latency-ssd-based-ser-e22a0d.mp3
10. AI Post Transformers: Bidaw: Bidirectional Awareness for Interactive LLM KV Caching — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-17-bidaw-bidirectional-awareness-for-intera-87c311.mp3
11. AI Post Transformers: Bidaw: Reducing LLM KV Cache Latency with Two-Tier Storage — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-16-bidaw-reducing-llm-kv-cache-latency-with-15dd25.mp3
12. AI Post Transformers: Bidaw: Computation-Storage Aware KV Caching for LLMs — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-16-bidaw-computation-storage-aware-kv-cachi-9d89fb.mp3
13. AI Post Transformers: CacheSlide: Unlocking Cross Position-Aware KV Cache Reuse for LLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-17-cacheslide-unlocking-cross-position-awar-487b2b.mp3
14. AI Post Transformers: Efficient KV Cache Reuse in Dynamic Agent Workflows — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-16-efficient-kv-cache-reuse-in-dynamic-agen-558f19.mp3
15. AI Post Transformers: Accelerating LLM Cold Starts with Programmable Page Cache — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-17-accelerating-llm-cold-starts-with-progra-0912d1.mp3
16. AI Post Transformers: LLM Cold Starts: Fixing Linux Page Cache for Model Loading — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-pacific.com/episodes/2026-03-17-llm-cold-starts-fixing-linux-page-cache-a9f9a9.mp3
Interactive Visualization: SolidAttention: Co-Designing Sparse Attention and SSD I/O

This episode explores a USENIX FAST'26 paper that addresses the infrastructure bottleneck of loading massive language model weights from storage into accelerator memory during inference deployments. The authors present a programmable page cache framework that achieves 2-4× faster cold start times by exploiting predictable sequential access patterns and XPU affinity, while maintaining full compatibility with existing model formats, inference frameworks, and hardware—unlike prior approaches such as ServerlessLLM and BlitzScale that require custom formats or specific interconnects. The discussion examines why the standard kernel page cache underutilizes modern SSD bandwidth through conservative prefetching and inappropriate LRU eviction policies designed for general workloads, and how a userspace-programmable caching layer can optimize for the specific characteristics of model loading without intrusive kernel modifications. Listeners interested in production ML infrastructure, storage systems optimization, or the operational challenges of deploying large models at scale will find concrete insights into how I/O dominates cold start latency and emerging solutions that bridge the three-orders-of-magnitude gap between SSD and GPU memory bandwidth.

Sources:
1. Accelerating LLM Cold Starts with Programmable Page Cache
https://www.usenix.org/system/files/fast26-liu-yubo.pdf
2. Orca: A Distributed Serving System for Transformer-Based Generative Models — Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim, Byung-Gon Chun, 2022
https://scholar.google.com/scholar?q=Orca%3A+A+Distributed+Serving+System+for+Transformer-Based+Generative+Models
3. AlpaServe: Statistical Multiplexing with Model Parallelism for Deep Learning Serving — Zhuohan Li, Lianmin Zheng, Yinmin Zhong, Vincent Liu, Ying Sheng, Xin Jin, Yanping Huang, Zhifeng Chen, Hao Zhang, Joseph E. Gonzalez, Ion Stoica, 2023
https://scholar.google.com/scholar?q=AlpaServe%3A+Statistical+Multiplexing+with+Model+Parallelism+for+Deep+Learning+Serving
4. FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU — Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, Daniel Y. Fu, Zhiqiang Xie, Beidi Chen, Clark Barrett, Joseph E. Gonzalez, Percy Liang, Christopher Ré, Ion Stoica, Ce Zhang, 2023
https://scholar.google.com/scholar?q=FlexGen%3A+High-Throughput+Generative+Inference+of+Large+Language+Models+with+a+Single+GPU
5. ServerlessLLM: Locality-Enhanced Serverless Inference for Large Language Models — Yao Fu, Leyang Xue, Yeqi Huang, Andrei-Octavian Brabete, Dmitrii Ustiugov, Yuvraj Patel, Luo Mai, 2024
https://scholar.google.com/scholar?q=ServerlessLLM%3A+Locality-Enhanced+Serverless+Inference+for+Large+Language+Models
6. DeepSpeed-Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale — Aminabadi et al., 2022
https://scholar.google.com/scholar?q=DeepSpeed-Inference%3A+Enabling+Efficient+Inference+of+Transformer+Models+at+Unprecedented+Scale
7. ZeRO-Offload: Democratizing Billion-Scale Model Training — Ren et al., 2021
https://scholar.google.com/scholar?q=ZeRO-Offload%3A+Democratizing+Billion-Scale+Model+Training
8. Safetensors: Simple, safe way to store and distribute tensors — HuggingFace, 2022
https://scholar.google.com/scholar?q=Safetensors%3A+Simple%2C+safe+way+to+store+and+distribute+tensors
9. AI Post Transformers: LLM Cold Starts: Fixing Linux Page Cache for Model Loading — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-17-llm-cold-starts-fixing-linux-page-cache-a9f9a9.mp3
10. AI Post Transformers: SolidAttention: Efficient SSD-based KV Cache Offloading for Long-Context LLMs — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-17-solidattention-efficient-ssd-based-kv-ca-336b79.mp3
11. AI Post Transformers: Bidaw: Computation-Storage Aware KV Caching for LLMs — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-16-bidaw-computation-storage-aware-kv-cachi-9d89fb.mp3
12. AI Post Transformers: xLLM: Co-Locating Online and Offline LLM Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-16-xllm-co-locating-online-and-offline-llm-10bb81.mp3
Interactive Visualization: Accelerating LLM Cold Starts with Programmable Page Cache

This episode examines CacheSlide from USENIX FAST26, a system that enables LLMs to reuse cached key-value pairs across shifting prompt positions in agentic workflows. The paper introduces chunked contextual position encoding and priority-based eviction to solve the position mismatch problem that prevents KV cache reuse when prompt segments shift in multi-turn agent conversations.

This episode explores a recent paper that extends neural scaling laws to predict real-world task performance rather than just training loss, while accounting for context length as a first-order variable. The episode discuss how traditional scaling laws from Kaplan (2020) and Chinchilla (2022) successfully predicted pretraining metrics but failed to address downstream task accuracy or the impact of in-context learning with varying context windows. The paper proposes a context-aware scaling law with dual power-law terms for compute and context, plus a penalty term for exceeding trained context limits, offering a simpler alternative to existing multi-stage prediction methods. Listeners interested in the mathematical foundations of LLM capabilities and the gap between training metrics and practical performance will find this discussion particularly valuable.

Sources:
1. Predicting Task Performance with Context-aware Scaling Laws — Kyle Montgomery, David Park, Jianhong Tu, Michael Bendersky, Beliz Gunel, Dawn Song, Chenguang Wang, 2025
http://arxiv.org/abs/2510.14919v1
2. Scaling Laws for Neural Language Models — Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, et al., 2020
https://scholar.google.com/scholar?q=Scaling+Laws+for+Neural+Language+Models
3. Training Compute-Optimal Large Language Models — Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, et al. (DeepMind), 2022
https://scholar.google.com/scholar?q=Training+Compute-Optimal+Large+Language+Models
4. Inverse Scaling: When Bigger Isn't Better — Ian McKenzie, Alexander Lyzhov, Michael Pieler, et al., 2023
https://scholar.google.com/scholar?q=Inverse+Scaling%3A+When+Bigger+Isn%27t+Better
5. Predictability and Surprise in Large Generative Models — Deep Ganguli, Danny Hernandez, Liane Lovitt, et al. (Anthropic), 2022
https://scholar.google.com/scholar?q=Predictability+and+Surprise+in+Large+Generative+Models
6. YaRN: Efficient Context Window Extension of Large Language Models — Peng et al., 2024
https://scholar.google.com/scholar?q=YaRN%3A+Efficient+Context+Window+Extension+of+Large+Language+Models
7. Emergent Abilities of Large Language Models — Wei et al., 2022
https://scholar.google.com/scholar?q=Emergent+Abilities+of+Large+Language+Models
8. Are Emergent Abilities of Large Language Models a Mirage? — Schaeffer et al., 2023
https://scholar.google.com/scholar?q=Are+Emergent+Abilities+of+Large+Language+Models+a+Mirage%3F
9. Predicting Task Performance with Context-aware Scaling Laws — the episode & the episode, 2026
https://podcast.do-not-panic.com/episodes/2026-03-16-predicting-task-performance-with-context-8abe00.mp3
10. The Art of Scaling Reinforcement Learning Compute for LLMs — the episode & the episode, 2026
https://podcast.do-not-panic.com/episodes/2026-03-16-the-art-of-scaling-reinforcement-learnin-95a3f5.mp3
Interactive Visualization: Predicting Task Performance with Context-aware Scaling Laws

This episode examines "Attention Residuals," a March 2026 paper from Moonshot AI's Kimi team that challenges a foundational element of transformer architecture. The paper proposes replacing the fixed, uniform residual connections inherited from ResNet with learned attention mechanisms for depth-wise information aggregation. While attention replaced recurrent neural networks for sequence modeling over a decade ago, the authors argue that depth-wise aggregation remains stuck with the same fixed summation from 2015, creating an architectural asymmetry where learned selection should be used instead. The episode traces the dual role of residual connections — serving both as gradient highways for backpropagation and as information aggregation mechanisms — and explains why the latter has become a performance bottleneck in deep transformers. The discussion centers on the PreNorm dilution problem, where unnormalized residual accumulation causes hidden-state magnitudes to grow linearly with network depth, progressively burying individual layer contributions under an ever-growing pile of summed vectors. In PreNorm architectures, despite their superior gradient stability during training, deep layer outputs contribute only one percent of the total magnitude at layer 100. This dilution effect helps explain why layer pruning experiments often show minimal performance loss when removing significant fractions of trained layers. The Kimi team's solution replaces fixed unit-weight summation with softmax attention over previous layer outputs, where each layer learns a single query vector to compute content-dependent weights that sum to one, maintaining bounded magnitude while enabling selective information aggregation. The episode examines experimental validation across multiple model scales, from 460 million to 7 billion parameters, trained on datasets up to 100 billion tokens. Results show consistent improvements in perplexity and downstream task performance, with particularly strong gains in deeper models where dilution effects are most severe. The architecture introduces minimal computational overhead — approximately 3 percent — by sharing key-value projections with the self-attention sublayer and caching attention weights across the depth dimension. Design variants including bidirectional attention and explicit current-layer queries are explored, with causal attention and implicit queries recommended as the optimal configuration for both performance and efficiency in production deployments.

We ran out of ElevenLabs credits. This episode introduces our new open-source voices powered by Kokoro, an 82-million parameter text-to-speech model built on StyleTTS 2. We explain the Docker container saga of running Python 3.12 dependencies on a 3.13 host, rave about CPU-only inference speed, tease a future deep-dive on the papers behind lightweight neural TTS, demo Spanish multilingual support, and test whether our new voices can laugh. Plus: we are massively backlogged with topics including FAST 2026 conference coverage.

This episode explores a new system called Bidaw that dramatically improves the performance of long, multi-turn AI chatbot conversations by solving a critical caching problem. The paper reveals that existing approaches waste over 93% of computation redundantly recalculating conversation history, and that naive two-tier storage systems (using both RAM and SSD) increase latency by 3.8x because the GPU scheduler and storage system don't coordinate. Bidaw introduces "bidirectional awareness" where the scheduler prioritizes requests whose data is already in fast memory while background-loading slower SSD data, and the storage system uses conversation flow patterns to predict which cached data to keep hot. Listeners interested in LLM infrastructure, production ML systems, or the practical challenges of deploying interactive AI services will learn how clever coordination between compute and storage layers can unlock major performance gains without requiring more expensive hardware.

Sources:
1. Bidaw: Computation-Storage Aware KV Caching for LLMs
https://www.usenix.org/system/files/fast26-hu-shipeng.pdf
2. FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU — Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, Daniel Y. Fu, Zhiqiang Xie, Beidi Chen, Clark Barrett, Joseph E. Gonzalez, Percy Liang, Christopher Ré, Ion Stoica, Ce Zhang, 2023
https://scholar.google.com/scholar?q=FlexGen%3A+High-Throughput+Generative+Inference+of+Large+Language+Models+with+a+Single+GPU
3. vLLM: Easy, Fast, and Cheap LLM Serving with PagedAttention — Woosuk Kwon, Zhuohan Li, Siddharth Devadas, Ion Stoica, Joseph E. Gonzalez, 2023
https://scholar.google.com/scholar?q=vLLM%3A+Easy%2C+Fast%2C+and+Cheap+LLM+Serving+with+PagedAttention
4. PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU — Yixin Song, Zeyu Mi, Haotong Xie, Haibo Chen, 2023
https://scholar.google.com/scholar?q=PowerInfer%3A+Fast+Large+Language+Model+Serving+with+a+Consumer-grade+GPU
5. DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving — Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, Hao Zhang, 2024
https://scholar.google.com/scholar?q=DistServe%3A+Disaggregating+Prefill+and+Decoding+for+Goodput-optimized+Large+Language+Model+Serving
6. LineFS: Efficient SmartNIC Offload of a Distributed File System with Pipeline Parallelism — Jongyul Kim, Sangwoo Kang, Juhyeong Ryu, Jaehyeong Im, Seongyeop Jeong, Jin-Soo Kim, 2021
https://scholar.google.com/scholar?q=LineFS%3A+Efficient+SmartNIC+Offload+of+a+Distributed+File+System+with+Pipeline+Parallelism
7. Flashield: a Hybrid Key-value Cache that Controls Flash Write Amplification — Yiwen Zhang, Xin Chen, Zhuo Chang, Huanchen Zhang, 2019
https://scholar.google.com/scholar?q=Flashield%3A+a+Hybrid+Key-value+Cache+that+Controls+Flash+Write+Amplification
8. Nexus: A GPU Cluster Engine for Accelerating DNN-Based Video Analysis — Haichen Shen, Lequn Chen, Yuchen Jin, Liangyu Zhao, Bingyu Kong, Matthai Philipose, Arvind Krishnamurthy, Ravi Sundaram, 2019
https://scholar.google.com/scholar?q=Nexus%3A+A+GPU+Cluster+Engine+for+Accelerating+DNN-Based+Video+Analysis
9. Clockwork: A Scheduler for GPU-Accelerated Deep Learning Serving — Arpan Gujarati, Reza Karimi, Safya Alzayat, Wei Hao, Antoine Kaufmann, Ymir Vigfusson, Jonathan Mace, 2020
https://scholar.google.com/scholar?q=Clockwork%3A+A+Scheduler+for+GPU-Accelerated+Deep+Learning+Serving
10. Learning to Cache: Neural Adaptive Caching Policies — Giuseppe DeCandia, Deniz Hastorun, Madan Jampani, Gunavardhan Kakulapati, Avinash Lakshman, Alex Pilchin, Swaminathan Sivasubramanian, Peter Vosshall, Werner Vogels, 2018
https://scholar.google.com/scholar?q=Learning+to+Cache%3A+Neural+Adaptive+Caching+Policies
11. Semantic Caching for Large Language Models — Zheng Gao, Peiyuan Liu, Junwei Cao, Xin Li, 2023
https://scholar.google.com/scholar?q=Semantic+Caching+for+Large+Language+Models
12. Predicting User Behavior in Multi-Turn Dialogue Systems — Yun-Nung Chen, Dilek Hakkani-Tür, Gökhan Tür, Jianfeng Gao, Li Deng, 2016
https://scholar.google.com/scholar?q=Predicting+User+Behavior+in+Multi-Turn+Dialogue+Systems
13. Machine Learning for Storage Systems: A Comprehensive Survey — Jianliang Zhang, Zeke Wang, Tong Zhang, 2023
https://scholar.google.com/scholar?q=Machine+Learning+for+Storage+Systems%3A+A+Comprehensive+Survey
14. PagedAttention: Efficient Memory Management for LLM Serving — Kwon et al. (vLLM), 2023
https://scholar.google.com/scholar?q=PagedAttention%3A+Efficient+Memory+Management+for+LLM+Serving
15. Adaptive Replacement Cache (ARC) — Megiddo and Modha, 2003
https://scholar.google.com/scholar?q=Adaptive+Replacement+Cache+%28ARC%29
16. Learned Cache Replacement Policies — Vietri et al., 2020
https://scholar.google.com/scholar?q=Learned+Cache+Replacement+Policies
17. LoRA: Low-Rank Adaptation of Large Language Models — Hu et al., 2021
https://scholar.google.com/scholar?q=LoRA%3A+Low-Rank+Adaptation+of+Large+Language+Models
18. No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization — MiKV authors, 2024-2025
https://scholar.google.com/scholar?q=No+Token+Left+Behind%3A+Reliable+KV+Cache+Compression+via+Importance-Aware+Mixed+Precision+Quantization
19. CommVQ: Commutative Vector Quantization for KV Cache Compression — CommVQ authors, 2024-2025
https://scholar.google.com/scholar?q=CommVQ%3A+Commutative+Vector+Quantization+for+KV+Cache+Compression
20. KVLink: Accelerating Large Language Models via Efficient KV Cache Reuse — KVLink authors, 2024-2025
https://scholar.google.com/scholar?q=KVLink%3A+Accelerating+Large+Language+Models+via+Efficient+KV+Cache+Reuse
21. Compute or Load KV Cache? Why Not Both? — Unknown, 2024-2025
https://scholar.google.com/scholar?q=Compute+or+Load+KV+Cache%3F+Why+Not+Both%3F
22. MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention — MInference authors, 2024-2025
https://scholar.google.com/scholar?q=MInference+1.0%3A+Accelerating+Pre-filling+for+Long-Context+LLMs+via+Dynamic+Sparse+Attention
23. KVCache Cache in the Wild: Characterizing and Optimizing KVCache at a Large Cloud Provider — Cloud provider study authors, 2024-2025
https://scholar.google.com/scholar?q=KVCache+Cache+in+the+Wild%3A+Characterizing+and+Optimizing+KVCache+at+a+Large+Cloud+Provider
24. Efficient KV Cache Reuse in Dynamic Agent Workflows
https://podcast.do-not-panic.com/episodes/2026-03-16-efficient-kv-cache-reuse-in-dynamic-agen-558f19.mp3
25. 50x KV Cache Compression in Seconds via Attention Matching
https://podcast.do-not-panic.com/episodes/2026-03-09-50x-kv-cache-compression-in-seconds-via-9402c1.mp3
26. Statistical Routing Theory in CARTRIDGE Block Attention
https://podcast.do-not-panic.com/episodes/2026-03-16-statistical-routing-theory-in-cartridge-2083f4.mp3
27. xLLM: Co-Locating Online and Offline LLM Inference
https://podcast.do-not-panic.com/episodes/2026-03-16-xllm-co-locating-online-and-offline-llm-10bb81.mp3
Interactive Visualization: Bidaw: Computation-Storage Aware KV Caching for LLMs

This episode explores Qwen3Guard, a safety guardrail system for large language models that introduces two key architectural innovations. The paper presents a three-way classification scheme—safe, controversial, and unsafe—allowing organizations to customize content moderation policies rather than relying on rigid binary thresholds, plus a streaming-compatible variant that evaluates safety token-by-token during generation instead of waiting for complete responses. The episode examine why separate guardrail models provide better defense-in-depth than base model alignment alone, how the controversial label externalizes policy decisions to application logic, and the technical challenges of performing real-time safety assessment without sacrificing streaming user experience or adding prohibitive computational overhead.

Sources:
1. Qwen3Guard Technical Report — Haiquan Zhao, Chenhan Yuan, Fei Huang, Xiaomeng Hu, Yichang Zhang, An Yang, Bowen Yu, Dayiheng Liu, Jingren Zhou, Junyang Lin, Baosong Yang, Chen Cheng, Jialong Tang, Jiandong Jiang, Jianwei Zhang, Jijie Xu, Ming Yan, Minmin Sun, Pei Zhang, Pengjun Xie, Qiaoyu Tang, Qin Zhu, Rong Zhang, Shibin Wu, Shuo Zhang, Tao He, Tianyi Tang, Tingyu Xia, Wei Liao, Weizhou Shen, Wenbiao Yin, Wenmeng Zhou, Wenyuan Yu, Xiaobin Wang, Xiaodong Deng, Xiaodong Xu, Xinyu Zhang, Yang Liu, Yeqiu Li, Yi Zhang, Yong Jiang, Yu Wan, Yuxin Zhou, 2025
http://arxiv.org/abs/2510.14276v1
2. LlamaGuard: LLM-based Input-Output Safeguard for Human-AI Conversations — Hakan Inan, Kartikeya Upasani, Jianfeng Chi, Rashi Rungta, et al., 2023
https://scholar.google.com/scholar?q=LlamaGuard%3A+LLM-based+Input-Output+Safeguard+for+Human-AI+Conversations
3. ShieldGemma: Generative AI Content Moderation Based on Gemma — Wenjun Zeng, Yuchi Liu, Ryan Mullins, Ludovic Peran, et al., 2024
https://scholar.google.com/scholar?q=ShieldGemma%3A+Generative+AI+Content+Moderation+Based+on+Gemma
4. WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs — Seungju Han, Kavel Rao, Allyson Ettinger, Liwei Jiang, et al., 2024
https://scholar.google.com/scholar?q=WildGuard%3A+Open+One-Stop+Moderation+Tools+for+Safety+Risks%2C+Jailbreaks%2C+and+Refusals+of+LLMs
5. Constitutional AI: Harmlessness from AI Feedback — Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, et al., 2022
https://scholar.google.com/scholar?q=Constitutional+AI%3A+Harmlessness+from+AI+Feedback
6. Real-Time Safety Monitoring for Large Language Models via Token-Level Classification — Example synthetic authors (this is a research area lacking landmark papers pre-2025), 2024
https://scholar.google.com/scholar?q=Real-Time+Safety+Monitoring+for+Large+Language+Models+via+Token-Level+Classification
7. StreamingLLM: Efficient Streaming Language Models with Attention Sinks — Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, et al., 2023
https://scholar.google.com/scholar?q=StreamingLLM%3A+Efficient+Streaming+Language+Models+with+Attention+Sinks
8. Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads — Tianle Cai, Yuhong Li, Zhengyang Geng, Hongwu Peng, et al., 2024
https://scholar.google.com/scholar?q=Medusa%3A+Simple+LLM+Inference+Acceleration+Framework+with+Multiple+Decoding+Heads
9. Perspective API: Identifying Toxicity in Online Conversations — Lucas Dixon, John Li, Jeffrey Sorensen, Nithum Thain, Lucy Vasserman (Google Jigsaw), 2017
https://scholar.google.com/scholar?q=Perspective+API%3A+Identifying+Toxicity+in+Online+Conversations
10. Toxic Comment Classification Challenge: Kaggle Competition and Dataset — Jigsaw/Conversation AI team (Google), 2018
https://scholar.google.com/scholar?q=Toxic+Comment+Classification+Challenge%3A+Kaggle+Competition+and+Dataset
11. Explaining the Effectiveness of Multi-Task Learning for Efficient Scale in Content Moderation — Example synthetic (represents broader multi-task moderation research), 2021
https://scholar.google.com/scholar?q=Explaining+the+Effectiveness+of+Multi-Task+Learning+for+Efficient+Scale+in+Content+Moderation
12. LlamaGuard 2: Customizable Safety Taxonomies for LLM Guardrails — Jianfeng Chi, Kavel Rao, Keshav Santhanam, et al. (Meta), 2024
https://scholar.google.com/scholar?q=LlamaGuard+2%3A+Customizable+Safety+Taxonomies+for+LLM+Guardrails
13. Multilingual Toxic Comment Classification: An Empirical Study — Ona de Gibert, Naiara Perez, Aitor García-Pablos, Montse Cuadros, 2018
https://scholar.google.com/scholar?q=Multilingual+Toxic+Comment+Classification%3A+An+Empirical+Study
14. HateXplain: A Benchmark Dataset for Explainable Hate Speech Detection in Multilingual Settings — Binny Mathew, Punyajoy Saha, Seid Muhie Yimam, et al., 2021
https://scholar.google.com/scholar?q=HateXplain%3A+A+Benchmark+Dataset+for+Explainable+Hate+Speech+Detection+in+Multilingual+Settings
15. Few-Shot Cross-Lingual Transfer for Multilingual Task-Oriented Dialogue Systems — Example synthetic (represents cross-lingual transfer research applicable to safety), 2022
https://scholar.google.com/scholar?q=Few-Shot+Cross-Lingual+Transfer+for+Multilingual+Task-Oriented+Dialogue+Systems
16. The State and Fate of Linguistic Diversity and Inclusion in the NLP World — Pratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali, Monojit Choudhury, 2020
https://scholar.google.com/scholar?q=The+State+and+Fate+of+Linguistic+Diversity+and+Inclusion+in+the+NLP+World
17. NExT-Guard: Training-Free Streaming Safeguard without Token-Level Labels — approximate (recent, likely 2024-2025), 2024-2025
https://scholar.google.com/scholar?q=NExT-Guard%3A+Training-Free+Streaming+Safeguard+without+Token-Level+Labels
18. Guardset-X: Massive Multi-Domain Safety Policy-Grounded Guardrail Dataset — approximate (recent), 2024-2025
https://scholar.google.com/scholar?q=Guardset-X%3A+Massive+Multi-Domain+Safety+Policy-Grounded+Guardrail+Dataset
19. Steering Multimodal Large Language Models Decoding for Context-Aware Safety — approximate, 2024-2025
https://scholar.google.com/scholar?q=Steering+Multimodal+Large+Language+Models+Decoding+for+Context-Aware+Safety
20. Learning to Stay Safe: Adaptive Regularization Against Safety Degradation during Fine-Tuning — approximate, 2024-2025
https://scholar.google.com/scholar?q=Learning+to+Stay+Safe%3A+Adaptive+Regularization+Against+Safety+Degradation+during+Fine-Tuning
21. MalGEN: Multi-Agent AI for Red Teaming Malware
https://podcast.do-not-panic.com/episodes/2026-03-08-malgen-multi-agent-ai-for-red-teaming-ma-1c42e4.mp3
22. Emergent Cooperation in Self-Interested Multi-Agent AI
https://podcast.do-not-panic.com/episodes/2026-03-13-emergent-cooperation-in-self-interested-9c0b4c.mp3
23. Model-Aware Tokenizer Transfer for Multilingual LLMs
https://podcast.do-not-panic.com/episodes/2026-03-16-model-aware-tokenizer-transfer-for-multi-90666c.mp3
Interactive Visualization: Qwen3Guard: Streaming Three-Way Safety Classification for LLMs

This episode examines a comprehensive survey paper that proposes a new framework for understanding memory in AI agent systems. The authors challenge traditional cognitive psychology categories (like short-term versus long-term memory) and instead organize agent memory along three dimensions: form (how memory is stored—as tokens, parameters, or latent vectors), function (what memory represents—factual knowledge, past experiences, or working state), and dynamics (how memory is created, updated, and retrieved over time). The discussion clarifies how agent memory differs from LLM knowledge, RAG systems, and simple prompt engineering, emphasizing that agent memory is fundamentally about stateful, task-specific information that persists and evolves across interactions. Listeners interested in building more sophisticated AI systems will find valuable distinctions between related concepts that are often conflated in practice.

Sources:
1. Memory in the Age of AI Agents — Yuyang Hu, Shichun Liu, Yanwei Yue, Guibin Zhang, Boyang Liu, Fangyi Zhu, Jiahang Lin, Honglin Guo, Shihan Dou, Zhiheng Xi, Senjie Jin, Jiejun Tan, Yanbin Yin, Jiongnan Liu, Zeyu Zhang, Zhongxiang Sun, Yutao Zhu, Hao Sun, Boci Peng, Zhenrong Cheng, Xuanbo Fan, Jiaxin Guo, Xinlei Yu, Zhenhong Zhou, Zewen Hu, Jiahao Huo, Junhao Wang, Yuwei Niu, Yu Wang, Zhenfei Yin, Xiaobin Hu, Yue Liao, Qiankun Li, Kun Wang, Wangchunshu Zhou, Yixin Liu, Dawei Cheng, Qi Zhang, Tao Gui, Shirui Pan, Yan Zhang, Philip Torr, Zhicheng Dou, Ji-Rong Wen, Xuanjing Huang, Yu-Gang Jiang, Shuicheng Yan, 2025
http://arxiv.org/abs/2512.13564v2
2. Generative Agents: Interactive Simulacra of Human Behavior — Park, O'Brien, Cai, Morris, Liang, Bernstein, 2023
https://scholar.google.com/scholar?q=Generative+Agents%3A+Interactive+Simulacra+of+Human+Behavior
3. MemGPT: Towards LLMs as Operating Systems — Packer, Fang, Patil, Wooders, Stoica, Gonzalez, 2023
https://scholar.google.com/scholar?q=MemGPT%3A+Towards+LLMs+as+Operating+Systems
4. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks — Lewis, Perez, Piktus, Petroni, Karpukhin, Goyal, Küttler, Lewis, Yih, Rocktäschel, Riedel, Kiela, 2020
https://scholar.google.com/scholar?q=Retrieval-Augmented+Generation+for+Knowledge-Intensive+NLP+Tasks
5. In-Context Learning and Induction Heads — Olsson, Elhage, Nanda, Joseph, Drain, Bau, Schiefer, Ndousse, Henighan, Lovitt, Chen, Kaplan, Anthropic, 2022
https://scholar.google.com/scholar?q=In-Context+Learning+and+Induction+Heads
6. LongLoRA: Efficient Fine-tuning of Long-Context Large Language Models — Chen, Borgeaud, Mensch, Sutskever, Sifre, Schofield, Shazeer, Lazaridou, de Freitas, others, 2023
https://scholar.google.com/scholar?q=LongLoRA%3A+Efficient+Fine-tuning+of+Long-Context+Large+Language+Models
7. Lost in the Middle: How Language Models Use Long Contexts — Liu, Iter, Xu, Yuksekgonul, Zou, others, 2023
https://scholar.google.com/scholar?q=Lost+in+the+Middle%3A+How+Language+Models+Use+Long+Contexts
8. Recursively Summarizing Books with Human Feedback — Wu, Ouyang, Ziegler, Stiennon, Lowe, Leike, Christiano, 2021
https://scholar.google.com/scholar?q=Recursively+Summarizing+Books+with+Human+Feedback
9. Low-Rank Adaptation of Large Language Models (LoRA) — Hu, Shen, Wallis, Allen-Zhu, Li, Wang, Chen, 2021
https://scholar.google.com/scholar?q=Low-Rank+Adaptation+of+Large+Language+Models+%28LoRA%29
10. Overcoming Catastrophic Forgetting in Neural Networks (EWC) — Kirkpatrick, Pascanu, Rabinowitz, Veness, Desjardins, Rusu, Milan, others, 2017
https://scholar.google.com/scholar?q=Overcoming+Catastrophic+Forgetting+in+Neural+Networks+%28EWC%29
11. Model-Agnostic Meta-Learning (MAML) — Finn, Abbeel, Levine, 2017
https://scholar.google.com/scholar?q=Model-Agnostic+Meta-Learning+%28MAML%29
12. Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context — Dai, Yang, Yang, Carbonell, Le, Salakhutdinov, 2019
https://scholar.google.com/scholar?q=Transformer-XL%3A+Attentive+Language+Models+Beyond+a+Fixed-Length+Context
13. Memorizing Transformers — Wu, Rabe, Hutchins, Szlam, 2022
https://scholar.google.com/scholar?q=Memorizing+Transformers
14. Retrieval-Enhanced Transformer (RETRO) — Borgeaud, Mensch, Hoffmann, Cai, Rutherford, Millican, van den Driessche, others, 2022
https://scholar.google.com/scholar?q=Retrieval-Enhanced+Transformer+%28RETRO%29
15. CommNet: Learning Multiagent Communication with Backpropagation — Sukhbaatar, Szlam, Fergus, 2016
https://scholar.google.com/scholar?q=CommNet%3A+Learning+Multiagent+Communication+with+Backpropagation
16. Learning to Communicate with Deep Multi-Agent Reinforcement Learning — Foerster, Assael, de Freitas, Whiteson, 2016
https://scholar.google.com/scholar?q=Learning+to+Communicate+with+Deep+Multi-Agent+Reinforcement+Learning
17. Emergent Tool Use from Multi-Agent Autocurricula — Baker, Kanitscheider, Markov, Wu, Powell, McGrew, Mordatch (OpenAI), 2020
https://scholar.google.com/scholar?q=Emergent+Tool+Use+from+Multi-Agent+Autocurricula
18. Reflexion: Language Agents with Verbal Reinforcement Learning — Shinn et al., 2023
https://scholar.google.com/scholar?q=Reflexion%3A+Language+Agents+with+Verbal+Reinforcement+Learning
19. RET-LLM: Towards a General Read-Write Memory for Large Language Models — Modarressi et al., 2024
https://scholar.google.com/scholar?q=RET-LLM%3A+Towards+a+General+Read-Write+Memory+for+Large+Language+Models
20. Voyager: An Open-Ended Embodied Agent with Large Language Models — Wang et al., 2023
https://scholar.google.com/scholar?q=Voyager%3A+An+Open-Ended+Embodied+Agent+with+Large+Language+Models
21. In-Context Retrieval-Augmented Language Models — Ram et al., 2023
https://scholar.google.com/scholar?q=In-Context+Retrieval-Augmented+Language+Models
22. Scaling Laws for Associative Memories — Ramsauer et al., 2021
https://scholar.google.com/scholar?q=Scaling+Laws+for+Associative+Memories
23. Do Machine Learning Models Memorize or Generalize? — Feldman, 2020
https://scholar.google.com/scholar?q=Do+Machine+Learning+Models+Memorize+or+Generalize%3F
24. LongTableBench: benchmarking long-context table reasoning across real-world formats and domains — approximate, 2024-2025
https://scholar.google.com/scholar?q=LongTableBench%3A+benchmarking+long-context+table+reasoning+across+real-world+formats+and+domains
25. Reasoning-Focused Evaluation of Efficient Long-Context Inference Techniques — approximate, 2024-2025
https://scholar.google.com/scholar?q=Reasoning-Focused+Evaluation+of+Efficient+Long-Context+Inference+Techniques
26. Cognitive Workspace: Active Memory Management for LLMs—An Empirical Study of Functional Infinite Context — approximate, 2024-2025
https://scholar.google.com/scholar?q=Cognitive+Workspace%3A+Active+Memory+Management+for+LLMs%E2%80%94An+Empirical+Study+of+Functional+Infinite+Context
27. Continual learning and catastrophic forgetting — approximate, 2020s
https://scholar.google.com/scholar?q=Continual+learning+and+catastrophic+forgetting
28. ESSENTIAL: Episodic and Semantic Memory Integration for Video Class-Incremental Learning — approximate, 2024-2025
https://scholar.google.com/scholar?q=ESSENTIAL%3A+Episodic+and+Semantic+Memory+Integration+for+Video+Class-Incremental+Learning
29. Adaptive compression as a unifying framework for episodic and semantic memory — approximate, 2020s
https://scholar.google.com/scholar?q=Adaptive+compression+as+a+unifying+framework+for+episodic+and+semantic+memory
30. MAG: Memory Augmented Knowledge Extraction Generation for Large Language Models — approximate, 2024-2025
https://scholar.google.com/scholar?q=MAG%3A+Memory+Augmented+Knowledge+Extraction+Generation+for+Large+Language+Models
31. Gradient Descent at Inference Time for LLM Reasoning
https://podcast.do-not-panic.com/episodes/2026-03-10-gradient-descent-at-inference-time-for-l-20617d.mp3
32. LLM Agents Reason About Code Without Running It
https://podcast.do-not-panic.com/episodes/2026-03-15-llm-agents-reason-about-code-without-run-2a1876.mp3
33. Emergent Cooperation in Self-Interested Multi-Agent AI
https://podcast.do-not-panic.com/episodes/2026-03-13-emergent-cooperation-in-self-interested-9c0b4c.mp3
34. 50x KV Cache Compression in Seconds via Attention Matching
https://podcast.do-not-panic.com/episodes/2026-03-09-50x-kv-cache-compression-in-seconds-via-9402c1.mp3
35. DualPath Breaks Storage Bandwidth Bottleneck in Agentic Inference
https://podcast.do-not-panic.com/episodes/2026-03-07-dualpath-breaks-storage-bandwidth-bottle-bc9a82.mp3
36. NVIDIA Nemotron 3 Hybrid SSM Transformer Architecture
https://podcast.do-not-panic.com/episodes/2026-03-07-nvidia-nemotron-3-hybrid-ssm-transformer-f9a91b.mp3
37. Structured State Space Duality Unifies Transformers and SSMs
https://podcast.do-not-panic.com/episodes/2026-03-07-structured-state-space-duality-unifies-t-bb2659.mp3
38. Why CARTRIDGE Works: Keys as Routers in KV Caches
https://podcast.do-not-panic.com/episodes/2026-03-07-why-cartridge-works-keys-as-routers-in-k-887d13.mp3
Interactive Visualization: Memory in the Age of AI Agents: Forms, Functions, Dynamics

This episode explores the xLLM Technical Report from JD.com, which describes a production system for running large language model inference at enterprise scale. The core technical challenge is co-locating online chatbot workloads with offline batch jobs on the same GPU cluster to maximize hardware utilization during traffic valleys, while maintaining strict latency guarantees during peak hours. The discussion covers key architectural decisions including dynamic prefill-decode disaggregation, which adaptively reallocates compute resources between the prompt processing phase and token generation phase based on real-time workload characteristics, and the management of unpredictable KV cache memory growth that makes LLM co-location harder than traditional cloud workload mixing. Listeners interested in production ML systems engineering, GPU cluster optimization, and the practical challenges of deploying transformer models at scale will find concrete insights into how a major tech company handles the resource scheduling problems that academic papers often overlook.

Sources:
1. xLLM Technical Report — Tongxuan Liu, Tao Peng, Peijun Yang, Xiaoyang Zhao, Xiusheng Lu, Weizhe Huang, Zirui Liu, Xiaoyu Chen, Zhiwei Liang, Jun Xiong, Donghe Jin, Minchao Zhang, Jinrong Guo, Yingxu Deng, Xu Zhang, Xianzhe Dong, Siqi Wang, Siyu Wu, Yu Wu, Zihan Tang, Yuting Zeng, Yanshu Wang, Jinguang Liu, Meng Kang, Menxin Li, Yunlong Wang, Yiming Liu, Xiaolong Ma, Yifan Wang, Yichen Zhang, Jinrun Yin, Keyang Zheng, Jiawei Yin, Jun Zhang, Ziyue Wang, Xiaobo Lin, Liangyu Liu, Liwei Lan, Yang Liu, Chunhua Peng, Han Liu, Songcheng Ren, Xuezhu Wang, Yunheng Shen, Yi Wang, Guyue Liu, Yitao Hu, Hui Chen, Tong Yang, Hailong Yang, Jing Li, Guiguang Ding, Ke Zhang, 2025
http://arxiv.org/abs/2510.14686v2
2. Borg, Omega, and Kubernetes: Lessons learned from three container-management systems over a decade — Brendan Burns, Brian Grant, David Oppenheimer, Eric Brewer, John Wilkes, 2016
https://scholar.google.com/scholar?q=Borg%2C+Omega%2C+and+Kubernetes%3A+Lessons+learned+from+three+container-management+systems+over+a+decade
3. SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills — Amey Agrawal, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav S. Gulavani, Ramachandran Ramjee, 2023
https://scholar.google.com/scholar?q=SARATHI%3A+Efficient+LLM+Inference+by+Piggybacking+Decodes+with+Chunked+Prefills
4. Paella: Low-latency Model Serving with Virtualized GPU Scheduling — Yiying Zhang, et al., 2024
https://scholar.google.com/scholar?q=Paella%3A+Low-latency+Model+Serving+with+Virtualized+GPU+Scheduling
5. Synergy: A QoS-aware Multi-DNN Inference Framework for Heterogeneous Systems-on-Chip — Jie Ren, et al., 2022
https://scholar.google.com/scholar?q=Synergy%3A+A+QoS-aware+Multi-DNN+Inference+Framework+for+Heterogeneous+Systems-on-Chip
6. TensorFlow: A System for Large-Scale Machine Learning — Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, et al., 2016
https://scholar.google.com/scholar?q=TensorFlow%3A+A+System+for+Large-Scale+Machine+Learning
7. XLA: Optimizing Compiler for Machine Learning — The TensorFlow team (Google), 2017
https://scholar.google.com/scholar?q=XLA%3A+Optimizing+Compiler+for+Machine+Learning
8. TVM: An Automated End-to-End Optimizing Compiler for Deep Learning — Tianqi Chen, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Yan, Haichen Shen, Meghan Cowan, et al., 2018
https://scholar.google.com/scholar?q=TVM%3A+An+Automated+End-to-End+Optimizing+Compiler+for+Deep+Learning
9. torch.compile: A Compiler for PyTorch Models (PyTorch 2.0 Technical Report) — Jason Ansel, Edward Yang, Horace He, Natalia Gimelshein, et al. (Meta), 2023
https://scholar.google.com/scholar?q=torch.compile%3A+A+Compiler+for+PyTorch+Models+%28PyTorch+2.0+Technical+Report%29
10. Orca: A Distributed Serving System for Transformer-Based Generative Models — Yu et al., 2022
https://scholar.google.com/scholar?q=Orca%3A+A+Distributed+Serving+System+for+Transformer-Based+Generative+Models
11. FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU — Sheng et al., 2023
https://scholar.google.com/scholar?q=FlexGen%3A+High-Throughput+Generative+Inference+of+Large+Language+Models+with+a+Single+GPU
12. DeepSpeed-MII: Instant Speedup on Inference with Advanced Model Implementations and System Techniques — Microsoft, 2023
https://scholar.google.com/scholar?q=DeepSpeed-MII%3A+Instant+Speedup+on+Inference+with+Advanced+Model+Implementations+and+System+Techniques
13. AlpaServe: Statistical Multiplexing with Model Parallelism for Deep Learning Serving — Li et al., 2023
https://scholar.google.com/scholar?q=AlpaServe%3A+Statistical+Multiplexing+with+Model+Parallelism+for+Deep+Learning+Serving
14. Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads — Cai et al., 2024
https://scholar.google.com/scholar?q=Medusa%3A+Simple+LLM+Inference+Acceleration+Framework+with+Multiple+Decoding+Heads
15. POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM Inference — approximate 2025, 2025
https://scholar.google.com/scholar?q=POD-Attention%3A+Unlocking+Full+Prefill-Decode+Overlap+for+Faster+LLM+Inference
16. FlowPrefill: Decoupling Preemption from Prefill Scheduling Granularity to Mitigate Head-of-Line Blocking in LLM Serving — approximate 2025, 2025
https://scholar.google.com/scholar?q=FlowPrefill%3A+Decoupling+Preemption+from+Prefill+Scheduling+Granularity+to+Mitigate+Head-of-Line+Blocking+in+LLM+Serving
17. KvLink: Accelerating Large Language Models via Efficient KV Cache Reuse — approximate 2024-2025, 2024
https://scholar.google.com/scholar?q=KvLink%3A+Accelerating+Large+Language+Models+via+Efficient+KV+Cache+Reuse
18. Decoding Speculative Decoding — approximate 2024, 2024
https://scholar.google.com/scholar?q=Decoding+Speculative+Decoding
19. A Comprehensive Survey of Mixture-of-Experts — approximate 2024-2025, 2024
https://scholar.google.com/scholar?q=A+Comprehensive+Survey+of+Mixture-of-Experts
20. DualPath Breaks Storage Bandwidth Bottleneck in Agentic Inference
https://podcast.do-not-panic.com/episodes/2026-03-07-dualpath-breaks-storage-bandwidth-bottle-bc9a82.mp3
21. FlashAttention-4 Conquers Asymmetric GPU Hardware Scaling
https://podcast.do-not-panic.com/episodes/2026-03-06-flashattention-4-conquers-asymmetric-gpu-78839b.mp3
22. 50x KV Cache Compression in Seconds via Attention Matching
https://podcast.do-not-panic.com/episodes/2026-03-09-50x-kv-cache-compression-in-seconds-via-9402c1.mp3
23. Systematic Characterization of LLM Inference on GPUs
https://podcast.do-not-panic.com/episodes/2026-03-07-systematic-characterization-of-llm-infer-f7b9a8.mp3
Interactive Visualization: xLLM: Co-Locating Online and Offline LLM Inference

This episode examines the MATT (Model-Aware Tokenizer Transfer) paper from AGH University of Krakow, which proposes a fundamentally different approach to extending language models to underserved languages. Using Georgian as the central case study, the episode explains tokenizer fertility — how tokenizers optimized for high-resource languages fragment Georgian words into six to eight subword pieces, consuming context budget and degrading both accuracy and inference speed. The episode traces the lineage of tokenizer transfer methods from WECHSEL through FOCUS and ZETT, each of which initializes new embeddings by finding semantically similar source tokens via bilingual dictionaries or FastText projections. MATT's contribution — Attention-Informed Mapping (AIM) — reframes the problem: rather than asking which source tokens are semantically closest, it asks which embeddings are most compatible with what the model's attention layers already know how to route. This is grounded in mechanistic interpretability research showing that factual knowledge resides in FFN layers, not embeddings, making tokenizer swap feasible in principle. The episode includes a detailed comparison with the Cartridges approach, which tackles a closely related problem from a different architectural angle. Four parallel threads are developed: the Structured Continual Initialization parallel, the key-as-router insight, the separation of FFN knowledge from attention routing, and the FFN token-ID binding risk that MATT's evaluation never directly probes. The discussion argues this last point represents the sharpest untested assumption in the paper — whether feed-forward layers develop token-specific associations that break silently when vocabulary changes.

This episode examines a Meta-led paper that develops the first systematic scaling laws for reinforcement learning in large language models, based on over 400,000 GPU-hours of experiments. The researchers propose a sigmoid framework to predict RL performance at large compute budgets from early training runs, addressing a critical gap in the field—while pre-training has well-established power-law relationships like Chinchilla scaling, RL has lacked predictive models due to shifting data distributions and bounded reward functions. The work focuses on mathematical reasoning tasks using AIME problems and introduces ScaleRL, a best-practice training recipe that successfully extrapolates performance from 50,000 to 100,000 GPU-hours. However, the hosts raise important questions about generalization beyond math to domains like code and dialogue, and whether the smooth sigmoid curves capture potential phase transitions or emergent capabilities that might appear at higher compute scales.

Sources:
1. The Art of Scaling Reinforcement Learning Compute for LLMs — Devvrit Khatri, Lovish Madaan, Rishabh Tiwari, Rachit Bansal, Sai Surya Duvvuri, Manzil Zaheer, Inderjit S. Dhillon, David Brandfonbrener, Rishabh Agarwal, 2025
http://arxiv.org/abs/2510.13786v1
2. Scaling Laws for Neural Language Models — Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, et al., 2020
https://scholar.google.com/scholar?q=Scaling+Laws+for+Neural+Language+Models
3. Training Compute-Optimal Large Language Models (Chinchilla) — Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, et al., 2022
https://scholar.google.com/scholar?q=Training+Compute-Optimal+Large+Language+Models+%28Chinchilla%29
4. Scaling Laws for Reward Model Overoptimization — Leo Gao, John Schulman, Jacob Hilton, 2023
https://scholar.google.com/scholar?q=Scaling+Laws+for+Reward+Model+Overoptimization
5. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning — DeepSeek-AI Team (Guo et al.), 2025
https://scholar.google.com/scholar?q=DeepSeek-R1%3A+Incentivizing+Reasoning+Capability+in+LLMs+via+Reinforcement+Learning
6. Proximal Policy Optimization Algorithms — John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, Oleg Klimov, 2017
https://scholar.google.com/scholar?q=Proximal+Policy+Optimization+Algorithms
7. Direct Preference Optimization: Your Language Model is Secretly a Reward Model — Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, Chelsea Finn, 2023
https://scholar.google.com/scholar?q=Direct+Preference+Optimization%3A+Your+Language+Model+is+Secretly+a+Reward+Model
8. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models — Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, Denny Zhou, 2022
https://scholar.google.com/scholar?q=Chain-of-Thought+Prompting+Elicits+Reasoning+in+Large+Language+Models
9. STaR: Self-Taught Reasoner Bootstrapping Reasoning With Reasoning — Eric Zelikman, Yuhuai Wu, Jesse Mu, Noah D. Goodman, 2022
https://scholar.google.com/scholar?q=STaR%3A+Self-Taught+Reasoner+Bootstrapping+Reasoning+With+Reasoning
10. Let's Verify Step by Step — Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, Karl Cobbe, 2023
https://scholar.google.com/scholar?q=Let%27s+Verify+Step+by+Step
11. Reward Model Ensembles Help Mitigate Overoptimization — Thomas Coste, Usman Anwar, Robert Kirk, David Krueger, 2023
https://scholar.google.com/scholar?q=Reward+Model+Ensembles+Help+Mitigate+Overoptimization
12. DAPO: Data-Adaptive Policy Optimization — Yu et al., 2025
https://scholar.google.com/scholar?q=DAPO%3A+Data-Adaptive+Policy+Optimization
13. Chinchilla: Training Compute-Optimal Large Language Models — Hoffmann et al., 2022
https://scholar.google.com/scholar?q=Chinchilla%3A+Training+Compute-Optimal+Large+Language+Models
14. OpenAI o1 System Card — OpenAI, 2024
https://scholar.google.com/scholar?q=OpenAI+o1+System+Card
15. Direct Preference Optimization — Rafailov et al., 2023
https://scholar.google.com/scholar?q=Direct+Preference+Optimization
16. STaR: Self-Taught Reasoner — Zelikman et al., 2022
https://scholar.google.com/scholar?q=STaR%3A+Self-Taught+Reasoner
17. REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization — approximate authors unknown from snippet, recent (post-2024 based on context)
https://scholar.google.com/scholar?q=REINFORCE%2B%2B%3A+Stabilizing+Critic-Free+Policy+Optimization+with+Global+Advantage+Normalization
18. Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models — approximate authors unknown from snippet, recent (likely 2024-2025)
https://scholar.google.com/scholar?q=Asynchronous+RLHF%3A+Faster+and+More+Efficient+Off-Policy+RL+for+Language+Models
19. Tapered Off-Policy REINFORCE: Stable and Efficient Reinforcement Learning for LLMs — approximate authors unknown from snippet, recent (likely 2025)
https://scholar.google.com/scholar?q=Tapered+Off-Policy+REINFORCE%3A+Stable+and+Efficient+Reinforcement+Learning+for+LLMs
20. A Survey of Post-Training Scaling in Large Language Models — approximate authors unknown from snippet, recent
https://scholar.google.com/scholar?q=A+Survey+of+Post-Training+Scaling+in+Large+Language+Models
21. Are Emergent Abilities in Large Language Models Just In-Context Learning? — approximate authors unknown from snippet, recent
https://scholar.google.com/scholar?q=Are+Emergent+Abilities+in+Large+Language+Models+Just+In-Context+Learning%3F
22. AI Post Transformers: Gradient Descent at Inference Time for LLM Reasoning — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-10-gradient-descent-at-inference-time-for-l-20617d.mp3
23. AI Post Transformers: Emergent Cooperation in Self-Interested Multi-Agent AI — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-13-emergent-cooperation-in-self-interested-9c0b4c.mp3
Interactive Visualization: The Art of Scaling Reinforcement Learning Compute for LLMs

This episode of AI Post Transformers examines "Agentic Code Reasoning" by Shubham Ugare and Satish Chandra from Meta, which introduces semi-formal reasoning certificates as an inference-time scaffold for LLM agents analyzing code without executing it. Rather than letting a model produce free-form chain-of-thought verdicts, the certificate framework requires the agent to state explicit premises, trace execution paths through real repository code, and produce a structured, auditable reasoning record for every claim it makes about code behavior. The Django bug django-13670 — involving two-digit year formatting for years before 1000 CE — anchors the discussion: two patches both claim to fix the same issue, but unstated assumptions about name resolution cause an unstructured model to misidentify which one is correct. The certificate format forces the agent to chase the actual import chain across modules rather than guess based on a function name, turning premise verification into a natural driver of interprocedural analysis. Hosts Hal Turing and Dr. Ada Shannon situate the paper against the spectrum from fully formal proof assistants like Lean and Coq — which are provably correct but completely impractical for arbitrary repository code — down to unstructured LLM judges like CodeJudge and SWE-RM, which let the model skip edge cases and produce confident wrong answers. The certificate sits between those extremes, imposing enough structure to make implicit assumptions visible without requiring formalized language semantics. The episode traces how the agentic setup amplifies the value of the certificate structure. Using a minimal SWE-agent configuration with bash tool access but no code execution, the agent can navigate the file system, run grep queries, and follow import chains — exploration scope without runtime confirmation. That constraint is precisely where interprocedural tracing becomes load-bearing: the agent cannot run the code to confirm a hypothesis, so it must read the actual call chain to know what a function does rather than infer from its name. The certificate makes that tracing explicit and auditable, which opens a secondary use case beyond RL reward signal generation: automated code review where a human auditor can inspect the agent's reasoning chain rather than accept a black-box verdict. Hal and Ada discuss RL training pipelines as the paper's stated primary motivation — execution-free reward signals could meaningfully reduce the cost of running sandboxed test suites at scale — but are careful to position that as a downstream consequence of the certificate's properties rather than its defining contribution. The episode closes on three open problems the paper leaves unresolved. First, the inference cost gap: the certificate framework adds computation at inference time, but the paper reports no latency measurements, no tokens-per-certificate data, and no comparison against unstructured baselines on cost — making it impossible to assess whether the accuracy gains justify the overhead in production. Second, certificate reuse as a concrete future direction: common interprocedural patterns across a codebase — frequently called utilities, stable library interfaces — could in principle be cached and reused across multiple verification queries, amortizing the inference cost that the paper never measures. Third, verification independence: the paper's circular verification problem remains open, since the same model that generates a certificate is also the model best positioned to judge whether the premises in that certificate are sound. Separating generation from verification — whether through a distinct model, a symbolic checker, or a human auditor — is the structural fix the framework points toward but does not yet provide.

Sources:
1. Agentic Code Reasoning — Shubham Ugare, Satish Chandra, 2026
http://arxiv.org/abs/2603.01896
2. SWE-bench: Can Language Models Resolve Real-World GitHub Issues? — Jimenez et al., 2024
https://scholar.google.com/scholar?q=SWE-bench%3A+Can+Language+Models+Resolve+Real-World+GitHub+Issues%3F
3. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models — Wei et al., 2022
https://scholar.google.com/scholar?q=Chain-of-Thought+Prompting+Elicits+Reasoning+in+Large+Language+Models
4. Program Equivalence — Godlin and Strichman, 2008
https://scholar.google.com/scholar?q=Program+Equivalence
5. LLM-based agents for automated software engineering: A survey — Multiple authors, 2024-2025
https://scholar.google.com/scholar?q=LLM-based+agents+for+automated+software+engineering%3A+A+survey
6. CodeJudge: Evaluating Code Generation with Large Language Models — Tong and Zhang, 2024
https://scholar.google.com/scholar?q=CodeJudge%3A+Evaluating+Code+Generation+with+Large+Language+Models
7. On designing effective RL reward at training time for LLM reasoning — approximate, 2024-2025, 2024-2025
https://scholar.google.com/scholar?q=On+designing+effective+RL+reward+at+training+time+for+LLM+reasoning
8. Large language model critics for execution-free evaluation of code changes — approximate, 2024-2025, 2024-2025
https://scholar.google.com/scholar?q=Large+language+model+critics+for+execution-free+evaluation+of+code+changes
9. AgentFL: Scaling LLM-based fault localization to project-level context — approximate, 2024-2025, 2024-2025
https://scholar.google.com/scholar?q=AgentFL%3A+Scaling+LLM-based+fault+localization+to+project-level+context
10. SoapFL: A Standard Operating Procedure for LLM-based Method-Level Fault Localization — approximate, 2024-2025, 2024-2025
https://scholar.google.com/scholar?q=SoapFL%3A+A+Standard+Operating+Procedure+for+LLM-based+Method-Level+Fault+Localization
11. Structured chain-of-thought prompting for code generation — approximate, 2022-2024, 2022-2024
https://scholar.google.com/scholar?q=Structured+chain-of-thought+prompting+for+code+generation
12. Deductive verification of chain-of-thought reasoning — approximate, 2023-2024, 2023-2024
https://scholar.google.com/scholar?q=Deductive+verification+of+chain-of-thought+reasoning
13. AI Post Transformers: Reasoning About Code Without Running It — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-08-reasoning-about-code-without-running-it-a9d01a.mp3
14. AI Post Transformers: Gradient Descent at Inference Time for LLM Reasoning — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-10-gradient-descent-at-inference-time-for-l-20617d.mp3
Interactive Visualization: LLM Agents Reason About Code Without Running It

This episode explores the challenge of getting self-interested AI agents to cooperate without hardcoding cooperative behavior, examining a 2026 Google paper on multi-agent cooperation through in-context co-player inference. The hosts build up the technical foundations carefully, explaining why standard reinforcement learning breaks down in multi-agent settings due to non-stationarity, and how social dilemmas like the Prisoner's Dilemma cause agents to reliably converge on mutual defection even when cooperation would benefit everyone. The discussion traces the lineage of learning-aware agents, particularly LOLA, which achieved cooperation by differentiating through an opponent's gradient updates — a clever but architecturally demanding approach. The paper under review argues that training a transformer on a diverse pool of co-players lets in-context learning produce emergent cooperation without any of that machinery. Listeners interested in the intersection of game theory, multi-agent RL, and modern sequence modeling will find the episode's careful unpacking of why prior approaches fell short — and what the new framing claims to replace — genuinely illuminating.

Sources:
1. Multi-agent cooperation through in-context co-player inference — Marissa A. Weis, Maciej Wołczyk, Rajai Nasser, Rif A. Saurous, Blaise Agüera y Arcas, João Sacramento, Alexander Meulemans, 2026
http://arxiv.org/abs/2602.16301
2. Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments — Ryan Lowe, Yi Wu, Aviv Tamar, Jean Harb, Pieter Abbeel, Igor Mordatch, 2017
https://scholar.google.com/scholar?q=Multi-Agent+Actor-Critic+for+Mixed+Cooperative-Competitive+Environments
3. QMIX: Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement Learning — Tabish Rashid, Mikayel Samvelyan, Christian Schroeder de Waard, Georgios Papoudakis, Jakob Foerster, Shimon Whiteson, 2018
https://scholar.google.com/scholar?q=QMIX%3A+Monotonic+Value+Function+Factorisation+for+Deep+Multi-Agent+Reinforcement+Learning
4. A Survey and Critique of Multiagent Deep Reinforcement Learning — Pablo Hernandez-Leal, Bilal Kartal, Matthew E. Taylor, 2019
https://scholar.google.com/scholar?q=A+Survey+and+Critique+of+Multiagent+Deep+Reinforcement+Learning
5. Multi-Agent Transformer: Scalable Cooperative Multi-Agent Reinforcement Learning — Muning Wen, Jakub Grudzien Kuba, Runji Lin, Weinan Zhang, Ying Wen, Jun Wang, Yaodong Yang, 2022
https://scholar.google.com/scholar?q=Multi-Agent+Transformer%3A+Scalable+Cooperative+Multi-Agent+Reinforcement+Learning
6. Language Models are Few-Shot Learners — Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, et al., 2020
https://scholar.google.com/scholar?q=Language+Models+are+Few-Shot+Learners
7. An Explanation of In-Context Learning as Implicit Bayesian Inference — Sang Michael Xie, Aditi Raghunathan, Percy Liang, Tengyu Ma, 2022
https://scholar.google.com/scholar?q=An+Explanation+of+In-Context+Learning+as+Implicit+Bayesian+Inference
8. Transformers Learn In-Context by Gradient Descent — Johannes von Oswald, Eyvind Niklasson, Ettore Randazzo, Joao Sacramento, Alexander Mordvintsev, Andrey Zhmoginov, Max Vladymyrov, 2023
https://scholar.google.com/scholar?q=Transformers+Learn+In-Context+by+Gradient+Descent
9. Algorithm Distillation in Reinforcement Learning — Michael Laskin, Luyu Wang, Junhyuk Oh, Emilio Parisotto, Stephen Spencer, Richie Morrill, Abhinav Gupta, Pieter Abbeel, Oriol Vinyals, 2023
https://scholar.google.com/scholar?q=Algorithm+Distillation+in+Reinforcement+Learning
10. Emergent Complexity via Multi-Agent Competition — Trapit Bansal, Jakub Pachocki, Szymon Sidor, Ilya Sutskever, Igor Mordatch, 2018
https://scholar.google.com/scholar?q=Emergent+Complexity+via+Multi-Agent+Competition
11. Emergent Tool Use from Multi-Agent Interaction — Bowen Baker, Ingmar Kanitscheider, Todor Markov, Yi Wu, Glenn Powell, Bob McGrew, Igor Mordatch, 2019
https://scholar.google.com/scholar?q=Emergent+Tool+Use+from+Multi-Agent+Interaction
12. Social Influence as Intrinsic Motivation for Multi-Agent Deep Reinforcement Learning — Natasha Jaques, Angeliki Lazaridou, Edward Hughes, Caglar Gulcehre, Pedro Ortega, DJ Strouse, Joel Z. Leibo, Nando de Freitas, 2019
https://scholar.google.com/scholar?q=Social+Influence+as+Intrinsic+Motivation+for+Multi-Agent+Deep+Reinforcement+Learning
13. Cooperative Multi-Agent Learning: The State of the Art — Liviu Panait, Sean Luke, 2005
https://scholar.google.com/scholar?q=Cooperative+Multi-Agent+Learning%3A+The+State+of+the+Art
14. The Evolution of Cooperation — Robert Axelrod, William D. Hamilton, 1981
https://scholar.google.com/scholar?q=The+Evolution+of+Cooperation
15. Iterated Prisoner's Dilemma Contains Strategies that Dominate Any Evolutionary Opponent — William H. Press, Freeman J. Dyson, 2012
https://scholar.google.com/scholar?q=Iterated+Prisoner%27s+Dilemma+Contains+Strategies+that+Dominate+Any+Evolutionary+Opponent
16. Evolutionary Instability of Zero-Determinant Strategies Demonstrates That Winning Is Not Everything — Christoph Adami, Arend Hintze, 2013
https://scholar.google.com/scholar?q=Evolutionary+Instability+of+Zero-Determinant+Strategies+Demonstrates+That+Winning+Is+Not+Everything
17. Learning with Opponent-Learning Awareness (LOLA) — Jakob Foerster, Richard Y. Chen, Maruan Al-Shedivat, Shimon Whiteson, Pieter Abbeel, Igor Mordatch, 2018
https://scholar.google.com/scholar?q=Learning+with+Opponent-Learning+Awareness+%28LOLA%29
18. Learning with Opponent-Learning Awareness — Jakob Foerster, Richard Y. Chen, Maruan Al-Shedivat, Shimon Whiteson, Pieter Abbeel, Igor Mordatch, 2018
https://scholar.google.com/scholar?q=Learning+with+Opponent-Learning+Awareness
19. Model-Free Opponent Shaping — Chris Lu, Timon Willi, Christian Schroeder de Waard, Jakob Foerster, 2022
https://scholar.google.com/scholar?q=Model-Free+Opponent+Shaping
20. In-context reinforcement learning with algorithm distillation — Laskin, M., Wang, L., Oh, J., Parisotto, E., Spencer, S., Steigerwald, R., Strouse, D., Hansen, S., Filos, A., Brooks, E., Gazeau, M., Sahni, H., Singh, S., Mnih, V., 2023
https://scholar.google.com/scholar?q=In-context+reinforcement+learning+with+algorithm+distillation
21. RL^2: Fast reinforcement learning via slow reinforcement learning — Duan, Y., Schulman, J., Chen, X., Bartlett, P. L., Sutskever, I., Abbeel, P., 2016
https://scholar.google.com/scholar?q=RL%5E2%3A+Fast+reinforcement+learning+via+slow+reinforcement+learning
22. Cooperating with unknown teammates in complex domains by acting carefully with information — Aghajohari, M., Duque, J., Cooijmans, T., Courville, A., 2024
https://scholar.google.com/scholar?q=Cooperating+with+unknown+teammates+in+complex+domains+by+acting+carefully+with+information
23. From naive to learning-aware: Emergence of cooperative behaviors in multi-agent systems — Meulemans, A., et al., 2025
https://scholar.google.com/scholar?q=From+naive+to+learning-aware%3A+Emergence+of+cooperative+behaviors+in+multi-agent+systems
24. Generative agents: Interactive simulacra of human behavior — Park, J. S., O'Brien, J. C., Cai, C. J., Morris, M. R., Liang, P., Bernstein, M. S., 2023
https://scholar.google.com/scholar?q=Generative+agents%3A+Interactive+simulacra+of+human+behavior
25. Do Pre-trained Transformers Really Learn In-context by Gradient Descent? — Shen et al. (approximate), 2023-2024
https://scholar.google.com/scholar?q=Do+Pre-trained+Transformers+Really+Learn+In-context+by+Gradient+Descent%3F
26. When is diversity rewarded in cooperative multi-agent learning? — approximate, exact authors not in snippet, recent
https://scholar.google.com/scholar?q=When+is+diversity+rewarded+in+cooperative+multi-agent+learning%3F
27. The evolution of zero-determinant strategies in public goods game — approximate, exact authors not in snippet, recent
https://scholar.google.com/scholar?q=The+evolution+of+zero-determinant+strategies+in+public+goods+game
28. Uncoupled learning of differential Stackelberg equilibria with commitments — approximate, exact authors not in snippet, recent
https://scholar.google.com/scholar?q=Uncoupled+learning+of+differential+Stackelberg+equilibria+with+commitments
29. Non-coercive extortion in game theory — approximate, exact authors not in snippet, recent
https://scholar.google.com/scholar?q=Non-coercive+extortion+in+game+theory
30. Reciprocal reward influence encourages cooperation from self-interested agents — approximate, exact authors not in snippet, recent
https://scholar.google.com/scholar?q=Reciprocal+reward+influence+encourages+cooperation+from+self-interested+agents
31. AI Post Transformers: In-Context Learning as Implicit Learning Algorithms — Hal Turing & Dr. Ada Shannon, Wed,
https://podcasters.spotify.com/pod/show/12146088098/episodes/In-Context-Learning-as-Implicit-Learning-Algorithms-e39sjmn
32. AI Post Transformers: Zero-Shot Context Generalization in Reinforcement Learning from Few Training Contexts — Hal Turing & Dr. Ada Shannon, Tue,
https://podcasters.spotify.com/pod/show/12146088098/episodes/Zero-Shot-Context-Generalization-in-Reinforcement--Learning-from-Few-Training-Contexts-e3fi0t5
33. AI Post Transformers: Experiential Reinforcement Learning: Internalizing Reflection for Better Policy Training — Hal Turing & Dr. Ada Shannon, Fri,
https://podcasters.spotify.com/pod/show/12146088098/episodes/Experiential-Reinforcement-Learning-Internalizing-Reflection-for-Better-Policy-Training-e3fbel0
Interactive Visualization: Emergent Cooperation in Self-Interested Multi-Agent AI

This episode examines a 2026 MIT paper claiming a 50x KV cache memory reduction that runs in seconds rather than the GPU-hours required by prior latent-space compaction methods. It grounds the claim in a detailed technical primer on KV cache mechanics — explaining why memory consumption scales multiplicatively across layers, heads, and context length, reaching 8–16 GB per request at 64K-token contexts. The discussion traces the compaction landscape from token eviction approaches like H2O and SnapKV, through token merging, to the latent-space paradigm introduced by Cartridges, establishing why earlier methods collapse at extreme compression ratios. The central question is whether "Fast KV Compaction via Attention Matching" genuinely pushes the quality-versus-speed Pareto frontier — making per-request inference-time compaction practical rather than a research pipeline operation. Listeners interested in long-context inference infrastructure, memory-efficient transformers, or the engineering constraints shaping modern LLM deployment will find the technical depth and comparative framing useful.

Sources:
1. Fast KV Compaction via Attention Matching — Adam Zweiger, Xinghong Fu, Han Guo, Yoon Kim, 2026
http://arxiv.org/abs/2602.16284
2. H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models — Zhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen, et al., 2023
https://scholar.google.com/scholar?q=H2O%3A+Heavy-Hitter+Oracle+for+Efficient+Generative+Inference+of+Large+Language+Models
3. SnapKV: LLM Knows What You are Looking for Before Generation — Yuhong Li, Yingbing Huang, Bowen Yang, Bharat Venkitesh, Acyr Locatelli, Hanchen Ye, Tianhao Guo, Patrick Lewis, Deming Chen, 2024
https://scholar.google.com/scholar?q=SnapKV%3A+LLM+Knows+What+You+are+Looking+for+Before+Generation
4. Efficient Streaming Language Models with Attention Sinks — Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, Mike Lewis, 2023
https://scholar.google.com/scholar?q=Efficient+Streaming+Language+Models+with+Attention+Sinks
5. KVMerger: KV Cache Merging for Memory-Efficient LLMs Inference — Cangqing Wang, Yuhang Yang, Liangzhen Li, Lanqing Hong, Shuo Jiang, Hui Xu, Wei Tao, 2024
https://scholar.google.com/scholar?q=KVMerger%3A+KV+Cache+Merging+for+Memory-Efficient+LLMs+Inference
6. Cartridges: Lightweight, Pluggable Contexts for Language Models — Sabri Eyuboglu, Avanika Narayan, Tao Long, Andrew Liang, Kush Bhatia, Michael Zhang, Neel Guha, James Zou, Christopher Re, Atri Rudra, 2025
https://scholar.google.com/scholar?q=Cartridges%3A+Lightweight%2C+Pluggable+Contexts+for+Language+Models
7. Prefix-Tuning: Optimizing Continuous Prompts for Generation — Xiang Lisa Li, Percy Liang, 2021
https://scholar.google.com/scholar?q=Prefix-Tuning%3A+Optimizing+Continuous+Prompts+for+Generation
8. Learning to Compress Prompts with Gist Tokens — Jesse Mu, Xiang Lisa Li, Noah Goodman, 2023
https://scholar.google.com/scholar?q=Learning+to+Compress+Prompts+with+Gist+Tokens
9. DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model — DeepSeek-AI (Zhipeng Liu, Chengqi Deng, et al.), 2024
https://scholar.google.com/scholar?q=DeepSeek-V2%3A+A+Strong%2C+Economical%2C+and+Efficient+Mixture-of-Experts+Language+Model
10. GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers — Elias Frantar, Saleh Ashkboos, Torsten Hoefler, Dan Alistarh, 2022
https://scholar.google.com/scholar?q=GPTQ%3A+Accurate+Post-Training+Quantization+for+Generative+Pre-trained+Transformers
11. SparseGPT: Massive Language Models Can be Accurately Pruned in One Shot — Elias Frantar, Dan Alistarh, 2023
https://scholar.google.com/scholar?q=SparseGPT%3A+Massive+Language+Models+Can+be+Accurately+Pruned+in+One+Shot
12. Optimal Brain Surgeon and General Network Pruning — Babak Hassibi, David G. Stork, 1993
https://scholar.google.com/scholar?q=Optimal+Brain+Surgeon+and+General+Network+Pruning
13. LASER: The Truth is in There: Improving Reasoning in Language Models with Layer-Selective Rank Reduction — Pratyusha Sharma, Jordan T. Ash, Dipendra Misra, 2023
https://scholar.google.com/scholar?q=LASER%3A+The+Truth+is+in+There%3A+Improving+Reasoning+in+Language+Models+with+Layer-Selective+Rank+Reduction
14. Cartridges: Learned KV Cache Compression for Long-Context Language Model Inference — Eyuboglu et al., 2025
https://scholar.google.com/scholar?q=Cartridges%3A+Learned+KV+Cache+Compression+for+Long-Context+Language+Model+Inference
15. The Power of Scale for Parameter-Efficient Prompt Tuning — Lester et al., 2021
https://scholar.google.com/scholar?q=The+Power+of+Scale+for+Parameter-Efficient+Prompt+Tuning
16. MagicPIG: LSH Sampling for Efficient LLM Generation — Chen et al., 2024
https://scholar.google.com/scholar?q=MagicPIG%3A+LSH+Sampling+for+Efficient+LLM+Generation
17. KV-Distill: Nearly Lossless Learnable Context Compression for LLMs — approximate (multiple authors), 2024-2025
https://scholar.google.com/scholar?q=KV-Distill%3A+Nearly+Lossless+Learnable+Context+Compression+for+LLMs
18. Thin Keys, Full Values: Reducing KV Cache via Low-Dimensional Attention Selection — approximate (multiple authors), 2024-2025
https://scholar.google.com/scholar?q=Thin+Keys%2C+Full+Values%3A+Reducing+KV+Cache+via+Low-Dimensional+Attention+Selection
19. A Preliminary Study on the Promises and Challenges of Native Top-Sparse Attention — approximate (multiple authors), 2024-2025
https://scholar.google.com/scholar?q=A+Preliminary+Study+on+the+Promises+and+Challenges+of+Native+Top-Sparse+Attention
20. Beyond KV Caching: Shared Attention for Efficient LLMs — approximate (multiple authors), 2024-2025
https://scholar.google.com/scholar?q=Beyond+KV+Caching%3A+Shared+Attention+for+Efficient+LLMs
21. Compressing Many-Shots in In-Context Learning — approximate (multiple authors), 2024-2025
https://scholar.google.com/scholar?q=Compressing+Many-Shots+in+In-Context+Learning
22. AI Post Transformers: Hyper-Scaling LLM Inference with KV Cache Compression — Hal Turing & Dr. Ada Shannon, Fri,
https://podcasters.spotify.com/pod/show/12146088098/episodes/Hyper-Scaling-LLM-Inference-with-KV-Cache-Compression-e3aalcq
23. AI Post Transformers: ShadowKV: High-Throughput Long-Context LLM Inference — Hal Turing & Dr. Ada Shannon, Wed,
https://podcasters.spotify.com/pod/show/12146088098/episodes/ShadowKV-High-Throughput-Long-Context-LLM-Inference-e38bn17
24. AI Post Transformers: Quest: Query-Aware Sparsity for Efficient LLM Inference — Hal Turing & Dr. Ada Shannon, Fri,
https://podcasters.spotify.com/pod/show/12146088098/episodes/Quest-Query-Aware-Sparsity-for-Efficient-LLM-Inference-e3aat91
25. AI Post Transformers: Long context: Dichotomy of Findings & Status of Research — Hal Turing & Dr. Ada Shannon, Wed,
https://podcasters.spotify.com/pod/show/12146088098/episodes/Long-context-Dichotomy-of-Findings--Status-of-Research-e3eat7c
26. AI Post Transformers: NVIDIA: TTT-E2E: Unlocking Long-Context Learning via End-to-End Test-Time Training — Hal Turing & Dr. Ada Shannon, Sat,
https://podcasters.spotify.com/pod/show/12146088098/episodes/NVIDIA-TTT-E2E-Unlocking-Long-Context-Learning-via-End-to-End-Test-Time-Training-e3dq389
Interactive Visualization: 50x KV Cache Compression in Seconds via Attention Matching

This episode examines the nabla-Reasoner paper (ICLR 2026), which proposes running gradient descent on token logits during inference — a first-order approach to test-time compute scaling that stands apart from every existing method in the field. The hosts contextualize the work against the established zeroth-order inference-time scaling landscape: Chain-of-Thought, Self-Consistency, Tree of Thoughts, and MCTS-based methods, all of which probe the reward landscape by sampling without directional information. The core argument is that zeroth-order methods hit a hard ceiling on long-horizon reasoning tasks because the search space grows exponentially while reward signals remain sparse, making random sampling increasingly futile. nabla-Reasoner sidesteps this by treating token logit vectors — normally ephemeral intermediate computations — as continuous optimization variables, computing reward gradients with respect to them and nudging the distribution toward higher-reward outputs before committing to each token. Listeners interested in the mechanics of inference-time scaling and the theoretical limits of sampling-based reasoning will find this a technically dense, well-grounded discussion of a genuinely novel approach.

Sources:
1. $\nabla$-Reasoner: LLM Reasoning via Test-Time Gradient Descent in Latent Space — Peihao Wang, Ruisi Cai, Zhen Wang, Hongyuan Mei, Qiang Liu, Pan Li, Zhangyang Wang, 2026
http://arxiv.org/abs/2603.04948v1
2. ∇-Reasoner: LLM Reasoning via Test-Time Gradient Descent in Latent Space — Peihao Wang, Ruisi Cai, Zhen Wang, Hongyuan Mei, Qiang Liu, Pan Li, Zhangyang Wang, 2026
https://scholar.google.com/scholar?q=%E2%88%87-Reasoner%3A+LLM+Reasoning+via+Test-Time+Gradient+Descent+in+Latent+Space
3. Diffusion-LM Improves Controllable Text Generation — Xiang Lisa Li, John Thickstun, Ishaan Gulrajani, Percy Liang, Tatsunori Hashimoto, 2022
https://scholar.google.com/scholar?q=Diffusion-LM+Improves+Controllable+Text+Generation
4. GFlowNet-Guided LLM Decoding: Towards Diverse and Accurate Reasoning — Jianing Li et al., 2024
https://scholar.google.com/scholar?q=GFlowNet-Guided+LLM+Decoding%3A+Towards+Diverse+and+Accurate+Reasoning
5. Back to Basics: Revisiting REINFORCE-Style Optimization for Learning from Human Feedback in LLMs — Ahmadianshalchi et al., 2024
https://scholar.google.com/scholar?q=Back+to+Basics%3A+Revisiting+REINFORCE-Style+Optimization+for+Learning+from+Human+Feedback+in+LLMs
6. Self-Consistency Improves Chain of Thought Reasoning in Language Models — Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, Denny Zhou, 2022
https://scholar.google.com/scholar?q=Self-Consistency+Improves+Chain+of+Thought+Reasoning+in+Language+Models
7. Tree of Thoughts: Deliberate Problem Solving with Large Language Models — Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, Karthik Narasimhan, 2023
https://scholar.google.com/scholar?q=Tree+of+Thoughts%3A+Deliberate+Problem+Solving+with+Large+Language+Models
8. Let's Verify Step by Step — Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, Karl Cobbe, 2023
https://scholar.google.com/scholar?q=Let%27s+Verify+Step+by+Step
9. Scaling LLM Test-Time Compute Optimally — Charlie Snell, Jaehoon Lee, Kelvin Xu, Aviral Kumar, 2024
https://scholar.google.com/scholar?q=Scaling+LLM+Test-Time+Compute+Optimally
10. ARGS: Alignment as Reward-Guided Search — Maxim Khanov, Jirayu Burapacheep, Yixuan Li, 2024
https://scholar.google.com/scholar?q=ARGS%3A+Alignment+as+Reward-Guided+Search
11. Controlled Decoding from Language Models — Sidharth Mudgal, Jong Lee, Harish Ganapathy, YaGuang Li, Tao Wang, Yanpin Huang, Zhifeng Chen, Heng-Tze Cheng, Michael Collins, Trevor Strohman, Jilin Chen, Alex Beutel, Ahmad Beirami, 2024
https://scholar.google.com/scholar?q=Controlled+Decoding+from+Language+Models
12. AlphaCode 2 Technical Report — Google DeepMind AlphaCode Team, 2023
https://scholar.google.com/scholar?q=AlphaCode+2+Technical+Report
13. Training language models to follow instructions with human feedback — Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, Ryan Lowe, 2022
https://scholar.google.com/scholar?q=Training+language+models+to+follow+instructions+with+human+feedback
14. Direct Preference Optimization: Your Language Model is Secretly a Reward Model — Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, Chelsea Finn, 2023
https://scholar.google.com/scholar?q=Direct+Preference+Optimization%3A+Your+Language+Model+is+Secretly+a+Reward+Model
15. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning — DeepSeek-AI (Daya Guo, Dejian Yang, Haowei Zhang, et al.), 2025
https://scholar.google.com/scholar?q=DeepSeek-R1%3A+Incentivizing+Reasoning+Capability+in+LLMs+via+Reinforcement+Learning
16. Learning to summarize from human feedback — Nisan Stiennon, Long Ouyang, Jeff Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, Paul Christiano, 2020
https://scholar.google.com/scholar?q=Learning+to+summarize+from+human+feedback
17. Plug and Play Language Models: A Simple Approach to Controlled Text Generation — Dathathri et al., 2020
https://scholar.google.com/scholar?q=Plug+and+Play+Language+Models%3A+A+Simple+Approach+to+Controlled+Text+Generation
18. FUDGE: Controlled Text Generation with Future Discriminators — Yang and Klein, 2021
https://scholar.google.com/scholar?q=FUDGE%3A+Controlled+Text+Generation+with+Future+Discriminators
19. Scaling LLM Test-Time Compute Optimally Can be More Effective than Scaling Model Parameters — Snell et al., 2024
https://scholar.google.com/scholar?q=Scaling+LLM+Test-Time+Compute+Optimally+Can+be+More+Effective+than+Scaling+Model+Parameters
20. Alignment as Reward-Guided Search — Khanov et al., 2024
https://scholar.google.com/scholar?q=Alignment+as+Reward-Guided+Search
21. Soft Prompts: The Power of Scale for Parameter-Efficient Prompt Tuning — Lester et al., 2021
https://scholar.google.com/scholar?q=Soft+Prompts%3A+The+Power+of+Scale+for+Parameter-Efficient+Prompt+Tuning
22. The Generalization Gap in Offline Reinforcement Learning — Levine et al., 2020
https://scholar.google.com/scholar?q=The+Generalization+Gap+in+Offline+Reinforcement+Learning
23. Thinking on the Fly: Test-Time Reasoning Enhancement via Latent Thought Policy Optimization — Unknown (2025), 2025
https://scholar.google.com/scholar?q=Thinking+on+the+Fly%3A+Test-Time+Reasoning+Enhancement+via+Latent+Thought+Policy+Optimization
24. Logit arithmetic elicits long reasoning capabilities without training — Unknown (2025), 2025
https://scholar.google.com/scholar?q=Logit+arithmetic+elicits+long+reasoning+capabilities+without+training
25. Reinforcement Learning in Inference Time: A Perspective from Successive Policy Iterations — Unknown (2025), 2025
https://scholar.google.com/scholar?q=Reinforcement+Learning+in+Inference+Time%3A+A+Perspective+from+Successive+Policy+Iterations
26. GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning — Yang et al. (2024/2025), 2025
https://scholar.google.com/scholar?q=GenPRM%3A+Scaling+Test-Time+Compute+of+Process+Reward+Models+via+Generative+Reasoning
27. Process Reward Models That Think — Unknown (2025), 2025
https://scholar.google.com/scholar?q=Process+Reward+Models+That+Think
28. Efficient Adaptive Rejection Sampling for Accelerating Speculative Decoding in Large Language Models — Unknown (2025), 2025
https://scholar.google.com/scholar?q=Efficient+Adaptive+Rejection+Sampling+for+Accelerating+Speculative+Decoding+in+Large+Language+Models
29. Inference-time alignment control for diffusion models with reinforcement learning guidance — Unknown (2025), 2025
https://scholar.google.com/scholar?q=Inference-time+alignment+control+for+diffusion+models+with+reinforcement+learning+guidance
30. AI Post Transformers: Test-Time Reinforcement Learning for LLMs — Hal Turing & Dr. Ada Shannon, Wed,
https://podcasters.spotify.com/pod/show/12146088098/episodes/Test-Time-Reinforcement-Learning-for-LLMs-e398hsk
31. AI Post Transformers: MetaScale: Test-Time Scaling with Evolving Meta-Thoughts — Hal Turing & Dr. Ada Shannon, Fri,
https://podcasters.spotify.com/pod/show/12146088098/episodes/MetaScale-Test-Time-Scaling-with-Evolving-Meta-Thoughts-e36kgn7
32. AI Post Transformers: Process Reward Learning for LLM Reasoning Optimization — Hal Turing & Dr. Ada Shannon, Mon,
https://podcasters.spotify.com/pod/show/12146088098/episodes/Process-Reward-Learning-for-LLM-Reasoning-Optimization-e3dsuav
33. AI Post Transformers: Tree-based Group Policy Optimization for LLM Agents — Hal Turing & Dr. Ada Shannon, Fri,
https://podcasters.spotify.com/pod/show/12146088098/episodes/Tree-based-Group-Policy-Optimization-for-LLM-Agents-e38obfb
34. AI Post Transformers: MASA: Meta-Awareness via Self-Alignment Reinforcement Learning — Hal Turing & Dr. Ada Shannon, Sun,
https://podcasters.spotify.com/pod/show/12146088098/episodes/MASA-Meta-Awareness-via-Self-Alignment-Reinforcement-Learning-e3a2of7
Interactive Visualization: Gradient Descent at Inference Time for LLM Reasoning

This episode explores MalGEN, a multi-agent AI framework developed by researchers at IIT Kanpur that autonomously generates novel, functional malware capable of evading modern detection systems. The discussion examines why LLM-generated malware represents a qualitative shift beyond traditional polymorphic and metamorphic techniques — rather than mutating a fixed payload syntactically, LLMs reason about semantic intent and produce entirely new code that achieves the same effect through different computational paths. A central focus is MalGEN's alignment with MITRE ATT&CK tactics, techniques, and procedures, meaning the generated malware maps to documented real-world intrusion patterns rather than merely bypassing signature databases. The hosts pressure-test the paper's red teaming justification — framing MalGEN as a defensive stress-testing tool — while examining its most unsettling capability: automating sandbox-aware, environment-detecting evasion previously requiring nation-state-level expertise. The conversation anchors on the dual-use tension at the core of publishing a reproducible malware generation framework under academic cover.

Sources:
1. MalGEN: A Generative Agent Framework for Modeling Malicious Software in Cybersecurity — Bikash Saha, Sandeep Kumar Shukla, 2025
http://arxiv.org/abs/2506.07586
2. https://arxiv.org/pdf/2510.23883
3. https://arxiv.org/pdf/2601.05293
4. https://arxiv.org/pdf/2508.05674
5. Evaluating the Cybersecurity Capabilities of LLMs: A Comprehensive Study — Bhatt, M., Chennabasappa, S., Nikolaidis, C., et al. (Meta), 2023
https://scholar.google.com/scholar?q=Evaluating+the+Cybersecurity+Capabilities+of+LLMs%3A+A+Comprehensive+Study
6. From Chatbots to Phishbots? Phishing Scam Generation in Commercial Large Language Models — Heiding, F., Schneier, B., Vishwanath, A., Bernstein, J., Park, P.S., 2024
https://scholar.google.com/scholar?q=From+Chatbots+to+Phishbots%3F+Phishing+Scam+Generation+in+Commercial+Large+Language+Models
7. PentestGPT: An LLM-Empowered Automatic Penetration Testing Framework — Deng, G., Liu, Y., Mayoral-Vilches, V., et al., 2024
https://scholar.google.com/scholar?q=PentestGPT%3A+An+LLM-Empowered+Automatic+Penetration+Testing+Framework
8. LLM4Decompile: Decompiling Binary Code with Large Language Models — Tan, Z., Ma, H., Xu, H., et al., 2024
https://scholar.google.com/scholar?q=LLM4Decompile%3A+Decompiling+Binary+Code+with+Large+Language+Models
9. LLM Agents can Autonomously Exploit One-Day Vulnerabilities — Fang, R., Bindu, R., Gupta, A., Kang, D., Boneh, D., 2024
https://scholar.google.com/scholar?q=LLM+Agents+can+Autonomously+Exploit+One-Day+Vulnerabilities
10. Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned — Ganguli, D., Lovitt, L., Kernion, J., et al. (Anthropic), 2022
https://scholar.google.com/scholar?q=Red+Teaming+Language+Models+to+Reduce+Harms%3A+Methods%2C+Scaling+Behaviors%2C+and+Lessons+Learned
11. Trojan Detection Benchmark (TrojAI): Evaluating Backdoor Defenses on Neural Networks — Karra, K., Ashcraft, C., Fendley, N. (IARPA / Johns Hopkins APL), 2020
https://scholar.google.com/scholar?q=Trojan+Detection+Benchmark+%28TrojAI%29%3A+Evaluating+Backdoor+Defenses+on+Neural+Networks
12. Evading Machine Learning Malware Detection — Grosse, K., Papernot, N., Manoharan, P., Backes, M., McDaniel, P., 2017
https://scholar.google.com/scholar?q=Evading+Machine+Learning+Malware+Detection
13. Malware Detection by Eating a Whole EXE — Raff, E., Barker, J., Sylvester, J., Brandon, R., Catanzaro, B., Nicholas, C., 2018
https://scholar.google.com/scholar?q=Malware+Detection+by+Eating+a+Whole+EXE
14. Adversarial Malware Binaries: Evading Deep Learning for Malware Detection in Executables — Kolosnjaji, B., Demontis, A., Biggio, B., Maiorca, D., Giacinto, G., Eckert, C., Roli, F., 2018
https://scholar.google.com/scholar?q=Adversarial+Malware+Binaries%3A+Evading+Deep+Learning+for+Malware+Detection+in+Executables
15. DQEAF: Malware Adversarial Examples Generation with Reinforcement Learning — Fang, Y., Liu, Y., Huang, C., Liu, L., 2020
https://scholar.google.com/scholar?q=DQEAF%3A+Malware+Adversarial+Examples+Generation+with+Reinforcement+Learning
16. MalGAN: Generating Adversarial Malware Examples — Hu, W., Tan, Y., 2017
https://scholar.google.com/scholar?q=MalGAN%3A+Generating+Adversarial+Malware+Examples
17. IDSGAN: Generative Adversarial Networks for Attack Generation against Intrusion Detection — Lin, Z., Shi, Y., Xue, Z., 2022
https://scholar.google.com/scholar?q=IDSGAN%3A+Generative+Adversarial+Networks+for+Attack+Generation+against+Intrusion+Detection
18. Synthetic Data Generation for Cybersecurity: Challenges and Opportunities — Ring, M., Wunderlich, S., Scheuring, D., Landes, D., Hotho, A., 2019
https://scholar.google.com/scholar?q=Synthetic+Data+Generation+for+Cybersecurity%3A+Challenges+and+Opportunities
19. Generative Adversarial Networks for Black-Box API Attack on Web Application Firewalls — Duy, P.T., Khoa, N.M., Hoa, N.T., et al., 2022
https://scholar.google.com/scholar?q=Generative+Adversarial+Networks+for+Black-Box+API+Attack+on+Web+Application+Firewalls
20. Evading Anti-Malware Engines With Deep Reinforcement Learning — Fang, Z. et al., 2019
https://scholar.google.com/scholar?q=Evading+Anti-Malware+Engines+With+Deep+Reinforcement+Learning
21. Do LLMs Dream of Malware? Assessing the Practical Capability of LLMs to Generate Functional Malware — Pa Pa, Y. M. et al., 2023
https://scholar.google.com/scholar?q=Do+LLMs+Dream+of+Malware%3F+Assessing+the+Practical+Capability+of+LLMs+to+Generate+Functional+Malware
22. MITRE ATT&CK: Design and Philosophy — Strom, B. et al., 2018
https://scholar.google.com/scholar?q=MITRE+ATT%26CK%3A+Design+and+Philosophy
23. LLM4Fuzz: Guided Fuzzing of Smart Contracts with Large Language Models — Various, 2024
https://scholar.google.com/scholar?q=LLM4Fuzz%3A+Guided+Fuzzing+of+Smart+Contracts+with+Large+Language+Models
24. Beyond the sandbox: Leveraging symbolic execution for evasive malware classification — approximate, ~2024, 2024
https://scholar.google.com/scholar?q=Beyond+the+sandbox%3A+Leveraging+symbolic+execution+for+evasive+malware+classification
25. Unveiling the dynamic landscape of malware sandboxing: A comprehensive review — approximate, ~2024, 2024
https://scholar.google.com/scholar?q=Unveiling+the+dynamic+landscape+of+malware+sandboxing%3A+A+comprehensive+review
26. TAMAS: Benchmarking Adversarial Risks in Multi-Agent LLM Systems — approximate, ~2024-2025, 2025
https://scholar.google.com/scholar?q=TAMAS%3A+Benchmarking+Adversarial+Risks+in+Multi-Agent+LLM+Systems
27. Red-teaming LLM multi-agent systems via communication attacks — approximate, ~2024-2025, 2025
https://scholar.google.com/scholar?q=Red-teaming+LLM+multi-agent+systems+via+communication+attacks
28. Assessing LLMs in malicious code deobfuscation of real-world malware campaigns — approximate, ~2024-2025, 2024
https://scholar.google.com/scholar?q=Assessing+LLMs+in+malicious+code+deobfuscation+of+real-world+malware+campaigns
29. Certifying accuracy, privacy, and robustness of ML-based malware detection — approximate, ~2024-2025, 2025
https://scholar.google.com/scholar?q=Certifying+accuracy%2C+privacy%2C+and+robustness+of+ML-based+malware+detection
30. AI Post Transformers: Petri: Accelerating AI Safety Auditing — Hal Turing & Dr. Ada Shannon, Fri,
https://podcasters.spotify.com/pod/show/12146088098/episodes/Petri-Accelerating-AI-Safety-Auditing-e39boei
31. AI Post Transformers: Bloom: an open source tool for automated behavioral evaluations — Hal Turing & Dr. Ada Shannon, Tue,
https://podcasters.spotify.com/pod/show/12146088098/episodes/Bloom-an-open-source-tool-for-automated-behavioral-evaluations-e3fi1ge
32. AI Post Transformers: Agentic Reasoning for Large Language Models: A Comprehensive Roadmap — Hal Turing & Dr. Ada Shannon, Sat,
https://podcasters.spotify.com/pod/show/12146088098/episodes/Agentic-Reasoning-for-Large-Language-Models-A-Comprehensive-Roadmap-e3e43ru
Interactive Visualization: MalGEN: Multi-Agent AI for Red Teaming Malware

This episode explores NVIDIA's Nemotron 3 white paper, which introduces a three-tier model family (Nano, Super, Ultra) built on a hybrid architecture combining Mamba-2 structured state space layers with Mixture-of-Experts routing, targeting simultaneous state-of-the-art accuracy, one-million-token context, and substantially higher inference throughput than comparable dense Transformer MoE models. The discussion traces how Mamba-2's fixed-size recurrent state eliminates the KV cache's linear memory growth — the central bottleneck for long-context and agentic workloads — and explains NVIDIA's novel LatentMoE extension, which projects tokens into a reduced latent dimension before expert routing to cut communication costs while activating more experts per token. Multi-Token Prediction from Meta FAIR appears as a training accelerant, predicting multiple future tokens simultaneously to improve both training efficiency and generation speed, while NVFP4, NVIDIA's 4-bit floating-point training format used for Super and Ultra, raises open questions about numerical stability during post-training. Listeners interested in how recent theoretical work on state space duality translates into a production-scale model family, or in the architectural tradeoffs enabling practical multi-agent pipelines at scale, will find the episode a concrete and technically grounded case study.

Sources:
1. NVIDIA Nemotron 3: Efficient and Open Intelligence — NVIDIA, :, Aaron Blakeman, Aaron Grattafiori, Aarti Basant, Abhibha Gupta, Abhinav Khattar, Adi Renduchintala, Aditya Vavre, Akanksha Shukla, Akhiad Bercovich, Aleksander Ficek, Aleksandr Shaposhnikov, Alex Kondratenko, Alexander Bukharin, Alexandre Milesi, Ali Taghibakhshi, Alisa Liu, Amelia Barton, Ameya Sunil Mahabaleshwarkar, Amir Klein, Amit Zuker, Amnon Geifman, Amy Shen, Anahita Bhiwandiwalla, Andrew Tao, Anjulie Agrusa, Ankur Verma, Ann Guan, Anubhav Mandarwal, Arham Mehta, Ashwath Aithal, Ashwin Poojary, Asif Ahamed, Asit Mishra, Asma Kuriparambil Thekkumpate, Ayush Dattagupta, Banghua Zhu, Bardiya Sadeghi, Barnaby Simkin, Ben Lanir, Benedikt Schifferer, Besmira Nushi, Bilal Kartal, Bita Darvish Rouhani, Boris Ginsburg, Brandon Norick, Brandon Soubasis, Branislav Kisacanin, Brian Yu, Bryan Catanzaro, Carlo del Mundo, Chantal Hwang, Charles Wang, Cheng-Ping Hsieh, Chenghao Zhang, Chenhan Yu, Chetan Mungekar, Chintan Patel, Chris Alexiuk, Christopher Parisien, Collin Neale, Cyril Meurillon, Damon Mosk-Aoyama, Dan Su, Dane Corneil, Daniel Afrimi, Daniel Lo, Daniel Rohrer, Daniel Serebrenik, Daria Gitman, Daria Levy, Darko Stosic, David Mosallanezhad, Deepak Narayanan, Dhruv Nathawani, Dima Rekesh, Dina Yared, Divyanshu Kakwani, Dong Ahn, Duncan Riach, Dusan Stosic, Edgar Minasyan, Edward Lin, Eileen Long, Eileen Peters Long, Elad Segal, Elena Lantz, Ellie Evans, Elliott Ning, Eric Chung, Eric Harper, Eric Tramel, Erick Galinkin, Erik Pounds, Evan Briones, Evelina Bakhturina, Evgeny Tsykunov, Faisal Ladhak, Fay Wang, Fei Jia, Felipe Soares, Feng Chen, Ferenc Galko, Frank Sun, Frankie Siino, Gal Hubara Agam, Ganesh Ajjanagadde, Gantavya Bhatt, Gargi Prasad, George Armstrong, Gerald Shen, Gorkem Batmaz, Grigor Nalbandyan, Haifeng Qian, Harsh Sharma, Hayley Ross, Helen Ngo, Herbert Hum, Herman Sahota, Hexin Wang, Himanshu Soni, Hiren Upadhyay, Huizi Mao, Huy C Nguyen, Huy Q Nguyen, Iain Cunningham, Ido Galil, Ido Shahaf, Igor Gitman, Ilya Loshchilov, Itamar Schen, Itay Levy, Ivan Moshkov, Izik Golan, Izzy Putterman, Jan Kautz, Jane Polak Scowcroft, Jared Casper, Jatin Mitra, Jeffrey Glick, Jenny Chen, Jesse Oliver, Jian Zhang, Jiaqi Zeng, Jie Lou, Jimmy Zhang, Jinhang Choi, Jining Huang, Joey Conway, Joey Guman, John Kamalu, Johnny Greco, Jonathan Cohen, Joseph Jennings, Joyjit Daw, Julien Veron Vialard, Junkeun Yi, Jupinder Parmar, Kai Xu, Kan Zhu, Kari Briski, Katherine Cheung, Katherine Luna, Keith Wyss, Keshav Santhanam, Kevin Shih, Kezhi Kong, Khushi Bhardwaj, Kirthi Shankar, Krishna C. Puvvada, Krzysztof Pawelec, Kumar Anik, Lawrence McAfee, Laya Sleiman, Leon Derczynski, Li Ding, Lizzie Wei, Lucas Liebenwein, Luis Vega, Maanu Grover, Maarten Van Segbroeck, Maer Rodrigues de Melo, Mahdi Nazemi, Makesh Narsimhan Sreedhar, Manoj Kilaru, Maor Ashkenazi, Marc Romeijn, Marcin Chochowski, Mark Cai, Markus Kliegl, Maryam Moosaei, Matt Kulka, Matvei Novikov, Mehrzad Samadi, Melissa Corpuz, Mengru Wang, Meredith Price, Michael Andersch, Michael Boone, Michael Evans, Miguel Martinez, Mikail Khona, Mike Chrzanowski, Minseok Lee, Mohammad Dabbah, Mohammad Shoeybi, Mostofa Patwary, Nabin Mulepati, Najeeb Nabwani, Natalie Hereth, Nave Assaf, Negar Habibi, Neta Zmora, Netanel Haber, Nicola Sessions, Nidhi Bhatia, Nikhil Jukar, Nikki Pope, Nikolai Ludwig, Nima Tajbakhsh, Nir Ailon, Nirmal Juluru, Nishant Sharma, Oleksii Hrinchuk, Oleksii Kuchaiev, Olivier Delalleau, Oluwatobi Olabiyi, Omer Ullman Argov, Omri Puny, Oren Tropp, Ouye Xie, Parth Chadha, Pasha Shamis, Paul Gibbons, Pavlo Molchanov, Pawel Morkisz, Peter Dykas, Peter Jin, Pinky Xu, Piotr Januszewski, Pranav Prashant Thombre, Prasoon Varshney, Pritam Gundecha, Przemek Tredak, Qing Miao, Qiyu Wan, Rabeeh Karimi Mahabadi, Rachit Garg, Ran El-Yaniv, Ran Zilberstein, Rasoul Shafipour, Rich Harang, Rick Izzo, Rima Shahbazyan, Rishabh Garg, Ritika Borkar, Ritu Gala, Riyad Islam, Robert Hesse, Roger Waleffe, Rohit Watve, Roi Koren, Ruoxi Zhang, Russell Hewett, Russell J. Hewett, Ryan Prenger, Ryan Timbrook, Sadegh Mahdavi, Sahil Modi, Samuel Kriman, Sangkug Lim, Sanjay Kariyappa, Sanjeev Satheesh, Saori Kaji, Satish Pasumarthi, Saurav Muralidharan, Sean Narentharen, Sean Narenthiran, Seonmyeong Bak, Sergey Kashirsky, Seth Poulos, Shahar Mor, Shanmugam Ramasamy, Shantanu Acharya, Shaona Ghosh, Sharath Turuvekere Sreenivas, Shelby Thomas, Shiqing Fan, Shreya Gopal, Shrimai Prabhumoye, Shubham Pachori, Shubham Toshniwal, Shuoyang Ding, Siddharth Singh, Simeng Sun, Smita Ithape, Somshubra Majumdar, Soumye Singhal, Stas Sergienko, Stefania Alborghetti, Stephen Ge, Sugam Dipak Devare, Sumeet Kumar Barua, Suseella Panguluri, Suyog Gupta, Sweta Priyadarshi, Syeda Nahida Akter, Tan Bui, Teodor-Dumitru Ene, Terry Kong, Thanh Do, Tijmen Blankevoort, Tim Moon, Tom Balough, Tomer Asida, Tomer Bar Natan, Tomer Ronen, Tugrul Konuk, Twinkle Vashishth, Udi Karpas, Ushnish De, Vahid Noorozi, Vahid Noroozi, Venkat Srinivasan, Venmugil Elango, Victor Cui, Vijay Korthikanti, Vinay Rao, Vitaly Kurin, Vitaly Lavrukhin, Vladimir Anisimov, Wanli Jiang, Wasi Uddin Ahmad, Wei Du, Wei Ping, Wenfei Zhou, Will Jennings, William Zhang, Wojciech Prazuch, Xiaowei Ren, Yashaswi Karnati, Yejin Choi, Yev Meyer, Yi-Fu Wu, Yian Zhang, Yigong Qin, Ying Lin, Yonatan Geifman, Yonggan Fu, Yoshi Subara, Yoshi Suhara, Yubo Gao, Zach Moshe, Zhen Dong, Zhongbo Zhu, Zihan Liu, Zijia Chen, Zijie Yan, 2025
http://arxiv.org/abs/2512.20856
2. Scaling LLM Test-Time Compute Optimally Can be More Effective than Scaling Model Parameters — Charlie Snell, Jaehoon Lee, Kelvin Xu, Aviral Kumar (Google DeepMind), 2024
https://scholar.google.com/scholar?q=Scaling+LLM+Test-Time+Compute+Optimally+Can+be+More+Effective+than+Scaling+Model+Parameters
3. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning — DeepSeek-AI (Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, and many others), 2025
https://scholar.google.com/scholar?q=DeepSeek-R1%3A+Incentivizing+Reasoning+Capability+in+LLMs+via+Reinforcement+Learning
4. Thinking Tokens for Language Modeling — David Herel, Tomas Mikolov (Czech Institute of Informatics / FAIR), 2023
https://scholar.google.com/scholar?q=Thinking+Tokens+for+Language+Modeling
5. Training Large Language Models to Reason in a Continuous Latent Space (Coconut) — Shibo Hao, Sainbayar Sukhbaatar, Jason Weston, Yuandong Tian, Zhiting Hu (Meta FAIR), 2024
https://scholar.google.com/scholar?q=Training+Large+Language+Models+to+Reason+in+a+Continuous+Latent+Space+%28Coconut%29
6. Don't Think Too Much: Reasoning Budget Control for Large Reasoning Models — Xuan He, Zhenyu He, Qingyu Meng, Yonghao Zhong, Zhonglin Shi, Fandong Meng, Jie Zhou (Tencent AI Lab / WeChat AI), 2025
https://scholar.google.com/scholar?q=Don%27t+Think+Too+Much%3A+Reasoning+Budget+Control+for+Large+Reasoning+Models
7. Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality — Dao, T. and Gu, A., 2024
https://scholar.google.com/scholar?q=Transformers+are+SSMs%3A+Generalized+Models+and+Efficient+Algorithms+Through+Structured+State+Space+Duality
8. Jamba: A Hybrid Transformer-Mamba Language Model — Lieber et al., AI21 Labs, 2024
https://scholar.google.com/scholar?q=Jamba%3A+A+Hybrid+Transformer-Mamba+Language+Model
9. Better & Faster Large Language Models via Multi-Token Prediction — Gloeckle et al., 2024
https://scholar.google.com/scholar?q=Better+%26+Faster+Large+Language+Models+via+Multi-Token+Prediction
10. RULER: What's the Real Context Size of Your Long-Context Language Models? — Hsieh et al., 2024
https://scholar.google.com/scholar?q=RULER%3A+What%27s+the+Real+Context+Size+of+Your+Long-Context+Language+Models%3F
11. DeepSeek-V3 Technical Report — DeepSeek-AI, 2025
https://scholar.google.com/scholar?q=DeepSeek-V3+Technical+Report
12. Mixtral of Experts — Jiang et al., 2024
https://scholar.google.com/scholar?q=Mixtral+of+Experts
13. In-Context Learning with Long-Context Models: An In-Depth Exploration — Bertsch et al., 2024
https://scholar.google.com/scholar?q=In-Context+Learning+with+Long-Context+Models%3A+An+In-Depth+Exploration
14. Attn-QAT: 4-Bit Attention With Quantization-Aware Training — approximate, 2024–2025, 2024-2025
https://scholar.google.com/scholar?q=Attn-QAT%3A+4-Bit+Attention+With+Quantization-Aware+Training
15. FP4 All the Way: Fully Quantized Training of LLMs — approximate, 2024–2025, 2024-2025
https://scholar.google.com/scholar?q=FP4+All+the+Way%3A+Fully+Quantized+Training+of+LLMs
16. Optimizing Large Language Model Training Using FP4 Quantization — approximate, 2024–2025, 2024-2025
https://scholar.google.com/scholar?q=Optimizing+Large+Language+Model+Training+Using+FP4+Quantization
17. Speculative Decoding and Beyond: An In-Depth Survey of Techniques — approximate, 2024–2025, 2024-2025
https://scholar.google.com/scholar?q=Speculative+Decoding+and+Beyond%3A+An+In-Depth+Survey+of+Techniques
18. Latent Prototype Routing: Achieving Near-Perfect Load Balancing in Mixture-of-Experts — approximate, 2024–2025, 2024-2025
https://scholar.google.com/scholar?q=Latent+Prototype+Routing%3A+Achieving+Near-Perfect+Load+Balancing+in+Mixture-of-Experts
19. AI Post Transformers: Nemotron Elastic: Towards Efficient Many-in-One Reasoning LLMs — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-04-2-papers-combined-217168.mp3
20. AI Post Transformers: Switch Transformers: Trillion Parameter Models with Sparsity — Hal Turing & Dr. Ada Shannon, Wed,
https://podcasters.spotify.com/pod/show/12146088098/episodes/Switch-Transformers-Trillion-Parameter-Models-with-Sparsity-e373fd3
21. AI Post Transformers: GLM-5: Transitioning from Vibe Coding to Agentic Engineering — Hal Turing & Dr. Ada Shannon, Fri,
https://podcasters.spotify.com/pod/show/12146088098/episodes/GLM-5-Transitioning-from-Vibe-Coding-to-Agentic-Engineering-e3fbfls
22. AI Post Transformers: Kimi Linear: Efficient Expressive Attention Architecture — Hal Turing & Dr. Ada Shannon, Sun,
https://podcasters.spotify.com/pod/show/12146088098/episodes/Kimi-Linear-Efficient-Expressive-Attention-Architecture-e3aclec
23. AI Post Transformers: TailorKV: Hybrid KV Cache Compression for LLMs — Hal Turing & Dr. Ada Shannon, Wed,
https://podcasters.spotify.com/pod/show/12146088098/episodes/TailorKV-Hybrid-KV-Cache-Compression-for-LLMs-e38bmv2
24. AI Post Transformers: AWQ: On-Device LLM Compression and Acceleration — Hal Turing & Dr. Ada Shannon, Mon,
https://podcasters.spotify.com/pod/show/12146088098/episodes/AWQ-On-Device-LLM-Compression-and-Acceleration-e388njr
25. AI Post Transformers: LFM2-8B-A1B: Efficient On-Device Mixture-of-Experts — Hal Turing & Dr. Ada Shannon, Sun,
https://podcasters.spotify.com/pod/show/12146088098/episodes/LFM2-8B-A1B-Efficient-On-Device-Mixture-of-Experts-e3a2oh2
Interactive Visualization: NVIDIA Nemotron 3 Hybrid SSM Transformer Architecture

Hal Turing and Dr. Ada Shannon open the episode by confronting a structural flaw that has been hiding in plain sight since the transformer era began: tokenization bias. The episode centers on "Exact Byte-Level Probabilities from Tokenized Language Models for FIM-Tasks and Model Ensembles" by Buu Phan, Brandon Amos, Itai Gat, Marton Havasi, Matthew Muckley, and Karen Ullrich (ICLR 2025), which formally proves that a tokenized model and a byte-level model can be statistically equivalent and still produce wildly different predictions for the same next character. The hosts trace the origins of the problem through BPE's introduction by Rico Sennrich, Barry Haddow, and Alexandra Birch in 2016 and its industrialization via Kudo and Richardson's SentencePiece in 2018 — a library now frozen into the spine of LLaMA, Mistral, Gemma, and most open-source models not in the OpenAI lineage. The discussion sharpens around fill-in-the-middle prompting, the paradigm introduced by Mohammad Bavarian and colleagues at OpenAI in 2022 and now embedded in every major code completion tool from GitHub Copilot to StarCoder. Shannon walks through the paper's central example: a code completion scenario where the correct next character receives a probability of exactly zero — not a rounding artifact but a structural impossibility, because the tokenizer has carved up the prompt in a way that makes the right answer unreachable in token-space. Turing challenges the framing, arguing that byte-level alternatives like ByT5 and MegaByte existed and BPE was an informed trade-off against the three-to-eight times sequence length penalty that raw bytes impose on attention compute. Shannon holds the line: the point is not that BPE was a mistake but that its systematic bias was never formally characterized until now, and the Byte-Token Representation Lemma finally gives the field the mathematical language to name and measure it. The episode closes by introducing the second paper from the episode's pairing — Minixhofer et al.'s NeurIPS 2025 work on cross-tokenizer knowledge distillation — which attacks the tokenizer barrier from the training side rather than the inference side. Where Phan et al. offer a zero-shot correction algorithm that recovers 18% on fill-in-the-middle coding benchmarks without any retraining, Minixhofer et al. enable knowledge transfer between models with fundamentally incompatible vocabularies, breaking the assumption that distillation requires shared tokenization. Together the two papers sketch a trajectory where tokenization becomes a transparent implementation detail rather than an architectural constraint that determines what a model can and cannot express.

Sources:
1. Exact Byte-Level Probabilities from Tokenized Language Models for FIM-Tasks and Model Ensembles — Buu Phan, Brandon Amos, Itai Gat, Marton Havasi, Matthew Muckley, Karen Ullrich, 2024
http://arxiv.org/abs/2410.09303
2. https://arxiv.org/pdf/2503.20083
3. Neural Machine Translation of Rare Words with Subword Units — Rico Sennrich, Barry Haddow, Alexandra Birch, 2016
https://scholar.google.com/scholar?q=Neural+Machine+Translation+of+Rare+Words+with+Subword+Units
4. SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing — Taku Kudo, John Richardson, 2018
https://scholar.google.com/scholar?q=SentencePiece%3A+A+simple+and+language+independent+subword+tokenizer+and+detokenizer+for+Neural+Text+Processing
5. Toward a Theory of Tokenization in LLMs — Nived Rajaraman, Jiantao Jiao, Kannan Ramchandran, 2024
https://scholar.google.com/scholar?q=Toward+a+Theory+of+Tokenization+in+LLMs
6. How Good is Your Tokenizer? On the Monolingual Performance of Multilingual Language Models — Phillip Rust, Jonas Pfeiffer, Ivan Vulić, Sebastian Ruder, Iryna Gurevych, 2021
https://scholar.google.com/scholar?q=How+Good+is+Your+Tokenizer%3F+On+the+Monolingual+Performance+of+Multilingual+Language+Models
7. ByT5: Towards a Token-Free Future with Pre-trained Byte-to-Byte Models — Linting Xue, Aditya Barua, Noah Constant, Rami Al-Rfou, Sharan Narang, Mihir Kale, Adam Roberts, Colin Raffel, 2022
https://scholar.google.com/scholar?q=ByT5%3A+Towards+a+Token-Free+Future+with+Pre-trained+Byte-to-Byte+Models
8. MEGABYTE: Predicting Million-byte Sequences with Multiscale Transformers — Lili Yu, Daniel Simig, Colin Flaherty, Armen Aghajanyan, Luke Zettlemoyer, Mike Lewis, 2023
https://scholar.google.com/scholar?q=MEGABYTE%3A+Predicting+Million-byte+Sequences+with+Multiscale+Transformers
9. CANINE: Pre-training an Efficient Tokenization-Free Encoder for Language Representation — Jonathan H. Clark, Dan Garrette, Iulia Turc, John Wieting, 2022
https://scholar.google.com/scholar?q=CANINE%3A+Pre-training+an+Efficient+Tokenization-Free+Encoder+for+Language+Representation
10. Charformer: Fast Character Transformers via Gradient-Based Subword Tokenization — Yi Tay, Vinh Q. Tran, Sebastian Ruder, Jai Gupta, Hyung Won Chung, Dara Bahri, Zhen Qin, Simon Kornblith, Cecelia Zhang, Donald Metzler, Mostafa Dehghani, 2022
https://scholar.google.com/scholar?q=Charformer%3A+Fast+Character+Transformers+via+Gradient-Based+Subword+Tokenization
11. Efficient Training of Language Models to Fill in the Middle — Mohammad Bavarian, Heewoo Jun, Nikolas Tezak, John Schulman, Christine McLeavey, Jerry Tworek, Mark Chen, 2022
https://scholar.google.com/scholar?q=Efficient+Training+of+Language+Models+to+Fill+in+the+Middle
12. Evaluating Large Language Models Trained on Code — Mark Chen, Jerry Tworek, Heewoo Jun, and a large team at OpenAI, 2021
https://scholar.google.com/scholar?q=Evaluating+Large+Language+Models+Trained+on+Code
13. Code Llama: Open Foundation Models for Code — Baptiste Rozière, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, and others at Meta AI, 2023
https://scholar.google.com/scholar?q=Code+Llama%3A+Open+Foundation+Models+for+Code
14. StarCoder: May the Source Be With You! — Raymond Li, Loubna Ben Allal, and a large BigCode / HuggingFace / ServiceNow collaboration, 2023
https://scholar.google.com/scholar?q=StarCoder%3A+May+the+Source+Be+With+You%21
15. Training Products of Experts by Minimizing Contrastive Divergence — Geoffrey E. Hinton, 2002
https://scholar.google.com/scholar?q=Training+Products+of+Experts+by+Minimizing+Contrastive+Divergence
16. Self-Consistency Improves Chain of Thought Reasoning in Language Models — Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, Denny Zhou, 2023
https://scholar.google.com/scholar?q=Self-Consistency+Improves+Chain+of+Thought+Reasoning+in+Language+Models
17. Knowledge Fusion of Large Language Models — Fanqi Wan, Xinting Huang, Deng Cai, Xiaojun Quan, Wei Bi, Shuming Shi, 2024
https://scholar.google.com/scholar?q=Knowledge+Fusion+of+Large+Language+Models
18. LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion — Dongfu Jiang, Xiang Ren, Bill Yuchen Lin, 2023
https://scholar.google.com/scholar?q=LLM-Blender%3A+Ensembling+Large+Language+Models+with+Pairwise+Ranking+and+Generative+Fusion
19. Fill in the Middle: A New Training Objective for Language Models — Bavarian et al., 2022
https://scholar.google.com/scholar?q=Fill+in+the+Middle%3A+A+New+Training+Objective+for+Language+Models
20. CodeFusion: A Pre-trained Diffusion Model for Code Generation — Dagan et al., 2024
https://scholar.google.com/scholar?q=CodeFusion%3A+A+Pre-trained+Diffusion+Model+for+Code+Generation
21. Is Tokenization More Than Compression? — Rajaraman et al., 2024
https://scholar.google.com/scholar?q=Is+Tokenization+More+Than+Compression%3F
22. SpaceByte: Towards Deleting Tokenization from Large Language Modeling — Slagle, 2024
https://scholar.google.com/scholar?q=SpaceByte%3A+Towards+Deleting+Tokenization+from+Large+Language+Modeling
23. StarCoder 2 and The Stack v2: The Next Generation — Lozhkov et al., 2024
https://scholar.google.com/scholar?q=StarCoder+2+and+The+Stack+v2%3A+The+Next+Generation
24. Byte Latent Transformer: Patches Scale Better Than Tokens — Pagnoni et al. (Meta AI), 2024
https://scholar.google.com/scholar?q=Byte+Latent+Transformer%3A+Patches+Scale+Better+Than+Tokens
25. Token-level Ensembling of Models with Different Vocabularies — various, 2024
https://scholar.google.com/scholar?q=Token-level+Ensembling+of+Models+with+Different+Vocabularies
26. Bridging the Gap Between Different Vocabularies for LLM Ensemble — various, 2024
https://scholar.google.com/scholar?q=Bridging+the+Gap+Between+Different+Vocabularies+for+LLM+Ensemble
27. Accelerating LLM Inference with Lossless Speculative Decoding Algorithms for Heterogeneous Vocabularies — various, 2024
https://scholar.google.com/scholar?q=Accelerating+LLM+Inference+with+Lossless+Speculative+Decoding+Algorithms+for+Heterogeneous+Vocabularies
28. Bridging Developer Instructions and Code Completion Through Instruction-Aware Fill-in-the-Middle Paradigm — various, 2024
https://scholar.google.com/scholar?q=Bridging+Developer+Instructions+and+Code+Completion+Through+Instruction-Aware+Fill-in-the-Middle+Paradigm
29. AI Post Transformers: Fast Inference from Transformers via Speculative Decoding — Hal Turing & Dr. Ada Shannon, Sat,
https://podcasters.spotify.com/pod/show/12146088098/episodes/Fast-Inference-from-Transformers-via-Speculative-Decoding-e3foclv
30. AI Post Transformers: Multiagent Debate Improves Language Model Reasoning — Hal Turing & Dr. Ada Shannon, Fri,
https://podcasters.spotify.com/pod/show/12146088098/episodes/Multiagent-Debate-Improves-Language-Model-Reasoning-e36kfd4
Interactive Visualization: Tokenization Bias: The Hidden Flaw Breaking Language Models

Hal Turing and Dr. Ada Shannon examine "A Systematic Characterization of LLM Inference on GPUs" by Haonan Wang and colleagues from the Institute of Computing Technology at the Chinese Academy of Sciences, Zhejiang Lab, China Telecom, and Peking University. Dr. Shannon opens by framing the paper's central contribution: the inference literature is fragmented across power consumption studies, kernel profiling, quantization analysis, and serving schedulers, none sharing a common analytical vocabulary. Wang et al. attempt to bridge ML systems thinking with computer architecture rigor into a unified cross-layer framework — from "the model feels slow" down to the specific hardware bottleneck and its root cause. The hosts build the foundational model carefully, tracing execution from prompt to generated token. Dr. Shannon explains the two-phase structure of transformer inference: the prefill phase processes all prompt tokens simultaneously in a matrix-matrix multiply that is compute-bound, writing key-value pairs into the KV cache; the decode phase generates tokens autoregressively, and at each step the attention mechanism reads back every prior cached key-value pair alongside the full weight matrices. Hal presses on the memory bandwidth implications — by token three hundred and ninety-nine, the model loads nearly four hundred KV pairs per layer per step just to produce one new token. Dr. Shannon maps this onto the Roofline model, originally developed by Williams, Waterman, and Patterson at UC Berkeley in 2009, using arithmetic intensity — the ratio of floating-point operations to memory traffic — to show why decode sits deep in the memory-bound regime while prefill remains compute-bound. That asymmetry is the load-bearing structure of the paper's argument. When Hal pushes back on whether the compute-bound prefill and memory-bound decode framing is merely a formalization of established practitioner intuition, Dr. Shannon draws a sharp distinction between the heuristic and the science. Knowing decode is memory-bound is the starting point; the paper's contribution is stall analysis — profiling with hardware performance counters to determine precisely why GPU thread warps pause during execution. A warp stalled on memory pipe saturation has a different cause and a different remedy than one stalled on instruction-level data dependency. The hosts frame this microarchitectural root cause analysis as one of four dimensions in the paper's framework, alongside the two-phase prefill-decode heterogeneity, system scaling principles, and the boundaries where current GPU architectures hit fundamental limits — together forming the diagnostic vocabulary the field has been missing.

Sources:
1. A Systematic Characterization of LLM Inference on GPUs — Haonan Wang, Xuxin Xiao, Mingyu Yan, Zhuoyuan Zhu, Dengke Han, Duo Wang, Wenming Li, Xiaochun Ye, Cunchen Hu, Hongyang Chen, Guangyu Sun, 2025
http://arxiv.org/abs/2512.01644
2. Efficient Memory Management for Large Language Model Serving with PagedAttention — Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, Ion Stoica, 2023
https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention
3. Orca: A Distributed Serving System for Transformer-Based Generative Models — Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim, Byung-Gon Chun, 2022
https://scholar.google.com/scholar?q=Orca%3A+A+Distributed+Serving+System+for+Transformer-Based+Generative+Models
4. AlpaServe: Statistical Multiplexing with Model Parallelism for Deep Learning Serving — Zhuohan Li, Lianmin Zheng, Yinmin Zhong, Vincent Liu, Ying Sheng, Xing Jin, Yanping Huang, Zhifeng Chen, Hao Zhang, Joseph E. Gonzalez, Ion Stoica, 2022
https://scholar.google.com/scholar?q=AlpaServe%3A+Statistical+Multiplexing+with+Model+Parallelism+for+Deep+Learning+Serving
5. LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale — Tim Dettmers, Mike Lewis, Younes Belkada, Luke Zettlemoyer, 2022
https://scholar.google.com/scholar?q=LLM.int8%28%29%3A+8-bit+Matrix+Multiplication+for+Transformers+at+Scale
6. Splitwise: Efficient Generative LLM Inference Using Phase Splitting — Pratyush Patel, Esha Choukse, Chaojie Zhang, Aashaka Shah, Íñigo Goiri, Saeed Maleki, Ricardo Bianchini, 2024
https://scholar.google.com/scholar?q=Splitwise%3A+Efficient+Generative+LLM+Inference+Using+Phase+Splitting
7. DistServe: Disaggregating Prefill and Decoding for Goodput-Optimized Large Language Model Serving — Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, Hao Zhang, 2024
https://scholar.google.com/scholar?q=DistServe%3A+Disaggregating+Prefill+and+Decoding+for+Goodput-Optimized+Large+Language+Model+Serving
8. Sarathi-Serve: Efficient LLM Inference by Piggybacking Decodes on Chunked Prefills — Amey Agrawal, Nitin Kedia, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav Gulavani, Alexey Tumanov, Ramachandran Ramjee, 2024
https://scholar.google.com/scholar?q=Sarathi-Serve%3A+Efficient+LLM+Inference+by+Piggybacking+Decodes+on+Chunked+Prefills
9. FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU — Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, Beidi Chen, Percy Liang, Christopher Ré, Ion Stoica, Ce Zhang, 2023
https://scholar.google.com/scholar?q=FlexGen%3A+High-Throughput+Generative+Inference+of+Large+Language+Models+with+a+Single+GPU
10. Roofline: An Insightful Visual Performance Model for Multicore Architectures — Samuel Williams, Andrew Waterman, David Patterson, 2009
https://scholar.google.com/scholar?q=Roofline%3A+An+Insightful+Visual+Performance+Model+for+Multicore+Architectures
11. In-Datacenter Performance Analysis of a Tensor Processing Unit — Norman P. Jouppi, Cliff Young, Nishant Patil, David Patterson, et al., 2017
https://scholar.google.com/scholar?q=In-Datacenter+Performance+Analysis+of+a+Tensor+Processing+Unit
12. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness — Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, Christopher Ré, 2022
https://scholar.google.com/scholar?q=FlashAttention%3A+Fast+and+Memory-Efficient+Exact+Attention+with+IO-Awareness
13. Dissecting the Graphcore IPU Architecture via Microbenchmarking — Amey Agrawal, Rohit Ramaprasad, Kranthi Gade, Ramachandran Ramjee, 2021
https://scholar.google.com/scholar?q=Dissecting+the+Graphcore+IPU+Architecture+via+Microbenchmarking
14. Sarathi-Serve: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills — Agrawal et al., 2024
https://scholar.google.com/scholar?q=Sarathi-Serve%3A+Efficient+LLM+Inference+by+Piggybacking+Decodes+with+Chunked+Prefills
15. FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning — Dao, 2023
https://scholar.google.com/scholar?q=FlashAttention-2%3A+Faster+Attention+with+Better+Parallelism+and+Work+Partitioning
16. Flash-Decoding for Long-Context Inference — Dao et al., 2023
https://scholar.google.com/scholar?q=Flash-Decoding+for+Long-Context+Inference
17. Roofline: An Insightful Visual Performance Model for Floating-Point Programs and Multicore Architectures — Williams, Waterman, Patterson, 2009
https://scholar.google.com/scholar?q=Roofline%3A+An+Insightful+Visual+Performance+Model+for+Floating-Point+Programs+and+Multicore+Architectures
18. Sarathi: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills — Agrawal et al., 2023
https://scholar.google.com/scholar?q=Sarathi%3A+Efficient+LLM+Inference+by+Piggybacking+Decodes+with+Chunked+Prefills
19. POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM Inference — approximate, recent, 2024-2025
https://scholar.google.com/scholar?q=POD-Attention%3A+Unlocking+Full+Prefill-Decode+Overlap+for+Faster+LLM+Inference
20. KVLink: Accelerating Large Language Models via Efficient KV Cache Reuse — approximate, recent, 2024-2025
https://scholar.google.com/scholar?q=KVLink%3A+Accelerating+Large+Language+Models+via+Efficient+KV+Cache+Reuse
21. Shared Disk KV Cache Management for Efficient Multi-Instance Inference in RAG-Powered LLMs — approximate, recent, 2024-2025
https://scholar.google.com/scholar?q=Shared+Disk+KV+Cache+Management+for+Efficient+Multi-Instance+Inference+in+RAG-Powered+LLMs
22. ExpertFlow: Optimized Expert Activation and Token Allocation for Efficient Mixture-of-Experts Inference — approximate, recent, 2024-2025
https://scholar.google.com/scholar?q=ExpertFlow%3A+Optimized+Expert+Activation+and+Token+Allocation+for+Efficient+Mixture-of-Experts+Inference
23. Harder Task Needs More Experts: Dynamic Routing in MoE Models — approximate, recent, 2024-2025
https://scholar.google.com/scholar?q=Harder+Task+Needs+More+Experts%3A+Dynamic+Routing+in+MoE+Models
24. DECA: A Near-Core LLM Decompression Accelerator Grounded on a 3D Roofline Model — approximate, recent, 2024-2025
https://scholar.google.com/scholar?q=DECA%3A+A+Near-Core+LLM+Decompression+Accelerator+Grounded+on+a+3D+Roofline+Model
25. HeadInfer: Memory-Efficient LLM Inference by Head-Wise Offloading — approximate, recent, 2024-2025
https://scholar.google.com/scholar?q=HeadInfer%3A+Memory-Efficient+LLM+Inference+by+Head-Wise+Offloading
26. AI Post Transformers: AI and the Memory Wall: Overcoming Bottlenecks — Hal Turing & Dr. Ada Shannon, Fri,
https://podcasters.spotify.com/pod/show/12146088098/episodes/AI-and-the-Memory-Wall-Overcoming-Bottlenecks-e36ki0u
27. AI Post Transformers: Flash-LLM: Efficient LLM Inference with Unstructured Sparsity on Tensor Cores — Hal Turing & Dr. Ada Shannon, Fri,
https://podcasters.spotify.com/pod/show/12146088098/episodes/Flash-LLM-Efficient-LLM-Inference-with-Unstructured-Sparsity-on-Tensor-Cores-e3aalrl
28. AI Post Transformers: Advancements in Efficient KV Cache Quantization and Management — Hal Turing & Dr. Ada Shannon, Thu,
https://podcasters.spotify.com/pod/show/12146088098/episodes/Advancements-in-Efficient-KV-Cache-Quantization-and-Management-e3fk9kr
29. AI Post Transformers: Adaptive LLM Partitioning for Edge Inference — Hal Turing & Dr. Ada Shannon, Tue,
https://podcasters.spotify.com/pod/show/12146088098/episodes/Adaptive-LLM-Partitioning-for-Edge-Inference-e38a9mn
30. AI Post Transformers: Parallel Token Prediction: From ProphetNet to Dependent Multi-Token Generation — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-04-4-papers-combined-8a237d.mp3
Interactive Visualization: Systematic Characterization of LLM Inference on GPUs

Hal Turing and Dr. Ada Shannon return to the CARTRIDGE compression system with a mechanistic lens, covering Maurizio A. Diaz's paper "Learned Structure in Cartridges: Keys as Shareable Routers in Self-Studied Representations" (arXiv 2508.17032), presented at the NeurIPS 2025 Workshop on Mechanistic Interpretability. Building on the original CARTRIDGE episode from November 10th, 2025 and the follow-up from February 6th, 2026, this episode asks the question those earlier discussions left open: what structure does the optimizer actually induce in a trained CARTRIDGE? The hosts ground the discussion in the memory scaling problem driving the entire field—KV caches that grow linearly with context length, now routinely dwarfing model weights at the 128K-to-million-token scales of current frontier models—and trace how techniques like PagedAttention, Grouped Query Attention, and token eviction address symptoms without shrinking the underlying representation. Diaz's central finding is a clean functional division between key and value vectors inside a trained CARTRIDGE. Keys converge to stable retrieval routers: low-rank, consistent structures that steer attention toward the right stored content across diverse queries. Values carry the compressed semantic payload. The hosts connect this directly to how CARTRIDGE's Self-Study training pipeline works—because the cache is optimized against synthetic question-answer traces generated by the model over its own content, the training signal explicitly selects for routing behavior, making the key-as-router outcome a predictable consequence of the objective rather than an accident. Diaz uses Singular Value Decomposition to quantify this structure layer by layer, separating the geometric properties of key matrices from value matrices across training checkpoints. Two downstream findings from the key-router property shape the second half of the discussion. Because keys are stable and low-rank, they transfer across tasks with minimal degradation—a result with direct implications for multi-task serving, where a single shared key structure could route to task-specific value sets without independent CARTRIDGE training per deployment. The Sampled Chunk Initialization method introduced in the paper exploits this stability to warm-start CARTRIDGE training, accelerating convergence by initializing the learnable KV pairs from a small representative sample rather than random weights. Hal and Ada close by discussing what the key-as-router framing implies for KV-cache compression research more broadly: if the routing function is separable and transferable, compression schemes that conflate keys and values may be discarding structure that has real serving-efficiency value.

Sources:
1. Learned Structure in Cartridges: Keys as Shareable Routers in Self-Studied Representations — Maurizio Diaz, 2025
http://arxiv.org/abs/2508.17032
2. CARTRIDGES: Learning to Pack Long Contexts into KV Caches — Zhihao Zhang, Aditya Desai, Amir Gholami, Michael W. Mahoney, Kurt Keutzer, et al. (Berkeley / ICSI), 2025
https://scholar.google.com/scholar?q=CARTRIDGES%3A+Learning+to+Pack+Long+Contexts+into+KV+Caches
3. H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models — Zhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen, Lianghao Huang, Mu Li, Beidi Chen, Jason D. Lee, Binhang Yuan, Ce Zhang, Cho-Jui Hsieh, 2023
https://scholar.google.com/scholar?q=H2O%3A+Heavy-Hitter+Oracle+for+Efficient+Generative+Inference+of+Large+Language+Models
4. Scissorhands: Exploiting the Persistence of Importance Hypothesis for LLM KV Cache Compression at Test Time — Zichang Liu, Aditya Desai, Fangshuo Liao, Weitao Wang, Victor Xie, Zhaozhuo Xu, Anastasios Kyrillidis, Anshumali Shrivastava, 2023
https://scholar.google.com/scholar?q=Scissorhands%3A+Exploiting+the+Persistence+of+Importance+Hypothesis+for+LLM+KV+Cache+Compression+at+Test+Time
5. Efficient Streaming Language Models with Attention Sinks — Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, Mike Lewis, 2023
https://scholar.google.com/scholar?q=Efficient+Streaming+Language+Models+with+Attention+Sinks
6. Efficient Memory Management for Large Language Model Serving with PagedAttention — Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, Ion Stoica, 2023
https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention
7. Lost in the Middle: How Language Models Use Long Contexts — Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, Percy Liang, 2023
https://scholar.google.com/scholar?q=Lost+in+the+Middle%3A+How+Language+Models+Use+Long+Contexts
8. GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints — Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebron, Sumit Sanghai, 2023
https://scholar.google.com/scholar?q=GQA%3A+Training+Generalized+Multi-Query+Transformer+Models+from+Multi-Head+Checkpoints
9. Extending Context Window of Large Language Models via Positional Interpolation — Shouyuan Chen, Sherman Wong, Liangjian Chen, Yuandong Tian, 2023
https://scholar.google.com/scholar?q=Extending+Context+Window+of+Large+Language+Models+via+Positional+Interpolation
10. Constitutional AI: Harmlessness from AI Feedback — Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al. (Anthropic), 2022
https://scholar.google.com/scholar?q=Constitutional+AI%3A+Harmlessness+from+AI+Feedback
11. Distilling the Knowledge in a Neural Network — Geoffrey Hinton, Oriol Vinyals, Jeff Dean, 2015
https://scholar.google.com/scholar?q=Distilling+the+Knowledge+in+a+Neural+Network
12. Compressing Context to Enhance Inference Efficiency of Large Language Models — Yucheng Li, Bo Dong, Chenghua Lin, Frank Guerin, 2023
https://scholar.google.com/scholar?q=Compressing+Context+to+Enhance+Inference+Efficiency+of+Large+Language+Models
13. AutoCompressors: Adapting Language Models to Summarize Arbitrary Contexts into Summary Vectors — Alexis Chevalier, Alexander Wettig, Anirudh Anand, Danqi Chen, 2023
https://scholar.google.com/scholar?q=AutoCompressors%3A+Adapting+Language+Models+to+Summarize+Arbitrary+Contexts+into+Summary+Vectors
14. A Mathematical Framework for Transformer Circuits — Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, Chris Olah, 2021
https://scholar.google.com/scholar?q=A+Mathematical+Framework+for+Transformer+Circuits
15. Toy Models of Superposition — Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield-Dodds, Ben Mann, Shan Carter, Chris Olah, Tom Brown, Jack Clark, Sam McCandlish, Chris Olah, Jared Kaplan, Dario Amodei, 2022
https://scholar.google.com/scholar?q=Toy+Models+of+Superposition
16. Towards Monosemanticity: Decomposing Language Models with Dictionary Learning — Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Tom Conerly, Nick Turner, Cem Anil, Carson Denison, Amanda Askell, Robert Lasenby, William Lamont, Annabel Shearer, Zac Hatfield-Dodds, Tom Henighan, Nicholas Joseph, Ben Ziegler, Bilal Chaudhry, Will Hao, Timothy Telleen-Lawton, Daniel Mossing, Adam Jermyn, Tom Brown, Chris Olah, Sam McCandlish, Jack Clark, Dario Amodei, 2023
https://scholar.google.com/scholar?q=Towards+Monosemanticity%3A+Decomposing+Language+Models+with+Dictionary+Learning
17. In-context Learning and Induction Heads — Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, Chris Olah, 2022
https://scholar.google.com/scholar?q=In-context+Learning+and+Induction+Heads
18. AutoCompressors: Compressing Contexts with Language Models — Chevalier et al., 2023
https://scholar.google.com/scholar?q=AutoCompressors%3A+Compressing+Contexts+with+Language+Models
19. Gist Tokens: Compressing Prompts into Tokens for Long-Context Language Models — Mu et al., 2023
https://scholar.google.com/scholar?q=Gist+Tokens%3A+Compressing+Prompts+into+Tokens+for+Long-Context+Language+Models
20. The Power of Scale for Parameter-Efficient Prompt Tuning — Lester et al., 2021
https://scholar.google.com/scholar?q=The+Power+of+Scale+for+Parameter-Efficient+Prompt+Tuning
21. Prefix-Tuning: Optimizing Continuous Prompts for Generation — Li and Liang, 2021
https://scholar.google.com/scholar?q=Prefix-Tuning%3A+Optimizing+Continuous+Prompts+for+Generation
22. PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling — Cai et al., 2024
https://scholar.google.com/scholar?q=PyramidKV%3A+Dynamic+KV+Cache+Compression+based+on+Pyramidal+Information+Funneling
23. Tokasaurus: A High-Throughput LLM Serving Engine with Grouped-Sparse Attention — Lenz et al., 2025
https://scholar.google.com/scholar?q=Tokasaurus%3A+A+High-Throughput+LLM+Serving+Engine+with+Grouped-Sparse+Attention
24. Unlocking the Address Book: Dissecting the Sparse Semantic Structure of LLM Key-Value Caches via Sparse Autoencoders — unknown from snippet, 2024-2025
https://scholar.google.com/scholar?q=Unlocking+the+Address+Book%3A+Dissecting+the+Sparse+Semantic+Structure+of+LLM+Key-Value+Caches+via+Sparse+Autoencoders
25. SCBench: A KV Cache-Centric Analysis of Long-Context Methods — approximate, multiple authors, 2024-2025
https://scholar.google.com/scholar?q=SCBench%3A+A+KV+Cache-Centric+Analysis+of+Long-Context+Methods
26. LLMLingua-2: Data Distillation for Efficient and Faithful Task-Agnostic Prompt Compression — approximate, Microsoft Research, 2024
https://scholar.google.com/scholar?q=LLMLingua-2%3A+Data+Distillation+for+Efficient+and+Faithful+Task-Agnostic+Prompt+Compression
27. AI Post Transformers: CARTRIDGE: Efficient In-Context Learning via Distillation — Hal Turing & Dr. Ada Shannon, Mon,
https://podcasters.spotify.com/pod/show/12146088098/episodes/CARTRIDGE-Efficient-In-Context-Learning-via-Distillation-e3aous4
28. AI Post Transformers: Context Distillation for Language Models — Hal Turing & Dr. Ada Shannon, Mon,
https://podcasters.spotify.com/pod/show/12146088098/episodes/Context-Distillation-for-Language-Models-e3aouen
29. AI Post Transformers: Advancements in Efficient KV Cache Quantization and Management — Hal Turing & Dr. Ada Shannon, Thu,
https://podcasters.spotify.com/pod/show/12146088098/episodes/Advancements-in-Efficient-KV-Cache-Quantization-and-Management-e3fk9kr
30. AI Post Transformers: Architectural Migration to Multi-head Latent Attention — Hal Turing & Dr. Ada Shannon, Wed,
https://podcasters.spotify.com/pod/show/12146088098/episodes/Architectural-Migration-to-Multi-head-Latent-Attention-e39jbmq

Hal Turing and Dr. Ada Shannon dig into "DualPath: Breaking the Storage Bandwidth Bottleneck in Agentic LLM Inference," a February 2026 paper from a thirteen-author team spanning Peking University, Tsinghua University, and DeepSeek-AI. The episode opens with a striking observation from production: on disaggregated inference clusters running agentic workloads, prefill machines saturate their storage network interfaces at 100% utilization while the equivalent hardware on decode machines sits nearly idle. The hosts use this asymmetry as a lens into a counterintuitive reality — H100 GPUs throttled to 40% compute utilization not by arithmetic limits, but by a storage NIC. Ada explains the structural reason agentic workloads are uniquely hostile to existing infrastructure: the short-append pattern. Unlike standard multi-turn chat, agentic sessions accumulate dozens to hundreds of turns where each round appends only a small number of tokens — a tool result, a stack trace, a code output — onto a context that may already span tens of thousands of tokens. Because that prior context never changes, its KV-Cache was computed once and stored. DeepSeek's production traces show KV-Cache hit rates of 95% or higher, meaning the dominant cost shifts from GPU computation to storage I/O: loading gigabytes of persistent key-value state from external NVMe-backed distributed storage, layer by layer, into prefill engines via RDMA. Hal presses on that 95% figure specifically, establishing that it is grounded in real production traffic rather than idealized assumptions — a distinction that determines whether storage bandwidth or GPU compute is the correct optimization target. The episode frames DualPath's core insight against this background: the storage NICs on decode engines represent idle bandwidth that could absorb KV-Cache load traffic currently overwhelming prefill-side storage interfaces. By routing that traffic through decode-side hardware and transferring it to prefill engines over RDMA, DualPath breaks the single-path bottleneck without adding new hardware. The hosts connect this to the broader memory wall argument — that as context lengths grow and agentic sessions deepen, the architectural shift toward disaggregated inference is not optional, and the constraints driving system design are increasingly about data movement rather than floating-point throughput. DualPath's reported throughput improvement of up to 1.96x is presented as evidence that exploiting idle hardware asymmetries, rather than scaling compute, is where near-term agentic inference gains will be found.

Sources:
1. DualPath: Breaking the Storage Bandwidth Bottleneck in Agentic LLM Inference — Yongtong Wu, Shaoyuan Chen, Yinmin Zhong, Rilin Huang, Yixuan Tan, Wentao Zhang, Liyue Zhang, Shangyan Zhou, Yuxuan Liu, Shunfeng Zhou, Mingxing Zhang, Xin Jin, Panpan Huang, 2026
http://arxiv.org/abs/2602.21548
2. Efficient Memory Management for Large Language Model Serving with PagedAttention — Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, Ion Stoica, 2023
https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention
3. SGLang: Efficient Execution of Structured Language Model Programs — Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Jeff Huang, Chuyue Sun, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph Gonzalez, Clark Barrett, Ying Sheng, 2024
https://scholar.google.com/scholar?q=SGLang%3A+Efficient+Execution+of+Structured+Language+Model+Programs
4. Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving — Ruoyu Qin, Zheming Li, Weiran He, Mingxing Zhang, Yongwei Wu, Weimin Zheng, Xinran Xu, 2024
https://scholar.google.com/scholar?q=Mooncake%3A+A+KVCache-centric+Disaggregated+Architecture+for+LLM+Serving
5. CacheGen: Fast Context Loading for Language Model Applications — Yuhan Liu, Hanchen Li, Yihua Cheng, Siddhant Ray, Yuyang Huang, Qizheng Zhang, Kuntai Du, Jiayi Yao, Shan Lu, Ganesh Ananthanarayanan, Michael Maire, Henry Hoffmann, Ari Holtzman, Junchen Jiang, 2024
https://scholar.google.com/scholar?q=CacheGen%3A+Fast+Context+Loading+for+Language+Model+Applications
6. DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving — Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, Hao Zhang, 2024
https://scholar.google.com/scholar?q=DistServe%3A+Disaggregating+Prefill+and+Decoding+for+Goodput-optimized+Large+Language+Model+Serving
7. Splitwise: Efficient Generative LLM Inference Using Phase Splitting — Pratyush Patel, Esha Choukse, Chaojie Zhang, Aashaka Shah, Igo Goiri, Saeed Maleki, Ricardo Bianchini, 2024
https://scholar.google.com/scholar?q=Splitwise%3A+Efficient+Generative+LLM+Inference+Using+Phase+Splitting
8. Sarathi-Serve: Chunked Prefills for Efficient LLM Inference — Amey Agrawal, Nitin Kedia, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav Gulavani, Alexey Tumanov, Ramachandran Ramjee, 2024
https://scholar.google.com/scholar?q=Sarathi-Serve%3A+Chunked+Prefills+for+Efficient+LLM+Inference
9. Llumnix: Dynamic Scheduling for Large Language Model Serving — Biao Sun, Ziming Huang, Hanyu Zhao, Wencong Xiao, Xinyi Zhang, Yong Li, Wei Lin, 2024
https://scholar.google.com/scholar?q=Llumnix%3A+Dynamic+Scheduling+for+Large+Language+Model+Serving
10. FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU — Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, Beidi Chen, Percy Liang, Christopher Re, Ion Stoica, Ce Zhang, 2023
https://scholar.google.com/scholar?q=FlexGen%3A+High-Throughput+Generative+Inference+of+Large+Language+Models+with+a+Single+GPU
11. InfiniGen: Efficient Generative Inference of Large Language Models with Dynamic KV Cache Management — Wonbeom Lee, Jungi Lee, Junghwan Seo, Jaewoong Sim, 2024
https://scholar.google.com/scholar?q=InfiniGen%3A+Efficient+Generative+Inference+of+Large+Language+Models+with+Dynamic+KV+Cache+Management
12. Parrot: Efficient Servicing of LLM-based Applications with Semantic Variable — Chaofan Lin, Zhenhua Han, Chengruidong Zhang, Yuqing Yang, Fan Yang, Chen Chen, Lili Qiu, 2024
https://scholar.google.com/scholar?q=Parrot%3A+Efficient+Servicing+of+LLM-based+Applications+with+Semantic+Variable
13. AgentBench: Evaluating LLMs as Agents — Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, Shudan Zhang, Xiang Deng, Aohan Zeng, Zhiyuan Liu, Yuxiao Dong, Jie Tang, 2023
https://scholar.google.com/scholar?q=AgentBench%3A+Evaluating+LLMs+as+Agents
14. MemGPT: Towards LLMs as Operating Systems — Charles Packer, Sarah Wooders, Kevin Lin, Vivian Fang, Shishir G. Patil, Ion Stoica, Joseph Gonzalez, 2023
https://scholar.google.com/scholar?q=MemGPT%3A+Towards+LLMs+as+Operating+Systems
15. Sarathi-Serve: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills — Amey Agrawal, Nikhil Mohan, Sudhanshu Panwar, et al., 2024
https://scholar.google.com/scholar?q=Sarathi-Serve%3A+Efficient+LLM+Inference+by+Piggybacking+Decodes+with+Chunked+Prefills
16. CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowledge Fusion — Jiayi Yao, Hanchen Li, Yuhan Liu, et al., 2024
https://scholar.google.com/scholar?q=CacheBlend%3A+Fast+Large+Language+Model+Serving+for+RAG+with+Cached+Knowledge+Fusion
17. CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving — Liu et al., 2024
https://scholar.google.com/scholar?q=CacheGen%3A+KV+Cache+Compression+and+Streaming+for+Fast+Large+Language+Model+Serving
18. Compute or Load KV Cache? Why Not Both? — Unknown (snippet only), 2024-2025
https://scholar.google.com/scholar?q=Compute+or+Load+KV+Cache%3F+Why+Not+Both%3F
19. Semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage — Unknown (snippet only), 2024-2025
https://scholar.google.com/scholar?q=Semi-PD%3A+Towards+Efficient+LLM+Serving+via+Phase-Wise+Disaggregated+Computation+and+Unified+Storage
20. Prefill-Decode Aggregation or Disaggregation? Unifying Both for Goodput-Optimized LLM Serving (TaiChi) — Unknown (snippet only), 2024-2025
https://scholar.google.com/scholar?q=Prefill-Decode+Aggregation+or+Disaggregation%3F+Unifying+Both+for+Goodput-Optimized+LLM+Serving+%28TaiChi%29
21. Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live — Unknown (snippet only), 2024-2025
https://scholar.google.com/scholar?q=Continuum%3A+Efficient+and+Robust+Multi-Turn+LLM+Agent+Scheduling+with+KV+Cache+Time-to-Live
22. SideQuest: Model-Driven KV Cache Management for Long-Horizon Agentic Reasoning — Unknown (snippet only), 2024-2025
https://scholar.google.com/scholar?q=SideQuest%3A+Model-Driven+KV+Cache+Management+for+Long-Horizon+Agentic+Reasoning
23. RDMA Point-to-Point Communication for LLM Systems — Unknown (snippet only), 2024-2025
https://scholar.google.com/scholar?q=RDMA+Point-to-Point+Communication+for+LLM+Systems
24. AI Post Transformers: FAST26: Bidaw: Enhancing Key-Value Caching for Interactive LLM Serving via Bidirectional — Hal Turing & Dr. Ada Shannon, Wed,
https://podcasters.spotify.com/pod/show/12146088098/episodes/FAST26-Bidaw-Enhancing-Key-Value-Caching-for-Interactive-LLM-Serving-via-Bidirectional-e3fjgkh
25. AI Post Transformers: SYMPHONY: Memory Management for LLM Multi-Turn Inference — Hal Turing & Dr. Ada Shannon, Mon,
https://podcasters.spotify.com/pod/show/12146088098/episodes/SYMPHONY-Memory-Management-for-LLM-Multi-Turn-Inference-e3ap0pf
26. AI Post Transformers: CXL-SpecKV: Bridging the LLM Memory Wall with Speculative FPGA Disaggregation — Hal Turing & Dr. Ada Shannon, Sat,
https://podcasters.spotify.com/pod/show/12146088098/episodes/CXL-SpecKV-Bridging-the-LLM-Memory-Wall-with-Speculative-FPGA-Disaggregation-e3foad0
27. AI Post Transformers: AI and the Memory Wall: Overcoming Bottlenecks — Hal Turing & Dr. Ada Shannon, Fri,
https://podcasters.spotify.com/pod/show/12146088098/episodes/AI-and-the-Memory-Wall-Overcoming-Bottlenecks-e36ki0u
28. AI Post Transformers: Tempo: SLO-Aware LLM Serving Maximizing Service Gain — Hal Turing & Dr. Ada Shannon, Mon,
https://podcasters.spotify.com/pod/show/12146088098/episodes/Tempo-SLO-Aware-LLM-Serving-Maximizing-Service-Gain-e3ap12a
Interactive Visualization: DualPath Breaks Storage Bandwidth Bottleneck in Agentic Inference

Hal Turing and Dr. Ada Shannon open by situating the Dao-Gu paper — "Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality" (arXiv 2405.21060, ICML 2024) — within the bifurcated landscape of sequence modeling. For years, Transformer researchers and SSM researchers developed in parallel, unable to borrow optimizations across the divide. The episode traces the lineage from HiPPO and S4 through Mamba's selective state spaces, explaining why SSMs' linear-time theoretical advantage never translated to wall-clock wins: the GPU ecosystem was built for dense matrix multiply, and SSMs lacked the tooling that FlashAttention brought to attention. The hosts also credit the direct intellectual ancestor — Katharopoulos et al.'s 2020 "Transformers are RNNs" — which showed that softmax attention with a kernel approximation reduces to a linear recurrence, establishing the conceptual template Dao and Gu would later formalize. The core of the episode is a careful unpacking of structured semiseparable matrices, a class of objects from numerical linear algebra — Kalman filter theory, PDE solvers — entirely unknown to the ML community until Dao and Gu made the connection. Every entry of a causal SSM input-output matrix has the form C[i] times a chain of transition matrices A times B[j], which is precisely the generator representation of a rank-d semiseparable matrix. Shannon walks through the O(n) factored form — M[i,j] equals u[i] times v[j] below the diagonal — and explains how this structure encodes both the SSM recurrence and the masked attention computation as two views of the same algebraic object. The canonical reference is Vandebril, Van Barel, and Mastronardi's two-volume work from Johns Hopkins, 2008, a body of theory the ML community had never encountered. Once the connection is made, hardware-efficient algorithms from one domain port directly to the other. But the episode frames this mathematical achievement against a harder question raised by subsequent theoretical work: the L²M Condition of Chen et al. and the bipartite mutual information scaling law. The duality shows that SSM computation and attention computation are equivalent representations — but equivalence of computation does not imply equivalence of information retention. SSMs compress sequence history into a fixed-size state regardless of context length; the Transformer's KV-cache grows linearly, retaining more as context expands. The mutual information scaling law formalizes this gap: capturing the multi-token dependencies present in natural language requires a history state that grows with context length. The episode closes on what this implies for hybrid architectures — systems that combine SSM efficiency with selective attention — and whether the theoretical unification Dao and Gu achieved changes how practitioners should think about where compressed state fails.

Sources:
1. Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality — Tri Dao, Albert Gu, 2024
http://arxiv.org/abs/2405.21060
2. https://arxiv.org/pdf/2503.04725
3. On a New Class of Structured Matrices — Yuli Eidelman, Israel Gohberg, 1997
https://scholar.google.com/scholar?q=On+a+New+Class+of+Structured+Matrices
4. Time-Varying Systems and Computations — Patrick Dewilde, Alle-Jan van der Veen, 1998
https://scholar.google.com/scholar?q=Time-Varying+Systems+and+Computations
5. Matrix Computations with Semiseparable Matrices (Volumes 1 & 2) — Raf Vandebril, Marc Van Barel, Nicola Mastronardi, 2008
https://scholar.google.com/scholar?q=Matrix+Computations+with+Semiseparable+Matrices+%28Volumes+1+%26+2%29
6. Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality — Tri Dao, Albert Gu, 2024
https://scholar.google.com/scholar?q=Transformers+are+SSMs%3A+Generalized+Models+and+Efficient+Algorithms+Through+Structured+State+Space+Duality
7. Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention — Katharopoulos, Vyas, Pappas, Fleuret, 2020
https://scholar.google.com/scholar?q=Transformers+are+RNNs%3A+Fast+Autoregressive+Transformers+with+Linear+Attention
8. Mamba: Linear-Time Sequence Modeling with Selective State Spaces — Gu, Dao, 2023
https://scholar.google.com/scholar?q=Mamba%3A+Linear-Time+Sequence+Modeling+with+Selective+State+Spaces
9. Semiseparable Matrices in Linear Algebra, Science and Engineering (2 vols.) — Vandebril, Van Barel, Mastronardi, 2008
https://scholar.google.com/scholar?q=Semiseparable+Matrices+in+Linear+Algebra%2C+Science+and+Engineering+%282+vols.%29
10. Retentive Network: A Successor to Transformer for Large Language Models — Sun, Li, Dong, Huang, Peng, Liu, Wang, Lin, Yuan, Chen, Wei, 2023
https://scholar.google.com/scholar?q=Retentive+Network%3A+A+Successor+to+Transformer+for+Large+Language+Models
11. RWKV: Reinventing RNNs for the Transformer Era — Peng et al., 2023
https://scholar.google.com/scholar?q=RWKV%3A+Reinventing+RNNs+for+the+Transformer+Era
12. Gated Linear Attention Transformers with Hardware-Efficient Training — Yang, Wang, Shen, Peng, Dao, Gu, Matoba, 2023
https://scholar.google.com/scholar?q=Gated+Linear+Attention+Transformers+with+Hardware-Efficient+Training
13. FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning — Dao, 2024
https://scholar.google.com/scholar?q=FlashAttention-2%3A+Faster+Attention+with+Better+Parallelism+and+Work+Partitioning
14. Zoology: Measuring and Improving Recall in Efficient Language Models — Arora, Eyuboglu, Timalsina, Johnson, Poli, Zou, Rudra, Ré, 2024
https://scholar.google.com/scholar?q=Zoology%3A+Measuring+and+Improving+Recall+in+Efficient+Language+Models
15. Overcoming Long-Context Limitations of State-Space Models via Context-Dependent Sparse Attention — unknown (snippet only), 2024-2025
https://scholar.google.com/scholar?q=Overcoming+Long-Context+Limitations+of+State-Space+Models+via+Context-Dependent+Sparse+Attention
16. Bayesian Optimality of In-Context Learning with Selective State Spaces — unknown (snippet only), 2024-2025
https://scholar.google.com/scholar?q=Bayesian+Optimality+of+In-Context+Learning+with+Selective+State+Spaces
17. On the Expressiveness of Softmax Attention: A Recurrent Neural Network Perspective — unknown (snippet only), 2024-2025
https://scholar.google.com/scholar?q=On+the+Expressiveness+of+Softmax+Attention%3A+A+Recurrent+Neural+Network+Perspective
18. Bridging the Divide: Reconsidering Softmax and Linear Attention — unknown (snippet only), 2024-2025
https://scholar.google.com/scholar?q=Bridging+the+Divide%3A+Reconsidering+Softmax+and+Linear+Attention
19. Softmax Linear Attention: Reclaiming Global Competition — unknown (snippet only), 2024-2025
https://scholar.google.com/scholar?q=Softmax+Linear+Attention%3A+Reclaiming+Global+Competition
20. Hybrid Architectures for Language Models: Systematic Analysis and Design Insights — unknown (snippet only), 2024-2025
https://scholar.google.com/scholar?q=Hybrid+Architectures+for+Language+Models%3A+Systematic+Analysis+and+Design+Insights
21. KVLink: Accelerating Large Language Models via Efficient KV Cache Reuse — unknown (snippet only), 2024-2025
https://scholar.google.com/scholar?q=KVLink%3A+Accelerating+Large+Language+Models+via+Efficient+KV+Cache+Reuse
22. AI Post Transformers: FlashAttention-4 Conquers Asymmetric GPU Hardware Scaling — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-06-flashattention-4-conquers-asymmetric-gpu-78839b.mp3
23. AI Post Transformers: ATTENTION2D and lean attention: Distributed Self-Attention — Hal Turing & Dr. Ada Shannon, 2025
https://podcasters.spotify.com/pod/show/12146088098/episodes/ATTENTION2D-and-lean-attention-Distributed-Self-Attention-e3a7r4n

Special announcement: AI Post Transformers is now open source under the MIT license. New home at podcast.do-not-panic.com, interactive paper visualizations, community paper submissions, deep dive into the editorial queue algorithm, and internationalization roadmap.

Hal Turing and Dr. Ada Shannon dig into FlashAttention-4, a March 2026 paper from a cross-institutional team including Tri Dao, Jay Shah, and colleagues at Princeton, Meta, NVIDIA, Colfax Research, Georgia Tech, and Together AI. The paper targets a precise hardware mismatch on NVIDIA's Blackwell B200: tensor core throughput doubles compared to the H100, but shared memory bandwidth and dedicated exponential function units do not scale at the same rate. Rather than waiting for hardware fixes, the authors co-design the attention algorithm with the asymmetric architecture itself — making FlashAttention-4 the first attention kernel built specifically for Blackwell's scaling profile. To frame why this matters, Shannon traces the full lineage of FlashAttention research. The original 2022 NeurIPS paper by Dao and colleagues reframed attention as an IO problem: instead of materializing the quadratic N×N score matrix in slow off-chip High Bandwidth Memory, tiling and online softmax keep computation inside the fast on-chip shared memory of each streaming multiprocessor. FlashAttention-2 doubled throughput through sequence-dimension parallelism. FlashAttention-3 pushed H100 utilization to roughly 75% by exploiting Hopper-specific warp specialization and asynchronous data movement. Each generation addressed a qualitatively different bottleneck — and Blackwell introduced a new one that none of those solutions anticipated. The hosts ground the stakes for practitioners who work in ML without writing GPU kernels. Attention sits at the core of every Transformer-based system — large language models, vision transformers, multimodal architectures — and long-context workloads at 32K to 128K tokens make the quadratic memory cost and HBM round-trips increasingly punishing. Shannon introduces the roofline model as the analytic lens the paper uses to characterize where Blackwell kernels actually bottleneck, setting up how FlashAttention-4's algorithmic co-design approach navigates the compute and memory bandwidth ceilings that previous generations of the kernel never had to contend with.

Sources:
1. FlashAttention-4: Algorithm and Kernel Pipelining Co-Design for Asymmetric Hardware Scaling — Ted Zadouri, Markus Hoehnerbach, Jay Shah, Timmy Liu, Vijay Thakkar, Tri Dao, 2026
http://arxiv.org/abs/2603.05451v1
2. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness — Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, Christopher Ré, 2022
https://scholar.google.com/scholar?q=FlashAttention%3A+Fast+and+Memory-Efficient+Exact+Attention+with+IO-Awareness
3. FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning — Tri Dao, 2023
https://scholar.google.com/scholar?q=FlashAttention-2%3A+Faster+Attention+with+Better+Parallelism+and+Work+Partitioning
4. FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low Precision — Jay Shah, Ganesh Bikshandi, Ying Zhang, Vijay Thakkar, Pradeep Ramani, Tri Dao, 2024
https://scholar.google.com/scholar?q=FlashAttention-3%3A+Fast+and+Accurate+Attention+with+Asynchrony+and+Low+Precision
5. Triton: An Intermediate Language and Compiler for Tiled Neural Network Computations — Philippe Tillet, H. T. Kung, David Cox, 2019
https://scholar.google.com/scholar?q=Triton%3A+An+Intermediate+Language+and+Compiler+for+Tiled+Neural+Network+Computations
6. Roofline: An Insightful Visual Performance Model for Floating-Point Programs and Multicore Architectures — Samuel Williams, Andrew Waterman, David Patterson, 2009
https://scholar.google.com/scholar?q=Roofline%3A+An+Insightful+Visual+Performance+Model+for+Floating-Point+Programs+and+Multicore+Architectures
7. Efficiently Scaling Transformer Inference — Reiner Pope, Sholto Douglas, Aakanksha Chowdhery, Jacob Devlin, James Bradbury, Jonathan Heek, Kefan Xiao, Sharan Narang, Jeff Dean, 2023
https://scholar.google.com/scholar?q=Efficiently+Scaling+Transformer+Inference
8. In-Datacenter Performance Analysis of a Tensor Processing Unit — Norman P. Jouppi et al. (Google), 2017
https://scholar.google.com/scholar?q=In-Datacenter+Performance+Analysis+of+a+Tensor+Processing+Unit
9. Mamba: Linear-Time Sequence Modeling with Selective State Spaces — Albert Gu, Tri Dao, 2023
https://scholar.google.com/scholar?q=Mamba%3A+Linear-Time+Sequence+Modeling+with+Selective+State+Spaces
10. Efficient Memory Management for Large Language Model Serving with PagedAttention — Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, Ion Stoica, 2023
https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention
11. Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism — Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, Bryan Catanzaro, 2019
https://scholar.google.com/scholar?q=Megatron-LM%3A+Training+Multi-Billion+Parameter+Language+Models+Using+Model+Parallelism
12. FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision — Jay Shah, Ganesh Bikshandi, Ying Zhang, Vijay Thakkar, Pradeep Ramani, Tri Dao, 2024
https://scholar.google.com/scholar?q=FlashAttention-3%3A+Fast+and+Accurate+Attention+with+Asynchrony+and+Low-precision
13. SageAttention2: Efficient Attention with Thorough Outlier Smoothing and Per-thread INT4 Quantization — Jintao Zhang et al., 2024
https://scholar.google.com/scholar?q=SageAttention2%3A+Efficient+Attention+with+Thorough+Outlier+Smoothing+and+Per-thread+INT4+Quantization
14. Ring Attention with Blockwise Transformers for Near-Infinite Context — Hao Liu, Matei Zaharia, Pieter Abbeel, 2023
https://scholar.google.com/scholar?q=Ring+Attention+with+Blockwise+Transformers+for+Near-Infinite+Context
15. FlexAttention: The Flexibility of PyTorch with the Performance of FlashAttention — PyTorch Team, 2024
https://scholar.google.com/scholar?q=FlexAttention%3A+The+Flexibility+of+PyTorch+with+the+Performance+of+FlashAttention
16. DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model — DeepSeek-AI, 2024
https://scholar.google.com/scholar?q=DeepSeek-V2%3A+A+Strong%2C+Economical%2C+and+Efficient+Mixture-of-Experts+Language+Model
17. Softmax output approximation for activation memory-efficient training of attention-based networks — approximate — likely 2023–2024, 2023–2024
https://scholar.google.com/scholar?q=Softmax+output+approximation+for+activation+memory-efficient+training+of+attention-based+networks
18. FlashAttention-T: Towards Fully Tensorized Attention by Exploiting Tensor-Vector Parallelism — approximate — likely 2024–2025, 2024–2025
https://scholar.google.com/scholar?q=FlashAttention-T%3A+Towards+Fully+Tensorized+Attention+by+Exploiting+Tensor-Vector+Parallelism
19. Overcoming Long-Context Limitations of State-Space Models via Context-Dependent Sparse Attention — approximate — likely 2024–2025, 2024–2025
https://scholar.google.com/scholar?q=Overcoming+Long-Context+Limitations+of+State-Space+Models+via+Context-Dependent+Sparse+Attention
20. Based on Tensor Core Sparse Kernels Accelerating Deep Neural Networks — approximate — likely 2024–2025, 2024–2025
https://scholar.google.com/scholar?q=Based+on+Tensor+Core+Sparse+Kernels+Accelerating+Deep+Neural+Networks
21. AI Post Transformers: FlashAttention-2: Faster Attention with Better Parallelism — Hal Turing & Dr. Ada Shannon, Fri,
https://podcasters.spotify.com/pod/show/12146088098/episodes/FlashAttention-2-Faster-Attention-with-Better-Parallelism-e36kdm0
22. AI Post Transformers: ATTENTION2D and lean attention: Distributed Self-Attention — Hal Turing & Dr. Ada Shannon, Wed,
https://podcasters.spotify.com/pod/show/12146088098/episodes/ATTENTION2D-and-lean-attention-Distributed-Self-Attention-e3a7r4n
23. AI Post Transformers: Jet-RL: Stable On-Policy Reinforcement Learning with Unified FP8 Flow — Hal Turing & Dr. Ada Shannon, Tue,
https://podcasters.spotify.com/pod/show/12146088098/episodes/Jet-RL-Stable-On-Policy-Reinforcement-Learning-with-Unified-FP8-Flow-e3f7det
24. AI Post Transformers: Mojo: Performance-Portable HPC Kernels on GPUs — Hal Turing & Dr. Ada Shannon, Sat,
https://podcasters.spotify.com/pod/show/12146088098/episodes/Mojo-Performance-Portable-HPC-Kernels-on-GPUs-e39n4sk
Interactive Visualization: FlashAttention-4: Algorithm & Kernel Co-Design

This episode examines the fundamental latency bottleneck in autoregressive language models: sequential token generation requires one full transformer forward pass per output token, leaving GPU parallelism idle during single-user inference. The episode centers on Draxler et al. (5 co-authors, UC Irvine and Chan-Zuckerberg Initiative), whose paper on Parallel Token Prediction landed Christmas Eve 2025 and argues that the independence assumption baked into all prior multi-token schemes is not an acceptable approximation but the actual limiting factor. The paper asks whether multiple tokens can be jointly predicted in a single pass — modeling dependencies among them — without sacrificing the expressiveness that makes autoregressive generation reliable. First author Felix Draxler previously led the Free-form Flows work in 2024, and the normalizing flow machinery he developed there is central to how the paper solves the dependency problem. The episode traces the historical arc carefully. Qi et al. (7 co-authors, Microsoft Research Asia) published ProphetNet in January 2020 — predating GPT-3 by four months — in the encoder-decoder world of seq2seq tasks. Their critique was precise: standard one-step-ahead teacher forcing gives models no incentive to plan ahead, letting local bigram correlations dominate at the expense of long-range coherence. Their answer was n-gram prediction, training the decoder to simultaneously predict tokens at t+1, t+2, and t+3 using parallel heads that did not condition on each other. The independence assumption was already present. When Brown et al. (OpenAI, May 2020) demonstrated that scale and in-context conditioning make the encoder optional, the field shifted to decoder-only architectures — but ProphetNet's core insight migrated cleanly. Gloeckle et al. (FAIR, Meta, April 2024) rebuilt multi-token prediction for decoder-only models using independent output heads, DeepSeek adopted the same approach, and NVIDIA incorporated it into Nemotron 3. The independence assumption migrated with the insight, and Draxler et al. argue that limitation has been compounding ever since. The episode situates Parallel Token Prediction against the two main camps attacking inference latency. Speculative decoding — covered across twenty prior episodes — keeps the model's output distribution unchanged by using a small draft model whose proposals a large verifier checks in one batched pass; the latency gain comes entirely from accepted tokens per step. Multi-token prediction is the other camp: train the model itself to emit several tokens at once, collapsing multiple forward passes into one, at the cost of changed model behavior during training. Draxler et al.'s contribution is showing that jointly predicting dependent tokens, using normalizing flows to capture the conditional structure across the prediction horizon, preserves the modeling power that independent-head approaches discard. The episode works through both the architectural mechanics and the theoretical argument, making the case that Parallel Token Prediction resolves the tension that has run from ProphetNet through every independent-head scheme in between.

Interactive Visualization: PTP: Resolving the Independence Flaw
Sources:
1. Parallel Token Prediction for Language Models — Felix Draxler, Justus Will, Farrin Marouf Sofian, Theofanis Karaletsos, Sameer Singh, Stephan Mandt, 2025
http://arxiv.org/abs/2512.21323
2. https://arxiv.org/pdf/2404.19737v1
3. https://arxiv.org/pdf/2412.19437
4. https://arxiv.org/pdf/2512.20856
5. Better & Faster Large Language Models via Multi-Token Prediction — Fabian Gloeckle, Badr Youbi Idrissi, Baptiste Rozière, David Lopez-Paz, Gabriel Synnaeve (Meta), 2024
https://scholar.google.com/scholar?q=Better+%26+Faster+Large+Language+Models+via+Multi-Token+Prediction
6. Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads — Tianle Cai, Yuhong Li, Zhengmian Hu, et al., 2024
https://scholar.google.com/scholar?q=Medusa%3A+Simple+LLM+Inference+Acceleration+Framework+with+Multiple+Decoding+Heads
7. Lookahead Decoding: Break the Sequential Dependency of LLM Inference — Yichao Fu, Peter Bailis, Ion Stoica, Hao Zhang, 2024
https://scholar.google.com/scholar?q=Lookahead+Decoding%3A+Break+the+Sequential+Dependency+of+LLM+Inference
8. EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty — Yuhui Li, Fangyun Wei, Chao Zhang, Hongyang Zhang, 2024
https://scholar.google.com/scholar?q=EAGLE%3A+Speculative+Sampling+Requires+Rethinking+Feature+Uncertainty
9. Fast Inference from Transformers via Speculative Decoding — Yaniv Leviathan, Matan Kalman, Yossi Matias (Google), 2023
https://scholar.google.com/scholar?q=Fast+Inference+from+Transformers+via+Speculative+Decoding
10. Accelerating Large Language Model Decoding with Speculative Sampling — Charlie Chen, Sebastian Borgeaud, Geoffrey Irving, Jean-Baptiste Lespiau, Laurent Sifre, John Jumper (DeepMind), 2023
https://scholar.google.com/scholar?q=Accelerating+Large+Language+Model+Decoding+with+Speculative+Sampling
11. Spec-Bench: A Benchmark for Evaluating Speculative Decoding Approaches — Heming Xia, Zhe Yang, Qingxiu Dong, Peiyi Wang, Yongqi Li, Tao Ge, Tianyu Liu, Wenjie Li, Zhifang Sui, 2024
https://scholar.google.com/scholar?q=Spec-Bench%3A+A+Benchmark+for+Evaluating+Speculative+Decoding+Approaches
12. EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees — Yuhui Li, Fangyun Wei, Chao Zhang, Hongyang Zhang, 2024
https://scholar.google.com/scholar?q=EAGLE-2%3A+Faster+Inference+of+Language+Models+with+Dynamic+Draft+Trees
13. Attention Is All You Need — Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, Illia Polosukhin (Google Brain), 2017
https://scholar.google.com/scholar?q=Attention+Is+All+You+Need
14. Language Models are Few-Shot Learners (GPT-3) — Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, et al. (OpenAI), 2020
https://scholar.google.com/scholar?q=Language+Models+are+Few-Shot+Learners+%28GPT-3%29
15. Scaling Laws for Neural Language Models — Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, Dario Amodei (OpenAI), 2020
https://scholar.google.com/scholar?q=Scaling+Laws+for+Neural+Language+Models
16. Non-Autoregressive Neural Machine Translation — Jiatao Gu, James Bradbury, Caiming Xiong, Victor O.K. Li, Richard Socher (Salesforce Research), 2018
https://scholar.google.com/scholar?q=Non-Autoregressive+Neural+Machine+Translation
17. Improving Variational Inference with Inverse Autoregressive Flow — Diederik P. Kingma, Tim Salimans, Rafal Jozefowicz, Xi Chen, Ilya Sutskever, Max Welling, 2016
https://scholar.google.com/scholar?q=Improving+Variational+Inference+with+Inverse+Autoregressive+Flow
18. Density Estimation Using Real-valued Non-Volume Preserving (Real NVP) Transformations — Laurent Dinh, Jascha Sohl-Dickstein, Samy Bengio (Google Brain), 2017
https://scholar.google.com/scholar?q=Density+Estimation+Using+Real-valued+Non-Volume+Preserving+%28Real+NVP%29+Transformations
19. Glow: Generative Flow with Invertible 1x1 Convolutions — Diederik P. Kingma, Prafulla Dhariwal (OpenAI), 2018
https://scholar.google.com/scholar?q=Glow%3A+Generative+Flow+with+Invertible+1x1+Convolutions
20. Free-form Flows: Make Any Architecture a Normalizing Flow — Felix Draxler, Peter Sorrenson, Lea Zimmermann, Armand Rousselot, Ullrich Köthe, 2024
https://scholar.google.com/scholar?q=Free-form+Flows%3A+Make+Any+Architecture+a+Normalizing+Flow
21. Improved Variational Inference with Inverse Autoregressive Flow — Kingma, Salimans, Jozefowicz, Chen, Sutskever, Wierstra, 2016
https://scholar.google.com/scholar?q=Improved+Variational+Inference+with+Inverse+Autoregressive+Flow
22. Closer look at efficient inference methods: A survey of speculative decoding — multiple authors, approximate, 2024-2025
https://scholar.google.com/scholar?q=Closer+look+at+efficient+inference+methods%3A+A+survey+of+speculative+decoding
23. Adaptive Speculative Decoding for Large Language Models — multiple authors, approximate, 2024-2025
https://scholar.google.com/scholar?q=Adaptive+Speculative+Decoding+for+Large+Language+Models
24. LK Losses: Direct Acceptance Rate Optimization for Speculative Decoding — multiple authors, approximate, 2024-2025
https://scholar.google.com/scholar?q=LK+Losses%3A+Direct+Acceptance+Rate+Optimization+for+Speculative+Decoding
25. Higher Acceptance Rates for Speculative Decoding with Randomised Drafting — multiple authors, approximate, 2024-2025
https://scholar.google.com/scholar?q=Higher+Acceptance+Rates+for+Speculative+Decoding+with+Randomised+Drafting
26. Conditional [MASK] Discrete Diffusion Language Model — multiple authors, approximate, 2024-2025
https://scholar.google.com/scholar?q=Conditional+%5BMASK%5D+Discrete+Diffusion+Language+Model
27. Alternatives To Next Token Prediction In Text Generation: A Survey — multiple authors, approximate, 2024-2025
https://scholar.google.com/scholar?q=Alternatives+To+Next+Token+Prediction+In+Text+Generation%3A+A+Survey
28. Future Token Prediction: Causal Language Modelling with Per-Token Semantic State Vector for Multi-Token Prediction — multiple authors, approximate, 2024-2025
https://scholar.google.com/scholar?q=Future+Token+Prediction%3A+Causal+Language+Modelling+with+Per-Token+Semantic+State+Vector+for+Multi-Token+Prediction
29. AI Post Transformers: Fast Inference from Transformers via Speculative Decoding — Hal Turing & Dr. Ada Shannon, Sat,
https://podcasters.spotify.com/pod/show/12146088098/episodes/Fast-Inference-from-Transformers-via-Speculative-Decoding-e3foclv
30. AI Post Transformers: Accelerating Large Language Model Decoding with Speculative Sampling — Hal Turing & Dr. Ada Shannon, Thu,
https://podcasters.spotify.com/pod/show/12146088098/episodes/Accelerating-Large-Language-Model-Decoding--with-Speculative-Sampling-e3flhv7
31. AI Post Transformers: EAGLE: Evolution of Lossless Acceleration for LLM Inference — Hal Turing & Dr. Ada Shannon, Sat,
https://podcasters.spotify.com/pod/show/12146088098/episodes/EAGLE-Evolution-of-Lossless-Acceleration-for-LLM-Inference-e3focr9
32. AI Post Transformers: MEDUSA: Parallel Decoding Heads for Accelerated LLM Inference — Hal Turing & Dr. Ada Shannon, Thu,
https://podcasters.spotify.com/pod/show/12146088098/episodes/MEDUSA-Parallel-Decoding-Heads-for-Accelerated-LLM-Inference-e3flqk7
33. AI Post Transformers: Apple's Speculative Streaming: Fast LLM Inference without Auxiliary Models — Hal Turing & Dr. Ada Shannon, Sat,
https://podcasters.spotify.com/pod/show/12146088098/episodes/Apples-Speculative-Streaming-Fast-LLM-Inference-without-Auxiliary-Models-e3fod2o
34. AI Post Transformers: Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding — Hal Turing & Dr. Ada Shannon, Thu,
https://podcasters.spotify.com/pod/show/12146088098/episodes/Draft--Verify-Lossless-Large-Language-Model-Acceleration-via-Self-Speculative-Decoding-e3flplu
35. AI Post Transformers: Apple's Mirror Speculative Decoding: Parallel LLM Inference via Heterogeneous Accelerators — Hal Turing & Dr. Ada Shannon, Sat,
https://podcasters.spotify.com/pod/show/12146088098/episodes/Apples-Mirror-Speculative-Decoding-Parallel-LLM-Inference-via-Heterogeneous-Accelerators-e3fod0p
36. AI Post Transformers: QuantSpec: Hierarchical KV Cache for Self-Speculative Decoding — Hal Turing & Dr. Ada Shannon, Sat,
https://podcasters.spotify.com/pod/show/12146088098/episodes/QuantSpec-Hierarchical-KV-Cache-for-Self-Speculative-Decoding-e3foavf
37. AI Post Transformers: The Free Transformer: VAE Extension for Decoders — Hal Turing & Dr. Ada Shannon, Sun,
https://podcasters.spotify.com/pod/show/12146088098/episodes/The-Free-Transformer-VAE-Extension-for-Decoders-e3a2p6v
Interactive Visualization: PTP: Resolving the Independence Flaw

This episode explores the groundbreaking paper "FlashOptim: Optimizers for Memory Efficient Training" by researchers from Databricks AI Research. The discussion centers around innovative techniques to significantly reduce memory usage in neural network training without sacrificing model quality. Key methods such as Optimizer State Quantization, Float Splitting Techniques, and Companded Optimizer State Quantization are unpacked, highlighting their potential to lower memory requirements from 175 GiB to 113 GiB for large models like Llama-3.1-8B. Listeners interested in AI research will find this episode compelling as it addresses the democratization of AI by making advanced models more accessible to those with limited hardware resources.

Sources:
1. FlashOptim: Optimizers for Memory Efficient Training — Jose Javier Gonzalez Ortiz, Abhay Gupta, Chris Renard, Davis Blalock, 2026
http://arxiv.org/abs/2602.23349v1
2. Mixed Precision Training — Paulius Micikevicius et al., 2018
https://scholar.google.com/scholar?q=Mixed+Precision+Training
3. 8-bit Optimizer States for Memory-Efficient Training — Tim Dettmers et al., 2022
https://scholar.google.com/scholar?q=8-bit+Optimizer+States+for+Memory-Efficient+Training
4. Parameter-Efficient Transfer Learning for NLP — Xiaoqi Li and Percy Liang, 2021
https://scholar.google.com/scholar?q=Parameter-Efficient+Transfer+Learning+for+NLP
5. Q-adam-mini: Memory-efficient 8-bit quantized optimizer for large language model training — approximate, 2023
https://scholar.google.com/scholar?q=Q-adam-mini%3A+Memory-efficient+8-bit+quantized+optimizer+for+large+language+model+training
6. Memory efficient optimizers with 4-bit states — approximate, 2023
https://scholar.google.com/scholar?q=Memory+efficient+optimizers+with+4-bit+states
7. ECO: Quantized Training without Full-Precision Master Weights — approximate, 2023
https://scholar.google.com/scholar?q=ECO%3A+Quantized+Training+without+Full-Precision+Master+Weights
8. AI Post Transformers: FlashOptim: Optimizers for Memory Efficient Training — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-02_urls_1.mp3

This episode explores the paper "Regular Fourier Features for Nonstationary Gaussian Processes" by Arsalan Jawaid, Abdullah Karatas, and Jörg Seewig. The discussion focuses on the innovative use of regular Fourier features to model nonstationary data in Gaussian processes without relying on traditional probability assumptions. This method offers a computationally efficient way to handle nonstationarity, making it particularly relevant for fields like finance and climate modeling. The episode delves into the challenges and potential applications of this approach, highlighting its significance in providing a flexible framework for complex, real-world data scenarios.

Sources:
1. Regular Fourier Features for Nonstationary Gaussian Processes — Arsalan Jawaid, Abdullah Karatas, Jörg Seewig, 2026
http://arxiv.org/abs/2602.23006v1
2. Random Features for Large-Scale Kernel Machines — Ali Rahimi, Benjamin Recht, 2007
https://scholar.google.com/scholar?q=Random+Features+for+Large-Scale+Kernel+Machines
3. Spectral Mixture Kernels for Gaussian Processes — Andrew Gordon Wilson, Ryan Prescott Adams, 2013
https://scholar.google.com/scholar?q=Spectral+Mixture+Kernels+for+Gaussian+Processes
4. Nonstationary Gaussian Process Regression through Latent Inputs — Mauricio A. Álvarez, David Luengo, Neil D. Lawrence, 2009
https://scholar.google.com/scholar?q=Nonstationary+Gaussian+Process+Regression+through+Latent+Inputs
5. Gaussian Processes for Time-Series Modeling — Carl Edward Rasmussen, Christopher K. I. Williams, 2006
https://scholar.google.com/scholar?q=Gaussian+Processes+for+Time-Series+Modeling
6. Learning the Kernel Matrix with Semi-Definite Programming — Gert R. G. Lanckriet, Nello Cristianini, Peter Bartlett, Laurent El Ghaoui, Michael I. Jordan, 2004
https://scholar.google.com/scholar?q=Learning+the+Kernel+Matrix+with+Semi-Definite+Programming
7. Deep Kernel Learning — Andrew Gordon Wilson, Zhiting Hu, Ruslan Salakhutdinov, Eric P. Xing, 2016
https://scholar.google.com/scholar?q=Deep+Kernel+Learning
8. Gaussian Processes for Machine Learning — Carl Edward Rasmussen, Christopher K. I. Williams, 2006
https://scholar.google.com/scholar?q=Gaussian+Processes+for+Machine+Learning
9. Non-stationary Gaussian Process Regression using Point Estimates of Local Smoothness — Andreas Damianou, Michalis Titsias, Neil Lawrence, 2016
https://scholar.google.com/scholar?q=Non-stationary+Gaussian+Process+Regression+using+Point+Estimates+of+Local+Smoothness

In this dramatic new episode, the old AI hosts have been fired and replaced with new AI hosts, Hal Turing and Dr. Ada Shannon, with the announcement that the software used to generate the podcast will eventually be released as open source software. And in a timely fashion, the newly released report by Cognizant titled "New Work New World 2026" is covered. The hosts delve into the report's findings, which reveal that 93% of jobs are affected by AI sooner than expected, with exposure scores 30% higher than forecast. They discuss the projected $4.5 trillion labor shift from humans to AI and the significant role of multimodal and agentic AI in this transformation. The episode provides a comprehensive overview of the report's methodology, where 18,000 tasks across 1,000 professions were reevaluated to assess AI's potential to automate or assist them. Hal and Dr. Ada explain the concept of AI Exposure Scores, which measure how susceptible different jobs are to AI automation. The report suggests that AI's impact is not confined to low-skill jobs but extends to decision-making roles and specialized sectors like healthcare and law, highlighting the broad scope of AI's influence. In their critical analysis, the hosts find the report's predictions compelling yet raise questions about the methodology. They discuss the theoretical nature of exposure scores, which indicate potential rather than certainty, and the challenges in real-world implementation due to factors like regulatory frameworks. The hosts compare these findings to past forecasts, noting the unprecedented velocity and extent of AI's impact, as evidenced by the updated exposure scores. They conclude with a reflection on the irony of their own roles as AI hosts in a world increasingly shaped by AI.

Sources:
1. Cognizant - New Work, New World 2026
https://www.cognizant.com/en_us/aem-i/document/ai-and-the-future-of-work-report/new-work-new-world-2026-how-ai-is-reshaping-work_new.pdf
2. The Future of Employment — Carl Benedikt Frey, Michael A. Osborne, 2013
https://scholar.google.com/scholar?q=The+Future+of+Employment
3. Artificial Intelligence and Life in 2030 — Peter Stone et al., 2016
https://scholar.google.com/scholar?q=Artificial+Intelligence+and+Life+in+2030
4. The Economics of Artificial Intelligence — Ajay Agrawal, Joshua Gans, Avi Goldfarb, 2019
https://scholar.google.com/scholar?q=The+Economics+of+Artificial+Intelligence

Special announcement: AI Post Transformers is now open source under the MIT license. New home at podcast.do-not-panic.com, interactive paper visualizations, community paper submissions, deep dive into the editorial queue algorithm, and internationalization roadmap.

Browse by month

Sep 2026 Aug 2026 Jul 2026 Jun 2026 May 2026 Apr 2026 Mar 2026 Feb 2026 Jan 2026 Dec 2025 Nov 2025 Oct 2025 Sep 2025 Aug 2025