This episode explores KVzap, a method for pruning transformer KV caches by learning a cheap surrogate for a much stronger oracle, with the goal of making cache eviction practical during both prompt prefilling and token-by-token decoding. It explains why KV caches dominate long-context inference costs, clarifies the difference between prefilling and decoding, and lays out why serving systems have favored quantization and paging over content-aware token deletion: removing the wrong token can quietly break later answers. The discussion places KVzap alongside KVzip, Expected Attention, and DMS, arguing that its key advance is a learned per-layer, per-head importance predictor trained to imitate a richer KVzip+ teacher that measures not just attention but actual contribution to the residual stream. Listeners would find it interesting because it ties together systems bottlenecks, adaptive eviction policies such as delayed eviction and sliding windows, and concrete training choices into a broader case for faster, more faithful long-context inference.

Sources:
1. KVzap: Fast, Adaptive, Faithful KV Cache Pruning
https://arxiv.org/pdf/2601.07891
2. H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models — Zhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen, Beidi Chen, et al., 2023
https://scholar.google.com/scholar?q=H2O%3A+Heavy-Hitter+Oracle+for+Efficient+Generative+Inference+of+Large+Language+Models
3. SnapKV: LLM Knows What You are Looking for Before Generation — Yuhong Li, Yingbing Huang, Bowen Yang, Bharat Venkitesh, Patrick Lewis, et al., 2024
https://scholar.google.com/scholar?q=SnapKV%3A+LLM+Knows+What+You+are+Looking+for+Before+Generation
4. Expected Attention: KV Cache Compression by Estimating Attention from Future Queries Distribution — Alessio Devoto, Maximilian Jeblick, Simon Jegou, 2025
https://scholar.google.com/scholar?q=Expected+Attention%3A+KV+Cache+Compression+by+Estimating+Attention+from+Future+Queries+Distribution
5. KVzip: Query-Agnostic KV Cache Compression with Context Reconstruction — Jang-Hyun Kim, Jinuk Kim, Sangwoo Kwon, Jae W. Lee, Sangdoo Yun, Hyun Oh Song, 2025
https://scholar.google.com/scholar?q=KVzip%3A+Query-Agnostic+KV+Cache+Compression+with+Context+Reconstruction
6. Inference-Time Hyper-Scaling with KV Cache Compression — Adrian Lancucki, Konrad Staniszewski, Piotr Nawrot, Edoardo M. Ponti, 2025
https://scholar.google.com/scholar?q=Inference-Time+Hyper-Scaling+with+KV+Cache+Compression
7. Compactor: Calibrated Query-Agnostic KV Cache Compression with Approximate Leverage Scores — Vivek Chari, Benjamin Van Durme, 2025
https://scholar.google.com/scholar?q=Compactor%3A+Calibrated+Query-Agnostic+KV+Cache+Compression+with+Approximate+Leverage+Scores
8. DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads — Guangxuan Xiao, Jiaming Tang, Jingwei Zuo, Junxian Guo, Shang Yang, Haotian Tang, Yao Fu, Song Han, 2024
https://scholar.google.com/scholar?q=DuoAttention%3A+Efficient+Long-Context+LLM+Inference+with+Retrieval+and+Streaming+Heads
9. Retrieval Head Mechanistically Explains Long-Context Factuality — Wenhao Wu et al., 2024
https://scholar.google.com/scholar?q=Retrieval+Head+Mechanistically+Explains+Long-Context+Factuality
10. Query-Focused Retrieval Heads Improve Long-Context Reasoning and Re-ranking — Wuwei Zhang et al., 2025
https://scholar.google.com/scholar?q=Query-Focused+Retrieval+Heads+Improve+Long-Context+Reasoning+and+Re-ranking
11. MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention — Huiqiang Jiang et al., 2024
https://scholar.google.com/scholar?q=MInference+1.0%3A+Accelerating+Pre-filling+for+Long-Context+LLMs+via+Dynamic+Sparse+Attention
12. KV-Compress: Paged KV-Cache Compression with Variable Compression Rates per Attention Head — Isaac Rehg, 2024
https://scholar.google.com/scholar?q=KV-Compress%3A+Paged+KV-Cache+Compression+with+Variable+Compression+Rates+per+Attention+Head
13. PagedEviction: Structured Block-wise KV Cache Pruning for Efficient Large Language Model Inference — Krishna Teja Chitty-Venkata et al., 2025
https://scholar.google.com/scholar?q=PagedEviction%3A+Structured+Block-wise+KV+Cache+Pruning+for+Efficient+Large+Language+Model+Inference
14. KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows — Zaifeng Pan et al., 2025
https://scholar.google.com/scholar?q=KVFlow%3A+Efficient+Prefix+Caching+for+Accelerating+LLM-Based+Multi-Agent+Workflows
15. AI Post Transformers: Affordable Large-Scale Decoding Through Model-System Co-Design — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-19-affordable-large-scale-decoding-through-e1d7ed.mp3
16. AI Post Transformers: TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-25-turboquant-online-vector-quantiz-1967b7.mp3
17. AI Post Transformers: Deep Kernel Fusion for Transformer Decoding — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-deep-kernel-fusion-for-transformer-decod-b1a703.mp3
18. AI Post Transformers: DeepSeek-V4 and Practical Million-Token Context — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-25-deepseek-v4-and-practical-million-token-6f4de1.mp3
19. AI Post Transformers: How Induction Heads Emerge in Transformers — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-03-how-induction-heads-emerge-in-transforme-a7bfcb.mp3
20. AI Post Transformers: When Many-Shot CoT Becomes Test-Time Learning — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-when-many-shot-cot-becomes-test-time-lea-c25bfe.mp3

This episode explores KVzip, a query-agnostic method for compressing long-context KV caches so a model can reuse a shared document, codebase, or memory bank across many later questions without optimizing for just one query. It explains why KV cache has become a major systems bottleneck, including the striking example that a 120,000-token context for Qwen2.5-14B can require more memory for cache than for the model weights themselves. The discussion contrasts KVzip with exact prefix caching and query-aware pruning methods like SnapKV, then breaks down KVzip’s core idea: replay the original context, measure which cached states receive the most attention during reconstruction, and keep those as durable memory. Listeners would find it interesting because the paper ties a clean systems insight to concrete gains, reporting roughly 394x smaller decoding-time KV caches and about 2x lower FlashAttention latency across LLaMA, Qwen, and Gemma models on very long contexts.

Sources:
1. KVzip for Query-Agnostic KV Cache Compression
https://arxiv.org/pdf/2505.23416
2. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding — Jacob Devlin, Ming-Wei Chang, Kenton Lee, Kristina Toutanova, 2018
https://scholar.google.com/scholar?q=BERT%3A+Pre-training+of+Deep+Bidirectional+Transformers+for+Language+Understanding
3. SnapKV: LLM Knows What You are Looking for Before Generation — Yuhong Li, Yingbing Huang, Bowen Yang, Bharat Venkitesh, Acyr Locatelli, Hanchen Ye, Tianle Cai, Patrick Lewis, Deming Chen, 2024
https://scholar.google.com/scholar?q=SnapKV%3A+LLM+Knows+What+You+are+Looking+for+Before+Generation
4. KVzip: Query-Agnostic KV Cache Compression with Context Reconstruction — Jang-Hyun Kim, Jinuk Kim, Sangwoo Kwon, Jae W. Lee, Sangdoo Yun, Hyun Oh Song, 2025
https://scholar.google.com/scholar?q=KVzip%3A+Query-Agnostic+KV+Cache+Compression+with+Context+Reconstruction
5. Rethinking Key-Value Cache Compression Techniques for Large Language Model Serving — Wei Gao, Xinyu Zhou, Peng Sun, Tianwei Zhang, Yonggang Wen, 2025
https://scholar.google.com/scholar?q=Rethinking+Key-Value+Cache+Compression+Techniques+for+Large+Language+Model+Serving
6. SCBench: A KV Cache-Centric Analysis of Long-Context Methods — Yudong Li, Hongkang Jiang, Qihui Wu, Xintong Luo, Sohee Ahn, Chen Zhang, and others, 2025
https://scholar.google.com/scholar?q=SCBench%3A+A+KV+Cache-Centric+Analysis+of+Long-Context+Methods
7. DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads — Guangxuan Xiao, Jiaming Tang, Jingwei Zuo, Junxian Guo, Shang Yang, Haotian Tang, Yao Fu, Song Han, 2025
https://scholar.google.com/scholar?q=DuoAttention%3A+Efficient+Long-Context+LLM+Inference+with+Retrieval+and+Streaming+Heads
8. Compactor: Calibrated Query-Agnostic KV Cache Compression with Approximate Leverage Scores — Vivek Chari, Benjamin Van Durme, 2025
https://scholar.google.com/scholar?q=Compactor%3A+Calibrated+Query-Agnostic+KV+Cache+Compression+with+Approximate+Leverage+Scores
9. No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization — June Yong Yang, Byeongwook Kim, Jeongin Bae, Beomseok Kwon, Gunho Park, Eunho Yang, Se Jung Kwon, Dongsoo Lee, 2024
https://scholar.google.com/scholar?q=No+Token+Left+Behind%3A+Reliable+KV+Cache+Compression+via+Importance-Aware+Mixed+Precision+Quantization
10. Safety Alignment Should Be Made More Than Just a Few Tokens Deep — Xiangyu Qi, Ashwinee Panda, Kaifeng Lyu, Xiao Ma, Subhrajit Roy, Ahmad Beirami, Prateek Mittal, Peter Henderson, 2025
https://scholar.google.com/scholar?q=Safety+Alignment+Should+Be+Made+More+Than+Just+a+Few+Tokens+Deep
11. The Residual Stream Is All You Need: On the Redundancy of the KV Cache in Transformer Inference — Kaleem Ullah Qasim et al., 2026
https://arxiv.org/abs/2603.19664
12. DeltaKV: Residual-Based KV Cache Compression via Long-Range Similarity — Jitai Hao et al., 2026
https://arxiv.org/abs/2602.08005
13. ParisKV: Fast and Drift-Robust KV-Cache Retrieval for Long-Context LLMs — Yanlin Qi et al., 2026
https://arxiv.org/abs/2602.07721
14. HeteroCache: A Dynamic Retrieval Approach to Heterogeneous KV Cache Compression for Long-Context LLM Inference — Zhiyuan Shi et al., 2026
https://arxiv.org/abs/2601.13684
15. R-KV: Redundancy-aware KV Cache Compression for Reasoning Models — Zefan Cai et al., 2025
https://arxiv.org/abs/2505.24133
16. Hold Onto That Thought: Assessing KV Cache Compression On Reasoning — Minghui Liu et al., 2025
https://arxiv.org/abs/2512.12008
17. SideQuest: Model-Driven KV Cache Management for Long-Horizon Agentic Reasoning — Sanjay Kariyappa and G. Edward Suh, 2026
https://arxiv.org/abs/2602.22603
18. AI Post Transformers: PackKV Lossy Compression for KV Caches — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-04-packkv-lossy-compression-for-kv-caches-b37bce.mp3
19. AI Post Transformers: Affordable Large-Scale Decoding Through Model-System Co-Design — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-19-affordable-large-scale-decoding-through-e1d7ed.mp3
20. AI Post Transformers: Deep Kernel Fusion for Transformer Decoding — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-deep-kernel-fusion-for-transformer-decod-b1a703.mp3
21. AI Post Transformers: DeepSeek-V4 and Practical Million-Token Context — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-25-deepseek-v4-and-practical-million-token-6f4de1.mp3
22. AI Post Transformers: TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-25-turboquant-online-vector-quantiz-1967b7.mp3

This episode explores a systems paper that extends GPU memory through CXL-attached DRAM and SSDs, asking whether accelerators can reach beyond on-board HBM without the usual overhead of software-driven memory migration. It explains CXL, memory disaggregation, and the difference between local GPU memory, host-managed memory, CXL memory, and storage-backed expansion, while grounding the discussion in earlier work such as Infiniswap, DirectCXL, and Microsoft’s Pond. The conversation focuses on the paper’s main technical claim: custom GPU-side hardware, including RTL CXL controllers, multiple root ports, and latency-hiding policies, could make expanded memory tiers more usable than approaches like UVM or GPUDirect Storage. It is interesting because the speakers both highlight the engineering ambition and press on a central unresolved question: whether these ideas truly help real transformer workloads, rather than only looking good on more conventional benchmark traces.

Interactive Visualization: CXL-GPU and Beyond Onboard Memory
Sources:
1. CXL-GPU: Pushing GPU Memory Boundaries with the Integration of CXL Technologies — Donghyun Gouk, Seungkwan Kang, Seungjun Lee, Jiseon Kim, Kyungkuk Nam, Eojin Ryu, Sangwon Lee, Dongpyung Kim, Junhyeok Jang, Hanyeoreum Bae, Myoungsoo Jung, 2025
http://arxiv.org/abs/2506.15601
2. Disaggregated Memory for Expansion and Sharing in Blade Servers — Kevin Lim, Jichuan Chang, Trevor Mudge, Parthasarathy Ranganathan, Steven K. Reinhardt, Thomas F. Wenisch, 2009
https://scholar.google.com/scholar?q=Disaggregated+Memory+for+Expansion+and+Sharing+in+Blade+Servers
3. Efficient Memory Disaggregation with Infiniswap — Juncheng Gu, Youngmoon Lee, Yiwen Zhang, Mosharaf Chowdhury, Kang G. Shin, 2017
https://scholar.google.com/scholar?q=Efficient+Memory+Disaggregation+with+Infiniswap
4. Direct Access, High-Performance Memory Disaggregation with DirectCXL — Donghyun Gouk, Sangwon Lee, Miryeong Kwon, Myoungsoo Jung, 2022
https://scholar.google.com/scholar?q=Direct+Access%2C+High-Performance+Memory+Disaggregation+with+DirectCXL
5. Pond: CXL-Based Memory Pooling Systems for Cloud Platforms — Huaicheng Li, Daniel S. Berger, Lisa Hsu, Daniel Ernst, Pantea Zardoshti, Stanko Novakovic, Monish Shah, Samir Rajadnya, Scott Lee, Ishwar Agarwal, Mark D. Hill, Marcus Fontoura, Ricardo Bianchini, 2023
https://scholar.google.com/scholar?q=Pond%3A+CXL-Based+Memory+Pooling+Systems+for+Cloud+Platforms
6. SMT: Software-Defined Memory Tiering for Heterogeneous Computing Systems with CXL Memory Expander — K. Kim, H. Kim, J. So, W. Lee, J. Im, S. Park, J. Cho, H. Song, 2023
https://scholar.google.com/scholar?q=SMT%3A+Software-Defined+Memory+Tiering+for+Heterogeneous+Computing+Systems+with+CXL+Memory+Expander
7. TPP: Transparent Page Placement for CXL-Enabled Tiered-Memory — Hasan Al Maruf, Hao Wang, Abhishek Dhanotia, Johannes Weiner, Niket Agarwal, Pallab Bhattacharya, Chris Petersen, Mosharaf Chowdhury, Shobhit Kanaujia, Prakash Chauhan, 2023
https://scholar.google.com/scholar?q=TPP%3A+Transparent+Page+Placement+for+CXL-Enabled+Tiered-Memory
8. NVMMU: A Non-volatile Memory Management Unit for Heterogeneous GPU-SSD Architectures — Jie Zhang, David Donofrio, John Shalf, Mahmut T. Kandemir, Myoungsoo Jung, 2015
https://scholar.google.com/scholar?q=NVMMU%3A+A+Non-volatile+Memory+Management+Unit+for+Heterogeneous+GPU-SSD+Architectures
9. Overcoming the Memory Wall with CXL-Enabled SSDs — Shao-Peng Yang, Minjae Kim, Sanghyun Nam, Juhyung Park, Jin-yong Choi, Eyee Hyun Nam, Eunji Lee, Sungjin Lee, Bryan S. Kim, 2023
https://scholar.google.com/scholar?q=Overcoming+the+Memory+Wall+with+CXL-Enabled+SSDs
10. NeoMem: Hardware/Software Co-Design for CXL-Native Memory Tiering — Zhe Zhou, Yiqi Chen, Tao Zhang, Yang Wang, Ran Shu, Shuotao Xu, Peng Cheng, Lei Qu, Yongqiang Xiong, Jie Zhang, Guangyu Sun, 2024
https://scholar.google.com/scholar?q=NeoMem%3A+Hardware%2FSoftware+Co-Design+for+CXL-Native+Memory+Tiering
11. ARIADNE: Adaptive UVM Management for Efficient GPU Memory Oversubscription — approx. recent systems authors, 2024/2025
https://scholar.google.com/scholar?q=ARIADNE%3A+Adaptive+UVM+Management+for+Efficient+GPU+Memory+Oversubscription
12. MOST: Memory Oversubscription-Aware Scheduling for Tensor Migration on GPU Unified Storage — approx. recent systems authors, 2024/2025
https://scholar.google.com/scholar?q=MOST%3A+Memory+Oversubscription-Aware+Scheduling+for+Tensor+Migration+on+GPU+Unified+Storage
13. Selective memory compression for GPU memory oversubscription management — approx. recent architecture authors, 2024/2025
https://scholar.google.com/scholar?q=Selective+memory+compression+for+GPU+memory+oversubscription+management
14. Phoenix: A Refactored I/O Stack for GPU Direct Storage without Phony Buffers — approx. recent storage/systems authors, 2024/2025
https://scholar.google.com/scholar?q=Phoenix%3A+A+Refactored+I%2FO+Stack+for+GPU+Direct+Storage+without+Phony+Buffers
15. Managing Scalable Direct Storage Accesses for GPUs with GoFS — approx. recent storage/systems authors, 2024/2025
https://scholar.google.com/scholar?q=Managing+Scalable+Direct+Storage+Accesses+for+GPUs+with+GoFS
16. CCCL: Node-Spanning GPU Collectives with CXL Memory Pooling — approx. recent distributed systems authors, 2024/2025
https://scholar.google.com/scholar?q=CCCL%3A+Node-Spanning+GPU+Collectives+with+CXL+Memory+Pooling
17. Efficient Tensor Offloading Based on CXL Memory Pool For Extreme Scale Deep Learning — approx. recent ML systems authors, 2024/2025
https://scholar.google.com/scholar?q=Efficient+Tensor+Offloading+Based+on+CXL+Memory+Pool+For+Extreme+Scale+Deep+Learning
18. UHM: Unified Transferring and Pooling over Heterogeneous GPU Memories — approx. recent memory-systems authors, 2024/2025
https://scholar.google.com/scholar?q=UHM%3A+Unified+Transferring+and+Pooling+over+Heterogeneous+GPU+Memories
19. GPUVM: GPU-driven unified virtual memory — approx. recent architecture authors, 2024/2025
https://scholar.google.com/scholar?q=GPUVM%3A+GPU-driven+unified+virtual+memory
20. Salus: Efficient security support for cxl-expanded gpu memory — approx. recent security/systems authors, 2024/2025
https://scholar.google.com/scholar?q=Salus%3A+Efficient+security+support+for+cxl-expanded+gpu+memory
21. AI Post Transformers: Vistara Brings CXL Memory to Hyperscale — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-11-vistara-brings-cxl-memory-to-hyperscale-b5199e.mp3
22. AI Post Transformers: FengHuang for Rack-Scale LLM Inference Memory — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-12-fenghuang-for-rack-scale-llm-inference-m-62708e.mp3
23. AI Post Transformers: Why LLM Serving Needs Mathematical Optimization — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-why-llm-serving-needs-mathematical-optim-647fc6.mp3
24. AI Post Transformers: Affordable Large-Scale Decoding Through Model-System Co-Design — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-19-affordable-large-scale-decoding-through-e1d7ed.mp3
Interactive Visualization: CXL-GPU and Beyond Onboard Memory

This episode explores Beluga, a systems paper that tackles one of the biggest practical bottlenecks in long-context LLM inference: how to store and retrieve massive KV caches when GPU memory is no longer enough. It explains why traditional RDMA-based memory disaggregation is cumbersome and how Beluga uses CXL-based pooled memory to give GPUs and CPUs more direct, load/store-style access to shared cache data, reducing copies, staging, and synchronization overhead. The discussion digs into the architecture’s tradeoffs, including the fact that CXL is still slower than local HBM or DRAM, but argues that its simpler access model can still deliver large gains in the right workload regime. Listeners would find it interesting for its concrete analysis of when the reported speedups, including major reductions in time to first token and large throughput gains, are real advances versus artifacts of favorable cache-reuse conditions.

Sources:
1. Beluga: A CXL-Based Memory Architecture for Scalable and Efficient LLM KVCache Management — Xinjun Yang, Qingda Hu, Junru Li, Feifei Li, Yicong Zhu, Yuqi Zhou, Qiuru Lin, Jian Dai, Yang Kong, Jiayu Zhang, Guoqiang Xu, Qiang Liu, 2025
http://arxiv.org/abs/2511.20172
2. Memory Pooling With CXL — Donghyun Gouk, Miryeong Kwon, Hanyeoreum Bae, Sangwon Lee, Myoungsoo Jung, 2023
https://scholar.google.com/scholar?q=Memory+Pooling+With+CXL
3. Demystifying CXL Memory with Genuine CXL-Ready Systems and Devices — Yan Sun, Yifan Yuan, Zeduo Yu, Reese Kuper, Chihun Song, Jinghan Huang, Houxiang Ji, Siddharth Agarwal, Jiaqi Lou, Ipoom Jeong, Ren Wang, Jung Ho Ahn, Tianyin Xu, Nam Sung Kim, 2023
https://scholar.google.com/scholar?q=Demystifying+CXL+Memory+with+Genuine+CXL-Ready+Systems+and+Devices
4. Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving — Ruoyu Qin, Zheming Li, Weiran He, Mingxing Zhang, Yongwei Wu, Weimin Zheng, Xinran Xu, 2025
https://scholar.google.com/scholar?q=Mooncake%3A+A+KVCache-centric+Disaggregated+Architecture+for+LLM+Serving
5. Exploring CXL-based KV Cache Storage for LLM Serving — Yupeng Tang, Runxiang Cheng, Ping Zhou, Tongping Liu, Fei Liu, Wei Tang, Kyoungryun Bae, Jianjun Chen, Wu Xiang, Rui Shi, 2024
https://scholar.google.com/scholar?q=Exploring+CXL-based+KV+Cache+Storage+for+LLM+Serving
6. MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool — Cunchen Hu, Heyang Huang, Junhao Hu, Jiang Xu, Xusheng Chen, Tao Xie, Chenxi Wang, Sa Wang, Yungang Bao, Ninghui Sun, Yizhou Shan, 2024
https://scholar.google.com/scholar?q=MemServe%3A+Context+Caching+for+Disaggregated+LLM+Serving+with+Elastic+Memory+Pool
7. Efficient Memory Management for Large Language Model Serving with PagedAttention — Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, Ion Stoica, 2023
https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention
8. CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving — Yuhan Liu, Hanchen Li, Yihua Cheng, Siddhant Ray, Yuyang Huang, Qizheng Zhang, Kuntai Du, Jiayi Yao, Shan Lu, Ganesh Ananthanarayanan, Michael Maire, Henry Hoffmann, Ari Holtzman, Junchen Jiang, 2024
https://scholar.google.com/scholar?q=CacheGen%3A+KV+Cache+Compression+and+Streaming+for+Fast+Large+Language+Model+Serving
9. PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference — Dongjie Yang et al., 2024
https://scholar.google.com/scholar?q=PyramidInfer%3A+Pyramid+KV+Cache+Compression+for+High-throughput+LLM+Inference
10. ChunkKV: Semantic-Preserving KV Cache Compression for Efficient Long-Context LLM Inference — Xiang Liu et al., 2025
https://scholar.google.com/scholar?q=ChunkKV%3A+Semantic-Preserving+KV+Cache+Compression+for+Efficient+Long-Context+LLM+Inference
11. Inference-Time Hyper-Scaling with KV Cache Compression — Adrian Łańcucki, Konrad Staniszewski, Piotr Nawrot, Edoardo M. Ponti, 2025
https://scholar.google.com/scholar?q=Inference-Time+Hyper-Scaling+with+KV+Cache+Compression
12. FlowKV: A Disaggregated Inference Framework with Low-Latency KV Cache Transfer and Load-Aware Scheduling — Weiqing Li et al., 2025
https://scholar.google.com/scholar?q=FlowKV%3A+A+Disaggregated+Inference+Framework+with+Low-Latency+KV+Cache+Transfer+and+Load-Aware+Scheduling
13. TokenLake: A Unified Segment-level Prefix Cache Pool for Fine-grained Elastic Long-Context LLM Serving — Bingyang Wu et al., 2025
https://scholar.google.com/scholar?q=TokenLake%3A+A+Unified+Segment-level+Prefix+Cache+Pool+for+Fine-grained+Elastic+Long-Context+LLM+Serving
14. KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows — Zaifeng Pan et al., 2025
https://scholar.google.com/scholar?q=KVFlow%3A+Efficient+Prefix+Caching+for+Accelerating+LLM-Based+Multi-Agent+Workflows
15. Learned Prefix Caching for Efficient LLM Inference — Dongsheng Yang, Austin Li, Kai Li, Wyatt Lloyd, 2025
https://scholar.google.com/scholar?q=Learned+Prefix+Caching+for+Efficient+LLM+Inference
16. AI Post Transformers: FengHuang for Rack-Scale LLM Inference Memory — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-12-fenghuang-for-rack-scale-llm-inference-m-62708e.mp3
17. AI Post Transformers: CXL Computational Memory Offloading for Lower Runtime — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-cxl-computational-memory-offloading-for-3b2124.mp3
18. AI Post Transformers: Prefill-as-a-Service for Cross-Datacenter KV Cache — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-19-prefill-as-a-service-for-cross-datacente-7560be.mp3
19. AI Post Transformers: SolidAttention: Co-Designing Sparse Attention and SSD I/O — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-18-solidattention-co-designing-sparse-atten-5a8622.mp3
20. AI Post Transformers: Stochastic KV Routing for Cache Sharing — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-29-stochastic-kv-routing-for-cache-sharing-5fef63.mp3
21. AI Post Transformers: Affordable Large-Scale Decoding Through Model-System Co-Design — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-19-affordable-large-scale-decoding-through-e1d7ed.mp3
Interactive Visualization: Beluga: CXL Memory Pooling for LLM KV Cache

This episode explores a paper on inference-time scaling for coding agents, asking whether extra test-time compute still helps when tasks are long, messy, and require multi-step tool use rather than a single code completion. It focuses on the paper’s main argument that the real bottleneck is not generating more rollout attempts, but representing prior attempts well enough to compare, select, and reuse them, with structured trajectory summaries serving as the key middle layer between raw transcripts and final patches. The discussion examines two mechanisms: a parallel “tournament” style selection method over summaries, and a sequential refinement method that conditions later attempts on distilled lessons from earlier ones. Listeners would find it interesting because the conversation connects agent performance gains to practical questions of context management, selection versus reuse, and whether the reported improvements reflect a deep scaling insight or simply better engineering around long-horizon coding workflows.

Sources:
1. Scaling Test-Time Compute for Agentic Coding — Joongwon Kim, Wannan Yang, Kelvin Niu, Hongming Zhang, Yun Zhu, Eryk Helenowski, Ruan Silva, Zhengxing Chen, Srinivasan Iyer, Manzil Zaheer, Daniel Fried, Hannaneh Hajishirzi, Sanjeev Arora, Gabriel Synnaeve, Ruslan Salakhutdinov, Anirudh Goyal, 2026
http://arxiv.org/abs/2604.16529
2. ReAct: Synergizing Reasoning and Acting in Language Models — Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, Yuan Cao, 2023
https://scholar.google.com/scholar?q=ReAct%3A+Synergizing+Reasoning+and+Acting+in+Language+Models
3. Reflexion: Language Agents with Verbal Reinforcement Learning — Noah Shinn, Federico Cassano, Edward Berman, Ashwin Gopinath, Karthik Narasimhan, Shunyu Yao, 2023
https://scholar.google.com/scholar?q=Reflexion%3A+Language+Agents+with+Verbal+Reinforcement+Learning
4. ExpeL: LLM Agents Are Experiential Learners — Andrew Zhao, Daniel Huang, Quentin Xu, Matthieu Lin, Yong-Jin Liu, Gao Huang, 2023
https://scholar.google.com/scholar?q=ExpeL%3A+LLM+Agents+Are+Experiential+Learners
5. Rethinking Thinking Tokens: LLMs as Improvement Operators — Lovish Madaan, Aniket Didolkar, Suchin Gururangan, John Quan, Ruan Silva, Ruslan Salakhutdinov, Manzil Zaheer, Sanjeev Arora, Anirudh Goyal, 2025
https://scholar.google.com/scholar?q=Rethinking+Thinking+Tokens%3A+LLMs+as+Improvement+Operators
6. CodeMonkeys: Scaling Test-Time Compute for Software Engineering — Ryan Ehrlich, Bradley Brown, Jordan Juravsky, Ronald Clark, Christopher Re, Azalia Mirhoseini, 2025
https://scholar.google.com/scholar?q=CodeMonkeys%3A+Scaling+Test-Time+Compute+for+Software+Engineering
7. S*: Test Time Scaling for Code Generation — Dacheng Li, Shiyi Cao, Chengkun Cao, Xiuyu Li, Shangyin Tan, Kurt Keutzer, Jiarong Xing, Joseph E. Gonzalez, Ion Stoica, 2025
https://scholar.google.com/scholar?q=S%2A%3A+Test+Time+Scaling+for+Code+Generation
8. Scaling Test-time Compute for LLM Agents — King Zhu, Hanhao Li, Siwei Wu, Tianshun Xing, Dehua Ma, Xiangru Tang, Minghao Liu, Jian Yang, Jiaheng Liu, Yuchen Eleanor Jiang, Changwang Zhang, Chenghua Lin, Jun Wang, Ge Zhang, Wangchunshu Zhou, 2025
https://scholar.google.com/scholar?q=Scaling+Test-time+Compute+for+LLM+Agents
9. Agentic Test-Time Scaling for WebAgents — Nicholas Lee, Lutfi Eren Erdogan, Chris Joseph John, Surya Krishnapillai, Michael W. Mahoney, Kurt Keutzer, Amir Gholami, 2026
https://scholar.google.com/scholar?q=Agentic+Test-Time+Scaling+for+WebAgents
10. Does SWE-Bench-Verified Test Agent Ability or Model Memory? — Thanosan Prathifkumar, Noble Saji Mathews, Meiyappan Nagappan, 2025
https://scholar.google.com/scholar?q=Does+SWE-Bench-Verified+Test+Agent+Ability+or+Model+Memory%3F
11. A Benchmark for Procedural Memory Retrieval in Language Agents — Ishant Kohar, Aswanth Krishnan, 2025
https://scholar.google.com/scholar?q=A+Benchmark+for+Procedural+Memory+Retrieval+in+Language+Agents
12. PROCED-MEM: Benchmarking Procedural Memory Retrieval in Language Agents Across Domains — Ishant Kohar, Aswanth Krishnan, 2026
https://scholar.google.com/scholar?q=PROCED-MEM%3A+Benchmarking+Procedural+Memory+Retrieval+in+Language+Agents+Across+Domains
13. G-Memory: Tracing Hierarchical Memory for Multi-Agent Systems — Guibin Zhang, Muxin Fu, Guancheng Wan, Miao Yu, Kun Wang, Shuicheng Yan, 2025
https://scholar.google.com/scholar?q=G-Memory%3A+Tracing+Hierarchical+Memory+for+Multi-Agent+Systems
14. Scaling Agentic Verifier for Competitive Coding — Zeyao Ma et al., 2026
https://scholar.google.com/scholar?q=Scaling+Agentic+Verifier+for+Competitive+Coding
15. AgentPro: Enhancing LLM Agents with Automated Process Supervision — Yuchen Deng, Shichen Fan, Naibo Wang, Xinkui Zhao, See-Kiong Ng, 2025
https://scholar.google.com/scholar?q=AgentPro%3A+Enhancing+LLM+Agents+with+Automated+Process+Supervision
16. Recursive Introspection: Teaching Language Model Agents How to Self-Improve — Yuxiao Qu, Tianjun Zhang, Naman Garg, Aviral Kumar, 2024
https://scholar.google.com/scholar?q=Recursive+Introspection%3A+Teaching+Language+Model+Agents+How+to+Self-Improve
17. Agentic Refactoring: An Empirical Study of AI Coding Agents — Kosei Horikawa, Hao Li, Yutaro Kashiwa, Bram Adams, Hajimu Iida, Ahmed E. Hassan, 2025
https://scholar.google.com/scholar?q=Agentic+Refactoring%3A+An+Empirical+Study+of+AI+Coding+Agents
18. AI Post Transformers: TMAS: Scaling Test-Time Compute with Multi-Agent Synergy — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-14-tmas-scaling-test-time-compute-with-mult-3abe7a.mp3
19. AI Post Transformers: Benchmarking Test-Time Scaling for General LLM Agents — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-22-benchmarking-test-time-scaling-for-gener-8f14f9.mp3
20. AI Post Transformers: MiA-Signature and Global Activation for Long Context — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-13-mia-signature-and-global-activation-for-5ad62f.mp3
21. AI Post Transformers: Explicit Information Transmission for Context Compression — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-explicit-information-transmission-for-co-24e3c2.mp3
Interactive Visualization: Trajectory Summaries for Long-Horizon Coding Agents

This episode explores the DFX system, a four-FPGA appliance designed to accelerate transformer-based text generation by targeting a key weakness of GPUs: low-batch, token-by-token decode. It explains the difference between prompt processing and sequential generation, connects the paper’s older terminology to today’s prefill/decode framing, and shows why autoregressive inference often leaves GPU hardware underused even when training runs efficiently in parallel. The discussion also breaks down how DFX uses hardware-aware model parallelism and end-to-end accelerator design, rather than only speeding up isolated transformer subcomponents, to argue for lower latency and better energy and cost efficiency than a four-V100 GPU server. Listeners would find it interesting for its clear historical perspective on transformer serving and for its skepticism about how much of the reported advantage comes from FPGA specialization versus the fairness of the GPU baseline.

Sources:
1. DFX: A Low-latency Multi-FPGA Appliance for Accelerating Transformer-based Text Generation — Seongmin Hong, Seungjae Moon, Junsoo Kim, Sungjae Lee, Minsub Kim, Dongsoo Lee, Joo-Young Kim, 2022
http://arxiv.org/abs/2209.10797
2. GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism — Yanping Huang, Youlong Cheng, Ankur Bapna, Quoc V. Le, Yonghui Wu, Zhifeng Chen, and others, 2019
https://scholar.google.com/scholar?q=GPipe%3A+Efficient+Training+of+Giant+Neural+Networks+using+Pipeline+Parallelism
3. Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism — Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, Bryan Catanzaro, 2020
https://scholar.google.com/scholar?q=Megatron-LM%3A+Training+Multi-Billion+Parameter+Language+Models+Using+Model+Parallelism
4. GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding — Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Noam Shazeer, Zhifeng Chen, and others, 2020
https://scholar.google.com/scholar?q=GShard%3A+Scaling+Giant+Models+with+Conditional+Computation+and+Automatic+Sharding
5. Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM — Deepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley, Mostofa Patwary, Bryan Catanzaro, Amar Phanishayee, Matei Zaharia, and others, 2021
https://scholar.google.com/scholar?q=Efficient+Large-Scale+Language+Model+Training+on+GPU+Clusters+Using+Megatron-LM
6. Attention Is All You Need — Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, Illia Polosukhin, 2017
https://scholar.google.com/scholar?q=Attention+Is+All+You+Need
7. FTRANS: Energy-Efficient Acceleration of Transformers using FPGA — Jingcheng Rao, Yuchen Shao, Ke Wang, Zhihao Zhu, Xuehai Qian, Yiyu Shi, 2020
https://scholar.google.com/scholar?q=FTRANS%3A+Energy-Efficient+Acceleration+of+Transformers+using+FPGA
8. Fast Inference from Transformers via Speculative Decoding — Yaniv Leviathan, Matan Kalman, Yossi Matias, 2022
https://scholar.google.com/scholar?q=Fast+Inference+from+Transformers+via+Speculative+Decoding
9. PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference — Dongjie Yang, XiaoDong Han, Yan Gao, Yao Hu, Shilin Zhang, Hai Zhao, 2024
https://scholar.google.com/scholar?q=PyramidInfer%3A+Pyramid+KV+Cache+Compression+for+High-throughput+LLM+Inference
10. ChunkKV: Semantic-Preserving KV Cache Compression for Efficient Long-Context LLM Inference — Xiang Liu, Zhenheng Tang, Peijie Dong, Zeyu Li, Bo Li, Xuming Hu, Xiaowen Chu, 2025
https://scholar.google.com/scholar?q=ChunkKV%3A+Semantic-Preserving+KV+Cache+Compression+for+Efficient+Long-Context+LLM+Inference
11. Cost-Optimal Grouped-Query Attention for Long-Context LLMs — Yingfa Chen, Yutong Wu, Xu Han, Zhiyuan Liu, Maosong Sun, 2025
https://scholar.google.com/scholar?q=Cost-Optimal+Grouped-Query+Attention+for+Long-Context+LLMs
12. Optimised Grouped-Query Attention Mechanism for Transformers — Yuang Chen, Cheng Zhang, Xitong Gao, Robert D. Mullins, George A. Constantinides, Yiren Zhao, 2024
https://scholar.google.com/scholar?q=Optimised+Grouped-Query+Attention+Mechanism+for+Transformers
13. Prefill-Decode Aggregation or Disaggregation? Unifying Both for Goodput-Optimized LLM Serving — Chao Wang, Pengfei Zuo, Zhangyu Chen, Yunkai Liang, Zhou Yu, Ming-Chang Yang, 2025
https://scholar.google.com/scholar?q=Prefill-Decode+Aggregation+or+Disaggregation%3F+Unifying+Both+for+Goodput-Optimized+LLM+Serving
14. Nexus: Proactive Intra-GPU Disaggregation of Prefill and Decode in LLM Serving — Xiaoxiang Shi, Colin Cai, Junjia Du, Zhihao Jia, 2025
https://scholar.google.com/scholar?q=Nexus%3A+Proactive+Intra-GPU+Disaggregation+of+Prefill+and+Decode+in+LLM+Serving
15. SPAD: Specialized Prefill and Decode Hardware for Disaggregated LLM Inference — Hengrui Zhang, Pratyush Patel, August Ning, David Wentzlaff, 2025
https://scholar.google.com/scholar?q=SPAD%3A+Specialized+Prefill+and+Decode+Hardware+for+Disaggregated+LLM+Inference
16. AI Post Transformers: Deep Kernel Fusion for Transformer Decoding — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-deep-kernel-fusion-for-transformer-decod-b1a703.mp3
17. AI Post Transformers: Affordable Large-Scale Decoding Through Model-System Co-Design — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-19-affordable-large-scale-decoding-through-e1d7ed.mp3
18. AI Post Transformers: LAPS for Length-Aware LLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-laps-for-length-aware-llm-serving-0c6149.mp3
19. AI Post Transformers: Prefill-as-a-Service for Cross-Datacenter KV Cache — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-19-prefill-as-a-service-for-cross-datacente-7560be.mp3
20. AI Post Transformers: Speculative Decoding in Real vLLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-speculative-decoding-in-real-vllm-servin-6f4e2b.mp3
21. AI Post Transformers: FengHuang for Rack-Scale LLM Inference Memory — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-12-fenghuang-for-rack-scale-llm-inference-m-62708e.mp3
22. AI Post Transformers: Caffeine: A Unified FPGA for CNNs — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-06-caffeine-a-unified-fpga-for-cnns-e8acbe.mp3
Interactive Visualization: DFX: Multi-FPGA Acceleration for Transformer Inference

This episode explores Generative Recursive Reasoning, a paper that asks whether models can reason more effectively by repeatedly refining an internal latent state instead of externalizing long chains of thought as tokens. It explains how recursive reasoning trades parameter growth for inference-time computation, and how this approach may be especially useful for tasks like Sudoku, ARC-style problems, graph coloring, and N-Queens that benefit from iterative constraint solving. A central focus is the paper’s argument that reasoning should be stochastic rather than locked into a single deterministic path, using variational methods to model multiple possible latent trajectories and improve coverage when problems have more than one valid answer. The discussion is especially interesting because it contrasts this elegant search-like mechanism with mainstream transformer practice, highlighting both the promise of branching internal hypotheses and the practical reasons industry has not adopted such architectures at scale.

Sources:
1. Generative Recursive Reasoning in Latent Space
https://arxiv.org/pdf/2605.19376v1
2. Universal Transformers — Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, Lukasz Kaiser, 2018
https://scholar.google.com/scholar?q=Universal+Transformers
3. Looped Transformers are Better at Learning Learning Algorithms — Liu Yang, Kangwook Lee, Robert Nowak, Dimitris Papailiopoulos, 2024
https://scholar.google.com/scholar?q=Looped+Transformers+are+Better+at+Learning+Learning+Algorithms
4. Hierarchical Reasoning Model — Guan Wang, Jin Li, Yuhao Sun, Xing Chen, Changling Liu, Yue Wu, Meng Lu, Sen Song, Yasin Abbasi Yadkori, 2025
https://scholar.google.com/scholar?q=Hierarchical+Reasoning+Model
5. Less is More: Recursive Reasoning with Tiny Networks — Alexia Jolicoeur-Martineau, 2025
https://scholar.google.com/scholar?q=Less+is+More%3A+Recursive+Reasoning+with+Tiny+Networks
6. Probabilistic Tiny Recursive Model — Amin Sghaier, Ali Parviz, Alexia Jolicoeur-Martineau, 2026
https://scholar.google.com/scholar?q=Probabilistic+Tiny+Recursive+Model
7. Auto-Encoding Variational Bayes — Diederik P. Kingma, Max Welling, 2013/2014
https://scholar.google.com/scholar?q=Auto-Encoding+Variational+Bayes
8. Stochastic Backpropagation and Approximate Inference in Deep Generative Models — Danilo Jimenez Rezende, Shakir Mohamed, Daan Wierstra, 2014
https://scholar.google.com/scholar?q=Stochastic+Backpropagation+and+Approximate+Inference+in+Deep+Generative+Models
9. Inference Suboptimality in Variational Autoencoders — Chris Cremer, Xuechen Li, David Duvenaud, 2018
https://scholar.google.com/scholar?q=Inference+Suboptimality+in+Variational+Autoencoders
10. Iterative Amortized Inference — Joseph Marino, Yisong Yue, Stephan Mandt, 2018
https://scholar.google.com/scholar?q=Iterative+Amortized+Inference
11. Amortized Variational Inference: A Systematic Review — Ankush Ganguly, Sanjana Jain, Ukrit Watchareeruetai, 2023
https://scholar.google.com/scholar?q=Amortized+Variational+Inference%3A+A+Systematic+Review
12. Reasoning with Latent Thoughts: On the Power of Looped Transformers — Nikunj Saunshi, Nishanth Dikkala, Zhiyuan Li, Sanjiv Kumar, Sashank J. Reddi, 2025
https://scholar.google.com/scholar?q=Reasoning+with+Latent+Thoughts%3A+On+the+Power+of+Looped+Transformers
13. Structured Denoising Diffusion Models in Discrete State-Spaces — Jacob Austin, Daniel D. Johnson, Jonathan Ho, Daniel Tarlow, Rianne van den Berg, 2021
https://scholar.google.com/scholar?q=Structured+Denoising+Diffusion+Models+in+Discrete+State-Spaces
14. One Step Forward and K Steps Back: Better Reasoning with Denoising Recursion Models — Chris Cameron, Wangzheng Wang, Nikita Ivanov, Ashmita Bhattacharyya, Didier Chetelat, Yingxue Zhang, 2026
https://scholar.google.com/scholar?q=One+Step+Forward+and+K+Steps+Back%3A+Better+Reasoning+with+Denoising+Recursion+Models
15. LASAR: Latent Adaptive Semantic Aligned Reasoning for Generative Recommendation — approx. recent recommendation/LLM reasoning authors, 2025
https://scholar.google.com/scholar?q=LASAR%3A+Latent+Adaptive+Semantic+Aligned+Reasoning+for+Generative+Recommendation
16. Towards Inference-time Scaling for Continuous Space Reasoning — approx. recent latent reasoning authors, 2025
https://scholar.google.com/scholar?q=Towards+Inference-time+Scaling+for+Continuous+Space+Reasoning
17. GTS: Inference-Time Scaling of Latent Reasoning with a Learnable Gaussian Thought Sampler — approx. recent latent reasoning authors, 2025
https://scholar.google.com/scholar?q=GTS%3A+Inference-Time+Scaling+of+Latent+Reasoning+with+a+Learnable+Gaussian+Thought+Sampler
18. Latent Chain-of-Thought for Visual Reasoning — approx. recent visual reasoning authors, 2025
https://scholar.google.com/scholar?q=Latent+Chain-of-Thought+for+Visual+Reasoning
19. AI Post Transformers: TMAS: Scaling Test-Time Compute with Multi-Agent Synergy — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-14-tmas-scaling-test-time-compute-with-mult-3abe7a.mp3
20. AI Post Transformers: Agentic Discovery for Test-Time Scaling — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-12-agentic-discovery-for-test-time-scaling-f9a81f.mp3
21. AI Post Transformers: Latent Space as a New Computational Paradigm — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-05-latent-space-as-a-new-computational-para-810f39.mp3
22. AI Post Transformers: Causal-JEPA for Object-Level World Models — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-causal-jepa-for-object-level-world-model-311a8b.mp3
23. AI Post Transformers: LeWorldModel: Stable Joint-Embedding World Models from Pixels — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-25-leworldmodel-stable-joint-embedding-worl-650f9f.mp3
Interactive Visualization: Generative Recursive Reasoning in Latent Space

This episode explores a 2024 paper on the LPU, a custom processor designed specifically for large language model inference, with an emphasis on reducing the per-token delay that users notice in interactive systems. It explains why autoregressive decoding is often limited by memory movement and synchronization rather than raw compute, making conventional GPU strengths less decisive in small-batch, user-facing generation. The discussion highlights the paper’s full-stack argument: a specialized chip, a supporting software stack called HyperDex, and a multi-device link meant to preserve low latency while scaling across processors. Listeners would find it interesting because it reframes AI hardware performance around real conversational responsiveness and digs into whether the paper’s bold efficiency and scaling claims actually hold up under careful comparison.

Interactive Visualization: LPU Chip for Low-Latency LLM Inference
Sources:
1. LPU Chip for Low-Latency LLM Inference
https://arxiv.org/pdf/2408.07326
2. DFX: A Low-latency Multi-FPGA Appliance for Accelerating Transformer-based Text Generation — Seongmin Hong, et al., 2022
https://scholar.google.com/scholar?q=DFX%3A+A+Low-latency+Multi-FPGA+Appliance+for+Accelerating+Transformer-based+Text+Generation
3. SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head Pruning — Hanrui Wang, Zhekai Zhang, Song Han, 2021
https://scholar.google.com/scholar?q=SpAtten%3A+Efficient+Sparse+Attention+Architecture+with+Cascade+Token+and+Head+Pruning
4. A Software-Defined Tensor Streaming Multiprocessor for Large-Scale Machine Learning — Dennis Abts, et al., 2022
https://scholar.google.com/scholar?q=A+Software-Defined+Tensor+Streaming+Multiprocessor+for+Large-Scale+Machine+Learning
5. Efficient Memory Management for Large Language Model Serving with PagedAttention — Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, et al., 2023
https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention
6. DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving — Yinmin Zhong, Shengyu Zhao, et al., 2024
https://scholar.google.com/scholar?q=DistServe%3A+Disaggregating+Prefill+and+Decoding+for+Goodput-optimized+Large+Language+Model+Serving
7. RazorAttention: Efficient KV Cache Compression Through Retrieval Heads — Hanlin Tang et al., 2024
https://scholar.google.com/scholar?q=RazorAttention%3A+Efficient+KV+Cache+Compression+Through+Retrieval+Heads
8. Not All Heads Matter: A Head-Level KV Cache Compression Method with Integrated Retrieval and Reasoning — Yu Fu et al., 2024
https://scholar.google.com/scholar?q=Not+All+Heads+Matter%3A+A+Head-Level+KV+Cache+Compression+Method+with+Integrated+Retrieval+and+Reasoning
9. Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction — Haoran Qiu et al., 2024
https://scholar.google.com/scholar?q=Efficient+Interactive+LLM+Serving+with+Proxy+Model-based+Sequence+Length+Prediction
10. Slice-Level Scheduling for High Throughput and Load Balanced LLM Serving — Ke Cheng et al., 2024
https://scholar.google.com/scholar?q=Slice-Level+Scheduling+for+High+Throughput+and+Load+Balanced+LLM+Serving
11. Deferred Continuous Batching in Resource-Efficient Large Language Model Serving — Yongjun He, Yao Lu, Gustavo Alonso, 2024
https://scholar.google.com/scholar?q=Deferred+Continuous+Batching+in+Resource-Efficient+Large+Language+Model+Serving
12. TokenWeave: Efficient Compute-Communication Overlap for Distributed LLM Inference — Raja Gond, Nipun Kwatra, Ramachandran Ramjee, 2025
https://scholar.google.com/scholar?q=TokenWeave%3A+Efficient+Compute-Communication+Overlap+for+Distributed+LLM+Inference
13. Characterizing Communication Patterns in Distributed Large Language Model Inference — Lang Xu et al., 2025
https://scholar.google.com/scholar?q=Characterizing+Communication+Patterns+in+Distributed+Large+Language+Model+Inference
14. Mind the Memory Gap: Unveiling GPU Bottlenecks in Large-Batch LLM Inference — Pol G. Recasens et al., 2025
https://scholar.google.com/scholar?q=Mind+the+Memory+Gap%3A+Unveiling+GPU+Bottlenecks+in+Large-Batch+LLM+Inference
15. AI Post Transformers: FengHuang for Rack-Scale LLM Inference Memory — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-12-fenghuang-for-rack-scale-llm-inference-m-62708e.mp3
16. AI Post Transformers: Deep Kernel Fusion for Transformer Decoding — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-deep-kernel-fusion-for-transformer-decod-b1a703.mp3
17. AI Post Transformers: Affordable Large-Scale Decoding Through Model-System Co-Design — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-19-affordable-large-scale-decoding-through-e1d7ed.mp3
18. AI Post Transformers: Speculative Decoding in Real vLLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-speculative-decoding-in-real-vllm-servin-6f4e2b.mp3
19. AI Post Transformers: Why LLM Serving Needs Mathematical Optimization — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-why-llm-serving-needs-mathematical-optim-647fc6.mp3
20. AI Post Transformers: LAPS for Length-Aware LLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-laps-for-length-aware-llm-serving-0c6149.mp3
21. AI Post Transformers: Splitwise: Phase-Split LLM Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-26-splitwise-phase-split-llm-inference-e8945b.mp3
22. AI Post Transformers: JANUS for Scalable MoE Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-janus-for-scalable-moe-inference-78ae30.mp3
Interactive Visualization: LPU Chip for Low-Latency LLM Inference

Episode title: After Titans: Behrouz on Nested Learning and Hope The followup to our Titans episode, by the same core team a year later. Behrouz, Razaviyayn, Zhong, and Mirrokni (Google Research) generalize the Titans bet — that long-term memory should be a learnable module updated at test time — into a broader paradigm they call Nested Learning, where a "deep" architecture is really a hierarchy of nested optimization problems each compressing its own context flow. The episode walks through their three core contributions: (1) reframing standard optimizers like Adam and SGD-with-Momentum as associative-memory modules that compress gradient information, then proposing more expressive optimizers with their own deep memory; (2) a self-modifying sequence model whose update rule is itself learned end-to-end — the natural generalization of the test-time-learnable memory module Titans introduced; (3) a continuum memory system that replaces the traditional short-term-vs-long-term dichotomy with a continuum across multiple update rates. Combining the self-modifying module with the continuum memory system produces Hope, a continual-learning architecture reported promising on language modeling, knowledge incorporation, few-shot generalization, continual learning, and long-context reasoning. The hosts treat Hope as the next concrete instance of the Titans → Nested Learning research arc, stress-test the novelty against established meta-learning and fast-weights literature, and distinguish what the paper actually shows from what its framing suggests.

Sources:
Nested Learning: The Illusion of Deep Learning Architectures —
Ali Behrouz, Meisam Razaviyayn, Peilin Zhong, Vahab Mirrokni
(Google Research, NeurIPS 2025) — arXiv 2512.24695
https://arxiv.org/pdf/2512.24695
Titans: Learning to Memorize at Test Time — Ali Behrouz, Peilin
Zhong, Vahab Mirrokni (Jan 2025) — arXiv 2501.00663
https://arxiv.org/pdf/2501.00663
AI Post Transformers — "Titans: Learning to Memorize at Test Time"
https://podcast.do-not-panic.com/episodes/2026-05-20-titans-learning-to-memorize-at-test-time-054662.mp3
Interactive Visualization: After Titans: Behrouz on Nested Learning and Hope

This episode explores the Titans paper’s proposal to pair standard attention with a separate learned long-term memory that updates during inference, aiming to preserve distant information without paying full quadratic attention costs across very long sequences. It situates that idea against earlier approaches such as Neural Turing Machines, Transformer-XL, Compressive Transformers, Memorizing Transformers, and linear-attention recurrent models, highlighting the recurring tradeoff between precise recall and scalable memory. The discussion focuses on the paper’s most distinctive claim: memory writes are driven by a loss-based notion of surprise, making test-time memory updates look more like small online learning steps than a simple cache. Listeners would find it interesting because it gets at a central open question in modern AI systems design: whether neural networks can gain durable, useful memory at inference time without becoming too unstable, expensive, or operationally awkward to deploy.

Sources:
1. Titans: Learning to Memorize at Test Time
https://arxiv.org/pdf/2501.00663
2. Neural Turing Machines — Alex Graves, Greg Wayne, Ivo Danihelka, 2014
https://arxiv.org/abs/1410.5401
3. Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context — Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V. Le, Ruslan Salakhutdinov, 2019
https://arxiv.org/abs/1901.02860
4. Compressive Transformers for Long-Range Sequence Modelling — Jack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, Chloe Hillier, Timothy P. Lillicrap, 2020
https://openreview.net/forum?id=SylKikSYDH
5. Memorizing Transformers — Yuhuai Wu, Markus N. Rabe, DeLesley Hutchins, Christian Szegedy, 2022
https://arxiv.org/abs/2203.08913
6. Learning to (learn at test time): RNNs with expressive hidden states — Yu Sun, Xinhao Li, Karan Dalal, Jiarui Xu, Arjun Vikram, Genghan Zhang, Yann Dubois, Xinlei Chen, Xiaolong Wang, and Sanmi Koyejo, 2024
https://scholar.google.com/scholar?q=Learning+to+%28learn+at+test+time%29%3A+RNNs+with+expressive+hidden+states
7. Gated Delta Networks: Improving Mamba2 with Delta Rule — Songlin Yang, Jan Kautz, and Ali Hatamizadeh, 2024
https://scholar.google.com/scholar?q=Gated+Delta+Networks%3A+Improving+Mamba2+with+Delta+Rule
8. RULER: What's the Real Context Size of Your Long-Context Language Models? — Cheng-Ping Hsieh, Simeng Sun, Samuel Kriman, Shantanu Acharya, Dima Rekesh, Fei Jia, and Boris Ginsburg, 2024
https://scholar.google.com/scholar?q=RULER%3A+What%27s+the+Real+Context+Size+of+Your+Long-Context+Language+Models%3F
9. BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack — Yuri Kuratov, Aydar Bulatov, Petr Anokhin, Ivan Rodkin, Dmitry Sorokin, Artyom Sorokin, and Mikhail Burtsev, 2024
https://scholar.google.com/scholar?q=BABILong%3A+Testing+the+Limits+of+LLMs+with+Long+Context+Reasoning-in-a-Haystack
10. ATLAS: Learning to Optimally Memorize the Context at Test Time — Ali Behrouz, Zeman Li, Praneeth Kacham, Majid Daliri, Yuan Deng, Peilin Zhong, Meisam Razaviyayn, and Vahab Mirrokni, 2025
https://scholar.google.com/scholar?q=ATLAS%3A+Learning+to+Optimally+Memorize+the+Context+at+Test+Time
11. KV-Fold: One-Step KV-Cache Recurrence for Long-Context Inference — Alireza Nadali, Patrick Cooper, Ashutosh Trivedi, Alvaro Velasquez, 2026
https://scholar.google.com/scholar?q=KV-Fold%3A+One-Step+KV-Cache+Recurrence+for+Long-Context+Inference
12. SCBench: A KV Cache-Centric Analysis of Long-Context Methods — Yucheng Li, Huiqiang Jiang, Qianhui Wu, Xufang Luo, Surin Ahn, Chengruidong Zhang, Amir H. Abdi, Dongsheng Li, Jianfeng Gao, Yuqing Yang, Lili Qiu, 2024/2025
https://scholar.google.com/scholar?q=SCBench%3A+A+KV+Cache-Centric+Analysis+of+Long-Context+Methods
13. Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language Modeling — Liliang Ren, Yang Liu, Yadong Lu, Yelong Shen, Chen Liang, Weizhu Chen, 2024
https://scholar.google.com/scholar?q=Samba%3A+Simple+Hybrid+State+Space+Models+for+Efficient+Unlimited+Context+Language+Modeling
14. Longhorn: State Space Models are Amortized Online Learners — Bo Liu, Rui Wang, Lemeng Wu, Yihao Feng, Peter Stone, Qiang Liu, 2024
https://scholar.google.com/scholar?q=Longhorn%3A+State+Space+Models+are+Amortized+Online+Learners
15. Retrieval meets Long Context Large Language Models — Peng Xu, Wei Ping, Xianchao Wu, Lawrence McAfee, Chen Zhu, Zihan Liu, Sandeep Subramanian, Evelina Bakhturina, Mohammad Shoeybi, Bryan Catanzaro, 2023
https://scholar.google.com/scholar?q=Retrieval+meets+Long+Context+Large+Language+Models
16. Augmenting Language Models with Long-Term Memory — Weizhi Wang, Li Dong, Hao Cheng, Xiaodong Liu, Xifeng Yan, Jianfeng Gao, Furu Wei, 2023
https://scholar.google.com/scholar?q=Augmenting+Language+Models+with+Long-Term+Memory
17. Test-Time Training Provably Improves Transformers as In-Context Learners — Halil Alperen Gozeten, M. Emrullah Ildiz, Xuechen Zhang, Mahdi Soltanolkotabi, Marco Mondelli, Samet Oymak, 2025
https://scholar.google.com/scholar?q=Test-Time+Training+Provably+Improves+Transformers+as+In-Context+Learners
18. AI Post Transformers: δ-mem and Online Memory for LLMs — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-13-d-mem-and-online-memory-for-llms-6622fa.mp3
19. AI Post Transformers: In-Place Test-Time Training for Transformers — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-09-in-place-test-time-training-for-transfor-d0b976.mp3
20. AI Post Transformers: Gated Linear Attention for Efficient Long Sequences — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-18-gated-linear-attention-for-efficient-lon-c858ab.mp3
21. AI Post Transformers: MELT: Decoupling Compute From Memory — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-13-melt-decoupling-compute-from-memory-26430c.mp3
22. AI Post Transformers: Long Context Pre-Training with Lighthouse Attention — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-13-long-context-pre-training-with-lighthous-e85bbe.mp3
23. AI Post Transformers: Training Million-Token LLMs Beyond the Memory Barrier — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-04-training-million-token-llms-beyond-the-m-324edc.mp3
24. AI Post Transformers: Recursive Language Models for Arbitrarily Long Prompts — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-recursive-language-models-for-arbitraril-fbcd1c.mp3
25. AI Post Transformers: How Induction Heads Emerge in Transformers — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-03-how-induction-heads-emerge-in-transforme-a7bfcb.mp3
26. AI Post Transformers: DeepSeek-V4 and Practical Million-Token Context — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-25-deepseek-v4-and-practical-million-token-6f4de1.mp3
Interactive Visualization: Titans: Learning to Memorize at Test Time

This episode explores the paper’s claim that decoding cost in large language models is driven less by raw parameter counts and more by hardware-level behavior during autoregressive generation, especially memory bandwidth pressure from the KV cache. It explains why metrics like total or activated parameters can be misleading cost proxies, and walks through the tradeoffs among standard attention, grouped-query variants, and newer approaches such as MFA that aim to preserve expressive power while reducing cache overhead. The discussion also highlights the paper’s central systems argument: attention and FFN layers have very different performance bottlenecks, so separating them through Attention-FFN Disaggregation can make large models cheaper to serve without sacrificing capability. A listener would find it interesting for its concrete, skeptical look at why inference efficiency depends on model-system co-design rather than headline model size alone.

Sources:
1. Step-3 is Large yet Affordable: Model-system Co-design for Cost-effective Decoding — StepFun, :, Bin Wang, Bojun Wang, Changyi Wan, Guanzhe Huang, Hanpeng Hu, Haonan Jia, Hao Nie, Mingliang Li, Nuo Chen, Siyu Chen, Song Yuan, Wuxun Xie, Xiaoniu Song, Xing Chen, Xingping Yang, Xuelin Zhang, Yanbo Yu, Yaoyu Wang, Yibo Zhu, Yimin Jiang, Yu Zhou, Yuanwei Lu, Houyi Li, Jingcheng Hu, Ka Man Lo, Ailin Huang, Binxing Jiao, Bo Li, Boyu Chen, Changxin Miao, Chang Lou, Chen Hu, Chen Xu, Chenfeng Yu, Chengyuan Yao, Daokuan Lv, Dapeng Shi, Deshan Sun, Ding Huang, Dingyuan Hu, Dongqing Pang, Enle Liu, Fajie Zhang, Fanqi Wan, Gulin Yan, Han Zhang, Han Zhou, Hanghao Wu, Hangyu Guo, Hanqi Chen, Hanshan Zhang, Hao Wu, Haocheng Zhang, Haolong Yan, Haoran Lv, Haoran Wei, Hebin Zhou, Heng Wang, Heng Wang, Hongxin Li, Hongyu Zhou, Hongyuan Wang, Huiyong Guo, Jia Wang, Jiahao Gong, Jialing Xie, Jian Zhou, Jianjian Sun, Jiaoren Wu, Jiaran Zhang, Jiayu Liu, Jie Cheng, Jie Luo, Jie Yan, Jie Yang, Jieyi Hou, Jinguang Zhang, Jinlan Cao, Jisheng Yin, Junfeng Liu, Junhao Huang, Junzhe Lin, Kaijun Tan, Kaixiang Li, Kang An, Kangheng Lin, Kenkun Liu, Lei Yang, Liang Zhao, Liangyu Chen, Lieyu Shi, Liguo Tan, Lin Lin, Lin Zhang, Lina Chen, Liwen Huang, Liying Shi, Longlong Gu, Mei Chen, Mengqiang Ren, Ming Li, Mingzhe Chen, Na Wang, Nan Wu, Qi Han, Qian Zhao, Qiang Zhang, Qianni Liu, Qiaohui Chen, Qiling Wu, Qinglin He, Qinyuan Tan, Qiufeng Wang, Qiuping Wu, Qiuyan Liang, Quan Sun, Rui Li, Ruihang Miao, Ruosi Wan, Ruyan Guo, Shangwu Zhong, Shaoliang Pang, Shengjie Fan, Shijie Shang, Shilei Jiang, Shiliang Yang, Shiming Hao, Shuli Gao, Siming Huang, Siqi Liu, Tiancheng Cao, Tianhao Cheng, Tianhao Peng, Wang You, Wei Ji, Wen Sun, Wenjin Deng, Wenqing He, Wenzhen Zheng, Xi Chen, Xiangwen Kong, Xianzhen Luo, Xiaobo Yang, Xiaojia Liu, Xiaoxiao Ren, Xin Han, Xin Li, Xin Wu, Xu Zhao, Yanan Wei, Yang Li, Yangguang Li, Yangshijie Xu, Yanming Xu, Yaqiang Shi, Yeqing Shen, Yi Yang, Yifei Yang, Yifeng Gong, Yihan Chen, Yijing Yang, Yinmin Zhang, Yizhuang Zhou, Yuanhao Ding, Yuantao Fan, Yuanzhen Yang, Yuchu Luo, Yue Peng, Yufan Lu, Yuhang Deng, Yuhe Yin, Yujie Liu, Yukun Chen, Yuling Zhao, Yun Mou, Yunlong Li, Yunzhou Ju, Yusheng Li, Yuxiang Yang, Yuxiang Zhang, Yuyang Chen, Zejia Weng, Zhe Xie, Zheng Ge, Zheng Gong, Zhenyi Lu, Zhewei Huang, Zhichao Chang, Zhiguo Huang, Zhirui Wang, Zidong Yang, Zili Wang, Ziqi Wang, Zixin Zhang, Binxing Jiao, Daxin Jiang, Heung-Yeung Shum, Xiangyu Zhang, 2025
http://arxiv.org/abs/2507.19427
2. Fast Transformer Decoding: One Write-Head is All You Need — Noam Shazeer, 2019
https://scholar.google.com/scholar?q=Fast+Transformer+Decoding%3A+One+Write-Head+is+All+You+Need
3. GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints — Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebron, Sumit Sanghai, 2023
https://scholar.google.com/scholar?q=GQA%3A+Training+Generalized+Multi-Query+Transformer+Models+from+Multi-Head+Checkpoints
4. DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model — Zhihong Shao and DeepSeek-AI et al., 2024
https://scholar.google.com/scholar?q=DeepSeek-V2%3A+A+Strong%2C+Economical%2C+and+Efficient+Mixture-of-Experts+Language+Model
5. Multi-matrix Factorization Attention — Jingcheng Hu, Houyi Li, Yinmin Zhang, Zili Wang, Shuigeng Zhou, Xiangyu Zhang, Heung-Yeung Shum, Daxin Jiang, 2024
https://scholar.google.com/scholar?q=Multi-matrix+Factorization+Attention
6. Splitwise: Efficient generative LLM inference using phase splitting — Pratyush Patel, Esha Choukse, Chaojie Zhang, Aashaka Shah, Inigo Goiri, Saeed Maleki, Ricardo Bianchini, 2023
https://scholar.google.com/scholar?q=Splitwise%3A+Efficient+generative+LLM+inference+using+phase+splitting
7. P/D-Serve: Serving Disaggregated Large Language Model at Scale — Yibo Jin, Tao Wang, Huimin Lin and Huawei colleagues, 2024
https://scholar.google.com/scholar?q=P%2FD-Serve%3A+Serving+Disaggregated+Large+Language+Model+at+Scale
8. MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism — Ruidong Zhu, Ziheng Jiang, Chao Jin and ByteDance colleagues, 2025
https://scholar.google.com/scholar?q=MegaScale-Infer%3A+Serving+Mixture-of-Experts+at+Scale+with+Disaggregated+Expert+Parallelism
9. Step-3 is Large yet Affordable: Model-system Co-design for Cost-effective Decoding — StepFun et al., 2025
https://scholar.google.com/scholar?q=Step-3+is+Large+yet+Affordable%3A+Model-system+Co-design+for+Cost-effective+Decoding
10. DeepSeek-V3 Technical Report — DeepSeek-AI et al., 2024
https://scholar.google.com/scholar?q=DeepSeek-V3+Technical+Report
11. Qwen3 MoE 235B — Qwen Team / Alibaba researchers, 2025
https://scholar.google.com/scholar?q=Qwen3+MoE+235B
12. Prefill-Decode Disaggregation — Relevant serving-systems authors cited as [18, 31], 2024-2025
https://scholar.google.com/scholar?q=Prefill-Decode+Disaggregation
13. Kimi K2 Technical Report — Moonshot AI et al., 2025
https://scholar.google.com/scholar?q=Kimi+K2+Technical+Report
14. MiniMax M1 — MiniMax researchers, 2025
https://scholar.google.com/scholar?q=MiniMax+M1
15. KVLink: Accelerating Large Language Models via Efficient KV Cache Reuse — Jingbo Yang et al., 2025
https://scholar.google.com/scholar?q=KVLink%3A+Accelerating+Large+Language+Models+via+Efficient+KV+Cache+Reuse
16. HyperRAG: Enhancing Quality-Efficiency Tradeoffs in Retrieval-Augmented Generation with Reranker KV-Cache Reuse — Yuwei An et al., 2025
https://scholar.google.com/scholar?q=HyperRAG%3A+Enhancing+Quality-Efficiency+Tradeoffs+in+Retrieval-Augmented+Generation+with+Reranker+KV-Cache+Reuse
17. ProphetKV: User-Query-Driven Selective Recomputation for Efficient KV Cache Reuse in Retrieval-Augmented Generation — Shihao Wang et al., 2026
https://scholar.google.com/scholar?q=ProphetKV%3A+User-Query-Driven+Selective+Recomputation+for+Efficient+KV+Cache+Reuse+in+Retrieval-Augmented+Generation
18. HyperAttention: Long-context Attention in Near-Linear Time — Insu Han et al., 2023
https://scholar.google.com/scholar?q=HyperAttention%3A+Long-context+Attention+in+Near-Linear+Time
19. Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention — Tsendsuren Munkhdalai et al., 2024
https://scholar.google.com/scholar?q=Leave+No+Context+Behind%3A+Efficient+Infinite+Context+Transformers+with+Infini-attention
20. Every Attention Matters: An Efficient Hybrid Architecture for Long-Context Reasoning — Ling Team et al., 2025
https://scholar.google.com/scholar?q=Every+Attention+Matters%3A+An+Efficient+Hybrid+Architecture+for+Long-Context+Reasoning
21. KVDirect: Distributed Disaggregated LLM Inference — Shiyang Chen et al., 2024
https://scholar.google.com/scholar?q=KVDirect%3A+Distributed+Disaggregated+LLM+Inference
22. HexGen-2: Disaggregated Generative Inference of LLMs in Heterogeneous Environment — Youhe Jiang et al., 2025
https://scholar.google.com/scholar?q=HexGen-2%3A+Disaggregated+Generative+Inference+of+LLMs+in+Heterogeneous+Environment
23. GRACE-MoE: Grouping and Replication with Locality-Aware Routing for Efficient Distributed MoE Inference — Yu Han et al., 2025
https://scholar.google.com/scholar?q=GRACE-MoE%3A+Grouping+and+Replication+with+Locality-Aware+Routing+for+Efficient+Distributed+MoE+Inference
24. AI Post Transformers: JANUS for Scalable MoE Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-janus-for-scalable-moe-inference-78ae30.mp3
25. AI Post Transformers: Prefill-as-a-Service for Cross-Datacenter KV Cache — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-19-prefill-as-a-service-for-cross-datacente-7560be.mp3
26. AI Post Transformers: Batch-Aware Expert Routing for Faster MoE Decoding — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-batch-aware-expert-routing-for-faster-mo-683ab6.mp3
27. AI Post Transformers: Deep Kernel Fusion for Transformer Decoding — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-deep-kernel-fusion-for-transformer-decod-b1a703.mp3
28. AI Post Transformers: NanoFlow and the Future of LLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-nanoflow-and-the-future-of-llm-serving-7429c9.mp3
29. AI Post Transformers: Why LLM Serving Needs Mathematical Optimization — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-why-llm-serving-needs-mathematical-optim-647fc6.mp3
30. AI Post Transformers: Speculative Decoding in Real vLLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-04-speculative-decoding-in-real-vllm-servin-6f4e2b.mp3
31. AI Post Transformers: Nemotron 3 Super Hybrid Mamba-Transformer MoE — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-19-nemotron-3-super-hybrid-mamba-transforme-31ac75.mp3
32. AI Post Transformers: FlatAttention for Tile-Based Accelerator Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-flatattention-for-tile-based-accelerator-56e6ca.mp3
Interactive Visualization: Affordable Large-Scale Decoding Through Model-System Co-Design

This episode explores MegaScale-Infer, a systems paper on serving large mixture-of-experts language models by separating the attention path from the expert feed-forward path and scheduling them independently. It explains why MoE models can look efficient on paper yet still waste GPU capacity in practice, especially during decode, where KV-cache-heavy attention and uneven expert routing create very different bottlenecks. The discussion focuses on the paper’s core argument for disaggregated expert parallelism and a ping-pong microbatch pipeline designed to keep both attention and expert GPUs busy instead of leaving one side idle. Listeners would find it interesting for its clear look at the gap between model architecture and real-world serving performance, including a pointed debate over whether strong decode benchmarks actually translate into better end-to-end user latency.

Sources:
1. MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism — Ruidong Zhu, Ziheng Jiang, Chao Jin, Peng Wu, Cesar A. Stuardo, Dongyang Wang, Xinlei Zhang, Huaping Zhou, Haoran Wei, Yang Cheng, Jianzhe Xiao, Xinyi Zhang, Lingjun Liu, Haibin Lin, Li-Wen Chang, Jianxi Ye, Xiao Yu, Xuanzhe Liu, Xin Jin, Xin Liu, 2025
http://arxiv.org/abs/2504.02263
2. GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism — Yanping Huang, Yonglong Cheng, Ankur Bapna, Orhan Firat, Dehao Chen, Miaosen Wang, HyoukJoong Lee, Jiquan Ngiam, Quoc V. Le, Zhifeng Chen, 2019
https://scholar.google.com/scholar?q=GPipe%3A+Efficient+Training+of+Giant+Neural+Networks+using+Pipeline+Parallelism
3. PipeDream: Generalized Pipeline Parallelism for DNN Training — Deepak Narayanan, Aaron Harlap, Amar Phanishayee, Vivek Seshadri, Nikhil Devanur, Greg Ganger, Phil Gibbons, Matei Zaharia, 2019
https://scholar.google.com/scholar?q=PipeDream%3A+Generalized+Pipeline+Parallelism+for+DNN+Training
4. Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM — Deepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley, Mostofa Patwary, Nitin Korthikanti, Dmitri Vainbrand, Prethvi Kashinkunti, Julie Bernauer, Bryan Catanzaro, Matei Zaharia, 2021
https://scholar.google.com/scholar?q=Efficient+Large-Scale+Language+Model+Training+on+GPU+Clusters+Using+Megatron-LM
5. Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache — Bin Lin, Chen Zhang, Tao Peng, Hanyu Zhao, Wencong Xiao, Minmin Sun, Anmin Liu, Zhipeng Zhang, Lanbo Li, Xiafei Qiu, Shen Li, Zhigang Ji, Tao Xie, Yong Li, Wei Lin, 2024
https://scholar.google.com/scholar?q=Infinite-LLM%3A+Efficient+LLM+Service+for+Long+Context+with+DistAttention+and+Distributed+KVCache
6. DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving — Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, Hao Zhang, 2024
https://scholar.google.com/scholar?q=DistServe%3A+Disaggregating+Prefill+and+Decoding+for+Goodput-optimized+Large+Language+Model+Serving
7. Splitwise: Efficient Generative LLM Inference using Phase Splitting — Pratyush Patel, Esha Choukse, Chaojie Zhang, Aashaka Shah, Inigo Goiri, Saeed Maleki, Ricardo Bianchini, 2023
https://scholar.google.com/scholar?q=Splitwise%3A+Efficient+Generative+LLM+Inference+using+Phase+Splitting
8. MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs — Shiyi Cao, Shu Liu, Tyler Griggs, Peter Schafhalter, Xiaoxuan Liu, Ying Sheng, Joseph E. Gonzalez, Matei Zaharia, Ion Stoica, 2024
https://scholar.google.com/scholar?q=MoE-Lightning%3A+High-Throughput+MoE+Inference+on+Memory-constrained+GPUs
9. DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model — DeepSeek-AI et al., 2024
https://scholar.google.com/scholar?q=DeepSeek-V2%3A+A+Strong%2C+Economical%2C+and+Efficient+Mixture-of-Experts+Language+Model
10. Toward Efficient Inference for Mixture of Experts — Haiyang Huang, Newsha Ardalani, Anna Sun, Liu Ke, Hsien-Hsin S. Lee, Shruti Bhosale, Carole-Jean Wu, Benjamin Lee, 2024
https://scholar.google.com/scholar?q=Toward+Efficient+Inference+for+Mixture+of+Experts
11. AdapMoE: Adaptive Sensitivity-based Expert Gating and Management for Efficient MoE Inference — Shuzhang Zhong, Ling Liang, Yuan Wang, Runsheng Wang, Ru Huang, Meng Li, 2024
https://scholar.google.com/scholar?q=AdapMoE%3A+Adaptive+Sensitivity-based+Expert+Gating+and+Management+for+Efficient+MoE+Inference
12. HybriMoE: Hybrid CPU-GPU Scheduling and Cache Management for Efficient MoE Inference — Shuzhang Zhong, Yanfan Sun, Ling Liang, Runsheng Wang, Ru Huang, Meng Li, 2025
https://scholar.google.com/scholar?q=HybriMoE%3A+Hybrid+CPU-GPU+Scheduling+and+Cache+Management+for+Efficient+MoE+Inference
13. Oracle-MoE: Locality-preserving Routing in the Oracle Space for Memory-constrained Large Language Model Inference — Jixian Zhou, Fang Dong, Ruijun Huang, Hengjie Cao, Mengyi Chen, Yifeng Yang, Anrui Chen, Mingzhi Dong, Yujiang Wang, Dongsheng Li, David A. Clifton, Qin Lv, Rui Zhu, Chun Zhang, Fan Yang, Tun Lu, Ning Gu, Li Shang, 2025
https://scholar.google.com/scholar?q=Oracle-MoE%3A+Locality-preserving+Routing+in+the+Oracle+Space+for+Memory-constrained+Large+Language+Model+Inference
14. Efficient MoE Inference with Fine-Grained Scheduling of Disaggregated Expert Parallelism — Xinglin Pan, Shaohuai Shi, Wenxiang Lin, Yuxin Wang, Zhenheng Tang, Wei Wang, Xiaowen Chu, 2025
https://scholar.google.com/scholar?q=Efficient+MoE+Inference+with+Fine-Grained+Scheduling+of+Disaggregated+Expert+Parallelism
15. AI Post Transformers: JANUS for Scalable MoE Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-janus-for-scalable-moe-inference-78ae30.mp3
16. AI Post Transformers: Splitwise: Phase-Split LLM Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-26-splitwise-phase-split-llm-inference-e8945b.mp3
17. AI Post Transformers: Prefill-as-a-Service for Cross-Datacenter KV Cache — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-19-prefill-as-a-service-for-cross-datacente-7560be.mp3
18. AI Post Transformers: NanoFlow and the Future of LLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-nanoflow-and-the-future-of-llm-serving-7429c9.mp3
19. AI Post Transformers: Why LLM Serving Needs Mathematical Optimization — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-why-llm-serving-needs-mathematical-optim-647fc6.mp3
20. AI Post Transformers: Batch-Aware Expert Routing for Faster MoE Decoding — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-batch-aware-expert-routing-for-faster-mo-683ab6.mp3
21. AI Post Transformers: LAPS for Length-Aware LLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-laps-for-length-aware-llm-serving-0c6149.mp3
22. AI Post Transformers: Deep Kernel Fusion for Transformer Decoding — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-deep-kernel-fusion-for-transformer-decod-b1a703.mp3
Interactive Visualization: Serving MoE Models with Disaggregated Expert Parallelism

Episode title: The Sparsity Wall: What Reiner Pope Told Dwarkesh About MoE and Sparse Attention A two-paper deep dive framed around the Dwarkesh Patel x Reiner Pope blackboard lecture on training and serving frontier LLMs. The hosts work through "Unified Scaling Laws for Routed Language Models" (Clark et al., DeepMind 2022, arXiv 2202.01169) for the mixture-of-experts side and the DeepSeek sparse-attention paper (arXiv 2512.02556) for the attention side, treating Pope's blackboard framing on the podcast as the pedagogical lens. The episode separates what the papers establish from what Pope's practitioner intuition adds on top, with particular attention to how MoE on the FFN side and sparse attention on the QK side attack independent cost pools and can compound rather than compete.

Sources:
arXiv 2202.01169 — "Unified Scaling Laws for Routed Language Models"
https://arxiv.org/pdf/2202.01169
arXiv 2512.02556 — DeepSeek sparse-attention paper
https://arxiv.org/pdf/2512.02556
Dwarkesh Podcast — "Reiner Pope: The math behind how LLMs are trained and served" (April 29 2026)
https://www.dwarkesh.com/p/reiner-pope
Transcript: https://gist.github.com/dwarkeshsp/79100f0fdeed69d76241903bb0604dbe
Older MoE context: GShard (arXiv 2006.16668), Switch Transformer (arXiv 2101.03961)
Chinchilla scaling laws (arXiv 2203.15556) — referenced in the Pope episode
Interactive Visualization: The Sparsity Wall: What Reiner Pope Told Dwarkesh About MoE and Sparse Attention

This episode explores how a line of recent systems papers culminates in NanoFlow, a serving approach that breaks LLM inference into very small “nano-batches” so different GPU-intensive operations can overlap instead of running in strict sequence. It explains the shift from thinking only about memory bottlenecks such as KV-cache movement and fragmentation toward a more nuanced claim: even if some kernels are memory-bound, overall serving throughput can still be limited by underused compute when prefill and decode are serialized. The discussion walks through the progression from micro-batching and iteration-level scheduling to chunked prefill, then shows how NanoFlow extends that logic with an auto-searched schedule that jointly chooses nano-batch size, operation ordering, and GPU resource allocation. A listener would find it interesting because it frames LLM serving not as a single-kernel optimization problem but as a broader question of hardware utilization, scheduling strategy, and the economics of running large models efficiently at scale.

Interactive Visualization: NanoFlow and the Future of LLM Serving
Sources:
1. NanoFlow and the Future of LLM Serving
https://www.usenix.org/system/files/osdi25-zhu-kan.pdf
2. 2601.11822v1
https://arxiv.org/html/2601.11822v1
3. 2410.18038v2
https://arxiv.org/html/2410.18038v2
4. https://www.usenix.org/system/files/osdi24-agrawal.pdf
https://www.usenix.org/system/files/osdi24-agrawal.pdf
5. 1811.06965
https://arxiv.org/pdf/1811.06965
6. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness — Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, Christopher Re, 2022
https://scholar.google.com/scholar?q=FlashAttention%3A+Fast+and+Memory-Efficient+Exact+Attention+with+IO-Awareness
7. Orca: A Distributed Serving System for Transformer-Based Generative Models — Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim, Byung-Gon Chun, 2022
https://scholar.google.com/scholar?q=Orca%3A+A+Distributed+Serving+System+for+Transformer-Based+Generative+Models
8. Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve — Amey Agrawal, Nitin Kedia, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav Gulavani, Alexey Tumanov, Ramachandran Ramjee, 2024
https://scholar.google.com/scholar?q=Taming+Throughput-Latency+Tradeoff+in+LLM+Inference+with+Sarathi-Serve
9. NanoFlow: Towards Optimal Large Language Model Serving Throughput — Kan Zhu, Yufei Gao, Yilong Zhao, Liangyu Zhao, Gefei Zuo, Yile Gu, Dedong Xie, Zihao Ye, Keisuke Kamahori, Chien-Yu Lin, Ziren Wang, Stephanie Wang, Arvind Krishnamurthy, Baris Kasikci, 2025
https://scholar.google.com/scholar?q=NanoFlow%3A+Towards+Optimal+Large+Language+Model+Serving+Throughput
10. Efficient Memory Management for Large Language Model Serving with PagedAttention — Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, Ion Stoica, 2023
https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention
11. DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving — Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, Hao Zhang, 2024
https://scholar.google.com/scholar?q=DistServe%3A+Disaggregating+Prefill+and+Decoding+for+Goodput-optimized+Large+Language+Model+Serving
12. ASPEN: Breaking Operator Barriers for Efficient Parallelization of Deep Neural Networks — Jongseok Park, Kyungmin Bin, Gibum Park, Sangtae Ha, Kyunghan Lee, 2023
https://scholar.google.com/scholar?q=ASPEN%3A+Breaking+Operator+Barriers+for+Efficient+Parallelization+of+Deep+Neural+Networks
13. vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention — Ramya Prabhu, Ajay Nayak, Jayashree Mohan, Ramachandran Ramjee, Ashish Panwar, 2024
https://scholar.google.com/scholar?q=vAttention%3A+Dynamic+Memory+Management+for+Serving+LLMs+without+PagedAttention
14. MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool — Cunchen Hu et al., 2024
https://scholar.google.com/scholar?q=MemServe%3A+Context+Caching+for+Disaggregated+LLM+Serving+with+Elastic+Memory+Pool
15. DynaServe: Unified and Elastic Execution for Dynamic Disaggregated LLM Serving — Chaoyi Ruan et al., 2025
https://scholar.google.com/scholar?q=DynaServe%3A+Unified+and+Elastic+Execution+for+Dynamic+Disaggregated+LLM+Serving
16. LServe: Efficient Long-sequence LLM Serving with Unified Sparse Attention — Shang Yang et al., 2025
https://scholar.google.com/scholar?q=LServe%3A+Efficient+Long-sequence+LLM+Serving+with+Unified+Sparse+Attention
17. PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference — Dongjie Yang et al., 2024
https://scholar.google.com/scholar?q=PyramidInfer%3A+Pyramid+KV+Cache+Compression+for+High-throughput+LLM+Inference
18. ChunkKV: Semantic-Preserving KV Cache Compression for Efficient Long-Context LLM Inference — Xiang Liu et al., 2025
https://scholar.google.com/scholar?q=ChunkKV%3A+Semantic-Preserving+KV+Cache+Compression+for+Efficient+Long-Context+LLM+Inference
19. Inference-Time Hyper-Scaling with KV Cache Compression — Adrian Łańcucki, Konrad Staniszewski, Piotr Nawrot, Edoardo M. Ponti, 2025
https://scholar.google.com/scholar?q=Inference-Time+Hyper-Scaling+with+KV+Cache+Compression
20. Slice-Level Scheduling for High Throughput and Load Balanced LLM Serving — Ke Cheng et al., 2024
https://scholar.google.com/scholar?q=Slice-Level+Scheduling+for+High+Throughput+and+Load+Balanced+LLM+Serving
21. Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction — Haoran Qiu et al., 2024
https://scholar.google.com/scholar?q=Efficient+Interactive+LLM+Serving+with+Proxy+Model-based+Sequence+Length+Prediction
22. Deferred Continuous Batching in Resource-Efficient Large Language Model Serving — Yongjun He, Yao Lu, Gustavo Alonso, 2024
https://scholar.google.com/scholar?q=Deferred+Continuous+Batching+in+Resource-Efficient+Large+Language+Model+Serving
23. AI Post Transformers: Why LLM Serving Needs Mathematical Optimization — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-why-llm-serving-needs-mathematical-optim-647fc6.mp3
24. AI Post Transformers: LAPS for Length-Aware LLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-laps-for-length-aware-llm-serving-0c6149.mp3
25. AI Post Transformers: Splitwise: Phase-Split LLM Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-26-splitwise-phase-split-llm-inference-e8945b.mp3
26. AI Post Transformers: Continuous Batching for LLM Inference: Throughput and Latency Gains — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/continuous-batching-for-llm-inference-throughput-and-latency-gains/
27. AI Post Transformers: Speculative Decoding in Real vLLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-speculative-decoding-in-real-vllm-servin-6f4e2b.mp3
28. AI Post Transformers: FengHuang for Rack-Scale LLM Inference Memory — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-12-fenghuang-for-rack-scale-llm-inference-m-62708e.mp3
29. AI Post Transformers: CacheFlow and 3D-Parallel KV Cache Restoration — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-01-cacheflow-and-3d-parallel-kv-cache-resto-8db883.mp3
30. AI Post Transformers: ScoutAttention for Efficient KV Cache Offloading — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-24-scoutattention-for-efficient-kv-cache-of-b26699.mp3
31. AI Post Transformers: FlashFuser and Hopper-Era FFN Kernel Fusion — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-flashfuser-and-hopper-era-ffn-kernel-fus-e1fce9.mp3
Interactive Visualization: NanoFlow and the Future of LLM Serving

Episode title: Air Force One, Jensen Huang, and Anthropic's 2028 Memo A close reading of Anthropic's policy post "2028: Two Scenarios for Global AI Leadership" as what it actually is — a carefully timed corporate advocacy document, not a peer-reviewed paper. The episode unpacks the post's central "distillation attacks" framing, distinguishes the four very different things that label gets used for, and weighs Anthropic's policy recommendations against the empirical literature on whether unauthorized knowledge distillation is technically deterrable (citing Trace Rewriting, arXiv 2602.15143, and Watermark Robustness Against Distillation, arXiv 2502.11598). It situates the post in the news cycle of President Trump's May 13–15 2026 state visit to Beijing, the inclusion of Nvidia's Jensen Huang in the delegation, and the H200 clearance for roughly ten Chinese firms — a policy direction that diverges from what the post advocates. Mistral AI's Ministral 3 cascade- distillation work serves as the empirical lens for what compact- model distillation actually transfers in practice. The episode acknowledges legitimate underlying concerns about frontier- capability spread while declining to treat the post as research evidence. Sources (selected; the full citation list will be folded into the script): Anthropic — "2028: Two Scenarios for Global AI Leadership" https://www.anthropic.com/research/2028-ai-leadership Anthropic — "Detecting and preventing distillation attacks" (Feb 2026) https://www.anthropic.com/news/detecting-and-preventing-distillation-attacks Nathan Lambert — "The distillation panic" https://www.interconnects.ai/p/the-distillation-panic TIME — "How A.I. Was the Elephant in the Room at the Trump-Xi Summit" https://time.com/article/2026/05/15/trump-xi-us-china-summit-ai-semiconductor-chips/ Bloomberg — "Nvidia's Huang Joins Trump's China Trip as Last-Minute Addition" https://www.bloomberg.com/news/articles/2026-05-13/nvidia-s-huang-joins-trump-s-china-trip-as-last-minute-addition CNBC — "Trump-Xi summit revives China tech rally hopes as U.S. clears Nvidia H200 sales" https://www.cnbc.com/2026/05/14/trump-xi-meeting-china-stocks-ai-rally.html CFR — "At the Trump-Xi Summit, China Will Have the Upper Hand" https://www.cfr.org/articles/at-the-trump-xi-summit-china-will-have-the-upper-hand CFR — "How Trump Should Approach AI Talks With China" https://www.cfr.org/articles/how-trump-should-approach-ai-talks-with-china-targeted-dialogue-maximum-pressure IAPS — "AI Distillation Attacks: The Case for Targeted Government Intervention" https://www.iaps.ai/research/ai-distillation-attacks Chatham House — "Anthropic's feud with the Pentagon reveals the limits of AI governance" https://www.chathamhouse.org/2026/03/anthropics-feud-pentagon-reveals-limits-ai-governance Small Wars Journal — "Selective Virtue: Anthropic, the Pentagon, and the Contradictions of AI Governance" https://smallwarsjournal.com/2026/04/29/selective-virtue-anthropic-the-pentagon-ai-governance/ arXiv 2602.15143 — "Protecting Language Models Against Unauthorized Distillation through Trace Rewriting" arXiv 2502.11598 — "Can LLM Watermarks Robustly Prevent Unauthorized Knowledge Distillation?" Ministral 3 — discussed in the AI Post Transformers episode "Ministral 3: Cascade Distillation for Long-Context Multimodal Models" AI Post Transformers — "Dario Amodei: Machines of Loving Grace" https://podcast.do-not-panic.com/episodes/dario-amodei-machines-of-loving-grace/ AI Post Transformers — "Dario Amodei: The Adolescence of Technology" https://podcast.do-not-panic.com/episodes/dario-amodei-the-adolescence-of-technology/ AI Post Transformers — "Trace Rewriting Against Unauthorized LLM Distillation" (covers arXiv 2602.15143 / Xinhang Ma et al. WashU, with the watermark-radioactivity literature as comparison)

This episode explores a 2026 paper on defending language models against unauthorized distillation by rewriting chain-of-thought traces before they are returned through an API. It explains the core idea of making reasoning outputs remain useful and correct for human users while becoming less effective as training data for a smaller model trying to copy the teacher, and it connects that strategy to data poisoning and watermarking. The discussion focuses on two defense families, especially LLM-based trace rewriting, and highlights reported results showing strong student degradation on reasoning-heavy tasks like MATH while often preserving or even improving teacher performance. It also digs into the paper’s main ambiguity: whether the defense truly poisons the student’s learning signal, or whether a stronger rewrite model is simply producing cleaner, differently structured reasoning that smaller distilled models fail to absorb well.

Sources:
1. Protecting Language Models Against Unauthorized Distillation through Trace Rewriting — Xinhang Ma, William Yeoh, Ning Zhang, Yevgeniy Vorobeychik, 2026
http://arxiv.org/abs/2602.15143
2. Can LLM Watermarks Robustly Prevent Unauthorized Knowledge Distillation? — Leyi Pan, Aiwei Liu, Shiyu Huang, Yijian Lu, Xuming Hu, Lijie Wen, Irwin King, Philip S. Yu, 2025
http://arxiv.org/abs/2502.11598
3. A Survey of Deep Neural Network Watermarking Techniques — Yusuke Li, Huili Wang, Mauro Barni, 2021
https://scholar.google.com/scholar?q=A+Survey+of+Deep+Neural+Network+Watermarking+Techniques
4. Distillation-Resistant Watermarking for Model Protection in NLP — Xuandong Zhao, Lei Li, Yu-Xiang Wang, 2022
https://scholar.google.com/scholar?q=Distillation-Resistant+Watermarking+for+Model+Protection+in+NLP
5. Protecting Language Generation Models via Invisible Watermarking — Xuandong Zhao, Yu-Xiang Wang, Lei Li, 2023
https://scholar.google.com/scholar?q=Protecting+Language+Generation+Models+via+Invisible+Watermarking
6. Scalable Watermarking for Identifying Large Language Model Outputs — Sumanth Dathathri, Abigail See, S. Ghaisas and colleagues, 2024
https://scholar.google.com/scholar?q=Scalable+Watermarking+for+Identifying+Large+Language+Model+Outputs
7. Poisoning Attacks against Support Vector Machines — Battista Biggio, Blaine Nelson, Pavel Laskov, 2012
https://scholar.google.com/scholar?q=Poisoning+Attacks+against+Support+Vector+Machines
8. Dataset Security for Machine Learning: Data Poisoning, Backdoor Attacks, and Defenses — Micah Goldblum and colleagues, 2020
https://scholar.google.com/scholar?q=Dataset+Security+for+Machine+Learning%3A+Data+Poisoning%2C+Backdoor+Attacks%2C+and+Defenses
9. Online Data Poisoning Attacks — Xuezhou Zhang, Xiaojin Zhu, Laurent Lessard, 2020
https://scholar.google.com/scholar?q=Online+Data+Poisoning+Attacks
10. Poisoning Language Models During Instruction Tuning — Alexander Wan, Eric Wallace, Sheng Shen, Dan Klein, 2023
https://scholar.google.com/scholar?q=Poisoning+Language+Models+During+Instruction+Tuning
11. Adversarial Training Methods for Semi-Supervised Text Classification — Takeru Miyato, Andrew M. Dai, Ian Goodfellow, 2017
https://scholar.google.com/scholar?q=Adversarial+Training+Methods+for+Semi-Supervised+Text+Classification
12. HotFlip: White-Box Adversarial Examples for Text Classification — Javid Ebrahimi, Anyi Rao, Daniel Lowd, Dejing Dou, 2018
https://scholar.google.com/scholar?q=HotFlip%3A+White-Box+Adversarial+Examples+for+Text+Classification
13. Universal Adversarial Triggers for Attacking and Analyzing NLP — Eric Wallace, Shi Feng, Nikhil Kandpal, Matt Gardner, Sameer Singh, 2019
https://scholar.google.com/scholar?q=Universal+Adversarial+Triggers+for+Attacking+and+Analyzing+NLP
14. Bridge the Gap Between CV and NLP! A Gradient-based Textual Adversarial Attack Framework — Lifan Yuan, Yichi Zhang, Yangyi Chen, Wei Wei, 2023
https://scholar.google.com/scholar?q=Bridge+the+Gap+Between+CV+and+NLP%21+A+Gradient-based+Textual+Adversarial+Attack+Framework
15. Antidistillation Sampling — Yash Savani, Asher Trockman, Zhili Feng, Yixuan Even Xu, Avi Schwarzschild, Alexander Robey, Marc Finzi, J. Zico Kolter, 2025
https://scholar.google.com/scholar?q=Antidistillation+Sampling
16. DOGe: Defensive Output Generation for LLM Protection Against Knowledge Distillation — Pingzhi Li, Zhen Tan, Mohan Zhang, Huaizhi Qu, Huan Liu, Tianlong Chen, 2025
https://scholar.google.com/scholar?q=DOGe%3A+Defensive+Output+Generation+for+LLM+Protection+Against+Knowledge+Distillation
17. Information-Preserving Reformulation of Reasoning Traces for Antidistillation — Jiayu Ding, Lei Cui, Li Dong, Nanning Zheng, Furu Wei, 2025
https://scholar.google.com/scholar?q=Information-Preserving+Reformulation+of+Reasoning+Traces+for+Antidistillation
18. Can LLM Watermarks Robustly Prevent Unauthorized Knowledge Distillation? — Leyi Pan, Aiwei Liu, Shiyu Huang, Yijian Lu, Xuming Hu, Lijie Wen, Irwin King, Philip S. Yu, 2025
https://scholar.google.com/scholar?q=Can+LLM+Watermarks+Robustly+Prevent+Unauthorized+Knowledge+Distillation%3F
19. Unified Attacks to Large Language Model Watermarks: Spoofing and Scrubbing in Unauthorized Knowledge Distillation — Xin Yi, Yue Li, Shunfan Zheng, Linlin Wang, Xiaoling Wang, Liang He, 2025
https://scholar.google.com/scholar?q=Unified+Attacks+to+Large+Language+Model+Watermarks%3A+Spoofing+and+Scrubbing+in+Unauthorized+Knowledge+Distillation
20. CATER: Intellectual Property Protection on Text Generation APIs via Conditional Watermarks — Xuanli He, Qiongkai Xu, Yi Zeng, Lingjuan Lyu, Fangzhao Wu, Jiwei Li, Ruoxi Jia, 2022
https://scholar.google.com/scholar?q=CATER%3A+Intellectual+Property+Protection+on+Text+Generation+APIs+via+Conditional+Watermarks
21. DistillGuard: Evaluating Defenses Against LLM Knowledge Distillation — Bo Jiang, 2026
https://scholar.google.com/scholar?q=DistillGuard%3A+Evaluating+Defenses+Against+LLM+Knowledge+Distillation
22. Dialogue Chain-of-Thought Distillation for Commonsense-aware Conversational Agents — Hyungjoo Chae, Yongho Song, Kai Tzu-iunn Ong, Taeyoon Kwon, Minjin Kim, Youngjae Yu, Dongha Lee, Dongyeop Kang, Jinyoung Yeo, 2023
https://scholar.google.com/scholar?q=Dialogue+Chain-of-Thought+Distillation+for+Commonsense-aware+Conversational+Agents
23. Towards Faithful Multi-step Reasoning through Fine-Grained Causal-aware Attribution Reasoning Distillation — Zheng Chu, Jingchang Chen, Zhongjie Wang, Guo Tang, Qianglong Chen, Ming Liu, Bing Qin, 2025
https://scholar.google.com/scholar?q=Towards+Faithful+Multi-step+Reasoning+through+Fine-Grained+Causal-aware+Attribution+Reasoning+Distillation
24. Large Language Model Watermark Stealing With Mixed Integer Programming — Zhaoxi Zhang, Xiaomei Zhang, Yanjun Zhang, Leo Yu Zhang, Chao Chen, Shengshan Hu, Asif Gill, Shirui Pan, 2024
https://scholar.google.com/scholar?q=Large+Language+Model+Watermark+Stealing+With+Mixed+Integer+Programming
25. Distilling the Thought, Watermarking the Answer: A Principle Semantic Guided Watermark for Large Reasoning Models — Shuliang Liu, Xingyu Li, Hongyi Liu, Dong Fang, Yibo Yan, Bingchen Duan, Qi Zheng, Lingfeng Su, Xuming Hu, 2026
https://scholar.google.com/scholar?q=Distilling+the+Thought%2C+Watermarking+the+Answer%3A+A+Principle+Semantic+Guided+Watermark+for+Large+Reasoning+Models
26. AI Post Transformers: Self-Improving Pretraining With Post-Trained Models — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-02-self-improving-pretraining-with-post-tra-e37460.mp3
27. AI Post Transformers: Distilling Multi-Agent Reasoning into a Single LLM — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-22-distilling-multi-agent-reasoning-into-a-143263.mp3
28. AI Post Transformers: AgenticQwen and Small Industrial Tool Agents — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-27-agenticqwen-and-small-industrial-tool-ag-dc676d.mp3
Interactive Visualization: Trace Rewriting Against Unauthorized LLM Distillation

This episode explores Anthropic’s “2028: Two Scenarios for Global AI Leadership” as a strategic argument that advanced AI leadership may hinge on control of compute, semiconductor supply chains, and the ability to slow near-frontier imitation through export controls and distillation defenses. It examines the report’s core claim that the United States and its allies could preserve a 12 to 24 month lead over Chinese labs, while questioning whether ideas like “near-frontier” capability or “model intelligence” are defined well enough to support that kind of forecast. The discussion connects those claims to broader work on the AI triad of compute, data, and algorithms, the geopolitical importance of chip bottlenecks, and the risks of reducing national AI strength to a single score. A listener would find it interesting because it links frontier model development to real policy choices about industrial capacity, national security, and who gets to shape the norms around transformative AI.

Sources:
1. 2028 Scenarios for Global AI Leadership
https://www.anthropic.com/research/2028-ai-leadership
2. The AI Triad and What It Means for National Security Strategy — Ben Buchanan, 2020
https://scholar.google.com/scholar?q=The+AI+Triad+and+What+It+Means+for+National+Security+Strategy
3. Artificial Intelligence with American Values and Chinese Characteristics: A Comparative Analysis of American and Chinese Governmental AI Policies — Emmie Hine, Luciano Floridi, 2022
https://scholar.google.com/scholar?q=Artificial+Intelligence+with+American+Values+and+Chinese+Characteristics%3A+A+Comparative+Analysis+of+American+and+Chinese+Governmental+AI+Policies
4. China's Current Capabilities, Policies, and Industrial Ecosystem in AI — Jeff Ding, 2019
https://scholar.google.com/scholar?q=China%27s+Current+Capabilities%2C+Policies%2C+and+Industrial+Ecosystem+in+AI
5. China's Access to Foreign AI Technology: An Assessment — William Hannas, Huey-Meei Chang, 2019
https://scholar.google.com/scholar?q=China%27s+Access+to+Foreign+AI+Technology%3A+An+Assessment
6. Maintaining China's Dependence on Democracies for Advanced Computer Chips — Saif M. Khan, Carrick Flynn, 2020
https://scholar.google.com/scholar?q=Maintaining+China%27s+Dependence+on+Democracies+for+Advanced+Computer+Chips
7. Assessing the New Semiconductor Export Controls — Matthew Reynolds, 2022
https://scholar.google.com/scholar?q=Assessing+the+New+Semiconductor+Export+Controls
8. Choking off China's Access to the Future of AI — Gregory C. Allen, 2022
https://scholar.google.com/scholar?q=Choking+off+China%27s+Access+to+the+Future+of+AI
9. Hardware-Level Governance of AI Compute: A Feasibility Taxonomy for Regulatory Compliance and Treaty Verification — Samar Ansari, 2026
https://scholar.google.com/scholar?q=Hardware-Level+Governance+of+AI+Compute%3A+A+Feasibility+Taxonomy+for+Regulatory+Compliance+and+Treaty+Verification
10. Global AI governance: barriers and pathways forward — Huw Roberts, Emmie Hine, Mariarosaria Taddeo, Luciano Floridi, 2023
https://scholar.google.com/scholar?q=Global+AI+governance%3A+barriers+and+pathways+forward
11. International governance of civilian AI: a jurisdictional certification approach — Robert Trager, Ben Harack, Anka Reuel, Allison Carnegie, Lennart Heim, Lewis Ho, Sarah Kreps, Ranjit Lall, Owen Larter, Sean O hEigeartaigh, Simon Staffell, Jose Jaime Villalobos, 2023
https://scholar.google.com/scholar?q=International+governance+of+civilian+AI%3A+a+jurisdictional+certification+approach
12. Bridging the Artificial Intelligence Governance Gap: The United States' and China's Divergent Approaches to Governing General-Purpose Artificial Intelligence — Oliver Guest, Kevin Wei, 2025
https://scholar.google.com/scholar?q=Bridging+the+Artificial+Intelligence+Governance+Gap%3A+The+United+States%27+and+China%27s+Divergent+Approaches+to+Governing+General-Purpose+Artificial+Intelligence
13. Concrete Problems in AI Safety — Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, Dan Mane, 2016
https://scholar.google.com/scholar?q=Concrete+Problems+in+AI+Safety
14. Concrete Problems in AI Safety, Revisited — Inioluwa Deborah Raji, Roel Dobbe, 2024
https://scholar.google.com/scholar?q=Concrete+Problems+in+AI+Safety%2C+Revisited
15. An Overview of Catastrophic AI Risks — Dan Hendrycks, Mantas Mazeika, Thomas Woodside, 2023
https://scholar.google.com/scholar?q=An+Overview+of+Catastrophic+AI+Risks
16. US-China perspectives on extreme AI risks and global governance — Akash Wasil, Tim Durgin, 2024
https://scholar.google.com/scholar?q=US-China+perspectives+on+extreme+AI+risks+and+global+governance
17. Machines of Loving Grace — Dario Amodei and Anthropic, 2024
https://scholar.google.com/scholar?q=Machines+of+Loving+Grace
18. The Adolescence of Technology — Anthropic, 2025
https://scholar.google.com/scholar?q=The+Adolescence+of+Technology
19. Chip War: The Fight for the World's Most Critical Technology — Chris Miller, 2022
https://scholar.google.com/scholar?q=Chip+War%3A+The+Fight+for+the+World%27s+Most+Critical+Technology
20. The Export Control and China AI Debate — Gregory C. Allen, 2024
https://scholar.google.com/scholar?q=The+Export+Control+and+China+AI+Debate
21. Tracking AI Compute Trends and the Hardware Bottleneck Literature — Epoch AI and related empirical compute-trend researchers, 2023-2026
https://scholar.google.com/scholar?q=Tracking+AI+Compute+Trends+and+the+Hardware+Bottleneck+Literature
22. Compute-Efficient Deep Learning: Algorithmic Trends and Opportunities — Brian R. Bartoldson, Bhavya Kailkhura, Davis Blalock, 2023
https://scholar.google.com/scholar?q=Compute-Efficient+Deep+Learning%3A+Algorithmic+Trends+and+Opportunities
23. On Efficient Training of Large-Scale Deep Learning Models: A Literature Review — Li Shen, Yan Sun, Zhiyuan Yu, Liang Ding, Xinmei Tian, Dacheng Tao, 2023
https://scholar.google.com/scholar?q=On+Efficient+Training+of+Large-Scale+Deep+Learning+Models%3A+A+Literature+Review
24. Knowledge Distillation and Dataset Distillation of Large Language Models: Emerging Trends, Challenges, and Future Directions — Luyang Fang, Xiaowei Yu, Jiazhang Cai, Yongkai Chen, et al., 2025
https://scholar.google.com/scholar?q=Knowledge+Distillation+and+Dataset+Distillation+of+Large+Language+Models%3A+Emerging+Trends%2C+Challenges%2C+and+Future+Directions
25. PanDa: Prompt Transfer Meets Knowledge Distillation for Efficient Model Adaptation — Qihuang Zhong, Liang Ding, Juhua Liu, Bo Du, Dacheng Tao, 2022
https://scholar.google.com/scholar?q=PanDa%3A+Prompt+Transfer+Meets+Knowledge+Distillation+for+Efficient+Model+Adaptation
26. Tulu 3: Pushing Frontiers in Open Language Model Post-Training — Nathan Lambert, Jacob Morrison, Valentina Pyatkin, Hamish Ivison, et al., 2024
https://scholar.google.com/scholar?q=Tulu+3%3A+Pushing+Frontiers+in+Open+Language+Model+Post-Training
27. AI Diffusion to Low- and Middle Income Countries; A Blessing or a Curse? — Rafael Andersson Lipcsey, 2024
https://scholar.google.com/scholar?q=AI+Diffusion+to+Low-+and+Middle+Income+Countries%3B+A+Blessing+or+a+Curse%3F
28. Exporting the Surveillance State via Trade in AI — Martin Beraja, Andrew Kao, David Y. Yang, Noam Yuchtman, 2023
https://scholar.google.com/scholar?q=Exporting+the+Surveillance+State+via+Trade+in+AI
29. AI Post Transformers: Simple Self-Distillation for Better Code Generation — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-02-simple-self-distillation-for-better-code-cc88e0.mp3
30. AI Post Transformers: Nemotron 3 Super Hybrid Mamba-Transformer MoE — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-19-nemotron-3-super-hybrid-mamba-transforme-31ac75.mp3
31. AI Post Transformers: DeepSeek-V4 and Practical Million-Token Context — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-25-deepseek-v4-and-practical-million-token-6f4de1.mp3
32. AI Post Transformers: Why LLM Serving Needs Mathematical Optimization — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-why-llm-serving-needs-mathematical-optim-647fc6.mp3

This episode explores a position paper arguing that agentic AI systems, built from task decomposition, routing, specialized components, and explicit graph-like workflows, may offer a more credible path to AGI than simply scaling a single monolithic model. It examines how the paper frames AGI through both broad competence across environments and efficient skill acquisition, then asks whether real-world tasks are structured enough for modular systems to outperform one-model-fits-all approaches. The discussion connects that claim to prior work on universal intelligence, compositional generalization, graph-based inductive biases, hierarchical planning, and modular prompting, while stressing that the core debate is about whether intelligence needs external structure rather than just more parameters. A listener would find it interesting for its sharp, theory-driven challenge to the dominant scaling narrative and its concrete attempt to formalize when multi-agent systems should have an advantage.

Interactive Visualization: Agentic AI as a Path to AGI
Sources:
1. Agentic AI as a Path to AGI
https://arxiv.org/pdf/2605.12966
2. HTN Planning: Complexity and Expressivity — Kutluhan Erol, James Hendler, Dana S. Nau, 1994
https://scholar.google.com/scholar?q=HTN+Planning%3A+Complexity+and+Expressivity
3. Hierarchical Reinforcement Learning with the MAXQ Value Function Decomposition — Thomas G. Dietterich, 2000
https://scholar.google.com/scholar?q=Hierarchical+Reinforcement+Learning+with+the+MAXQ+Value+Function+Decomposition
4. Decomposed Prompting: A Modular Approach for Solving Complex Tasks — Tushar Khot, Harsh Trivedi, Matthew Finlayson, Yao Fu, Kyle Richardson, Peter Clark, Ashish Sabharwal, 2022
https://scholar.google.com/scholar?q=Decomposed+Prompting%3A+A+Modular+Approach+for+Solving+Complex+Tasks
5. HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face — Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weiming Lu, Yueting Zhuang, 2023
https://scholar.google.com/scholar?q=HuggingGPT%3A+Solving+AI+Tasks+with+ChatGPT+and+its+Friends+in+Hugging+Face
6. Graph of Thoughts: Solving Elaborate Problems with Large Language Models — Maciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger, Michal Podstawski, Lukas Gianinazzi, Joanna Gajda, Tomasz Lehmann, Hubert Niewiadomski, Piotr Nyczyk, Torsten Hoefler, 2023
https://scholar.google.com/scholar?q=Graph+of+Thoughts%3A+Solving+Elaborate+Problems+with+Large+Language+Models
7. TDAG: A Multi-Agent Framework based on Dynamic Task Decomposition and Agent Generation — Yaoxiang Wang, Zhiyong Wu, Junfeng Yao, Jinsong Su, 2024
https://scholar.google.com/scholar?q=TDAG%3A+A+Multi-Agent+Framework+based+on+Dynamic+Task+Decomposition+and+Agent+Generation
8. DAWN: Distributed LLM Multi-Agent Workflow Synthesis — Guancheng Wan, Mo Zhou, Ziyi Wang, Xiaoran Shang, Eric Hanchen Jiang, Guibin Zhang, Jinhe Bi, Yunpu Ma, Zaixi Zhang, Ke Liang, Wenke Huang, 2026
https://scholar.google.com/scholar?q=DAWN%3A+Distributed+LLM+Multi-Agent+Workflow+Synthesis
9. From Static Templates to Dynamic Runtime Graphs: A Survey of Workflow Optimization for LLM Agents — Ling Yue, Kushal Raj Bhandari, Ching-Yun Ko, Dhaval Patel, Shuxin Lin, Nianjun Zhou, Jianxi Gao, Pin-Yu Chen, Shaowu Pan, 2026
https://scholar.google.com/scholar?q=From+Static+Templates+to+Dynamic+Runtime+Graphs%3A+A+Survey+of+Workflow+Optimization+for+LLM+Agents
10. A Generalist Agent — Scott Reed et al., 2022
https://scholar.google.com/scholar?q=A+Generalist+Agent
11. The Measure of Intelligence — François Chollet, 2019
https://scholar.google.com/scholar?q=The+Measure+of+Intelligence
12. On the Measure of Intelligence — Shane Legg and Marcus Hutter, 2007
https://scholar.google.com/scholar?q=On+the+Measure+of+Intelligence
13. Relational Inductive Biases, Deep Learning, and Graph Networks — Peter W. Battaglia et al., 2018
https://scholar.google.com/scholar?q=Relational+Inductive+Biases%2C+Deep+Learning%2C+and+Graph+Networks
14. No Free Lunch Theorems for Optimization — David H. Wolpert and William G. Macready, 1997
https://scholar.google.com/scholar?q=No+Free+Lunch+Theorems+for+Optimization
15. Scaling can lead to compositional generalization — Florian Redhardt, Yassir Akram, Simon Schug, 2025
https://scholar.google.com/scholar?q=Scaling+can+lead+to+compositional+generalization
16. Single-agent or Multi-agent Systems? Why Not Both? — Mingyan Gao et al., 2025
https://scholar.google.com/scholar?q=Single-agent+or+Multi-agent+Systems%3F+Why+Not+Both%3F
17. When Single-Agent with Skills Replace Multi-Agent Systems and When They Fail — Xiaoxiao Li, 2026
https://scholar.google.com/scholar?q=When+Single-Agent+with+Skills+Replace+Multi-Agent+Systems+and+When+They+Fail
18. Decomposition Dilemmas: Does Claim Decomposition Boost or Burden Fact-Checking Performance? — Qisheng Hu, Quanyu Long, Wenya Wang, 2024/2025
https://scholar.google.com/scholar?q=Decomposition+Dilemmas%3A+Does+Claim+Decomposition+Boost+or+Burden+Fact-Checking+Performance%3F
19. Unifying Language Agent Algorithms with Graph-based Orchestration Engine for Reproducible Agent Research — Qianqian Zhang et al., 2025
https://scholar.google.com/scholar?q=Unifying+Language+Agent+Algorithms+with+Graph-based+Orchestration+Engine+for+Reproducible+Agent+Research
20. AI Post Transformers: ASI-Evolve for Data, Architectures, and RL — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-05-asi-evolve-for-data-architectures-and-rl-197b2b.mp3
21. AI Post Transformers: Kimi K2.5 and Visual Agent Swarms — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-24-kimi-k25-and-visual-agent-swarms-7d04d7.mp3
22. AI Post Transformers: AI Co-Mathematician for Mathematical Research — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-14-ai-co-mathematician-for-mathematical-res-4aa2d4.mp3
23. AI Post Transformers: TMAS: Scaling Test-Time Compute with Multi-Agent Synergy — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-14-tmas-scaling-test-time-compute-with-mult-3abe7a.mp3
24. AI Post Transformers: Agentic Discovery for Test-Time Scaling — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-12-agentic-discovery-for-test-time-scaling-f9a81f.mp3
25. AI Post Transformers: AgenticQwen and Small Industrial Tool Agents — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-27-agenticqwen-and-small-industrial-tool-ag-dc676d.mp3
26. AI Post Transformers: MEMSEARCHER: Reinforcement Learning for LLM Memory Management — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-memsearcher-reinforcement-learning-for-l-e9ad84.mp3
27. AI Post Transformers: Neural Computers as Learned Latent Runtimes — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-11-neural-computers-as-learned-latent-runti-9fa282.mp3
Interactive Visualization: Agentic AI as a Path to AGI

This episode explores a May 13, 2026 arXiv paper arguing that many-shot chain-of-thought prompting can act less like simple retrieval and more like a form of test-time learning, where structured context helps a model reason during inference without changing its weights. It examines the paper’s main claims that adding many reasoning demonstrations does not reliably help across all settings, that semantically similar examples can fail when their reasoning procedures are not actually usable, and that the order of demonstrations becomes more important as prompts grow longer. The discussion also focuses on the paper’s Curvilinear Demonstration Selection approach, which treats prompt construction more like designing a lesson plan than doing nearest-neighbor search, with especially notable gains on geometry tasks. Listeners would find it interesting because it challenges a common assumption behind huge context windows: more examples are not automatically better, and effective prompting may depend on curriculum design, model capabilities, and the compatibility of reasoning traces.

Sources:
1. When Many-Shot CoT Becomes Test-Time Learning
https://arxiv.org/pdf/2605.13511
2. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models — Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, Denny Zhou, 2022
https://arxiv.org/abs/2201.11903
3. Large Language Models are Zero-Shot Reasoners — Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, Yusuke Iwasawa, 2022
https://arxiv.org/abs/2205.11916
4. Self-Consistency Improves Chain of Thought Reasoning in Language Models — Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, Denny Zhou, 2022
https://arxiv.org/abs/2203.11171
5. Automatic Chain of Thought Prompting in Large Language Models — Zhuosheng Zhang, Aston Zhang, Mu Li, Alex Smola, 2022
https://arxiv.org/abs/2210.03493
6. Test-Time Training with Self-Supervision for Generalization under Distribution Shifts — Yu Sun, Xiaolong Wang, Zhuang Liu, John Miller, Alexei A. Efros, Moritz Hardt, 2020
https://arxiv.org/abs/1909.13231
7. What Learning Algorithm is In-Context Learning? Investigations with Linear Models — Ekin Akyurek, Dale Schuurmans, Jacob Andreas, Tengyu Ma, Denny Zhou, 2022
https://arxiv.org/abs/2211.15661
8. Transformers Learn In-Context by Gradient Descent — Johannes von Oswald, Eyvind Niklasson, Ettore Randazzo, Joao Sacramento, Alexander Mordvintsev, Andrey Zhmoginov, Max Vladymyrov, 2022
https://arxiv.org/abs/2212.07677
9. Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters — Charlie Snell, Jaehoon Lee, Kelvin Xu, Aviral Kumar, 2024
https://arxiv.org/abs/2408.03314
10. What Makes Good In-Context Examples for GPT-3? — Jiachang Liu, Dinghan Shen, Yizhe Zhang, Bill Dolan, Lawrence Carin, Weizhu Chen, 2021
https://arxiv.org/abs/2101.06804
11. Learning To Retrieve Prompts for In-Context Learning — Ohad Rubin, Jonathan Herzig, Jonathan Berant, 2021
https://arxiv.org/abs/2112.08633
12. Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity — Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, Pontus Stenetorp, 2021
https://arxiv.org/abs/2104.08786
13. Many-Shot CoT-ICL: Making In-Context Learning Truly Learn — Tsz Ting Chung, Lemao Liu, Mo Yu, Dit-Yan Yeung, 2026
https://arxiv.org/abs/2605.13511
14. Rethinking the Role of Demonstrations: What Makes In-Context Learning Work? — Sewon Min, Mike Lewis, Hannaneh Hajishirzi, Luke Zettlemoyer, 2022
https://scholar.google.com/scholar?q=Rethinking+the+Role+of+Demonstrations%3A+What+Makes+In-Context+Learning+Work%3F
15. Test-Time Compute: Scaling Language Models with More Thinking — OpenAI, 2024
https://scholar.google.com/scholar?q=Test-Time+Compute%3A+Scaling+Language+Models+with+More+Thinking
16. ALR2: A Retrieve-then-Reason Framework for Long-context Question Answering — Huayang Li, Pat Verga, Priyanka Sen, Bowen Yang, Vijay Viswanathan, Patrick Lewis, Taro Watanabe, Yixuan Su, 2024
https://scholar.google.com/scholar?q=ALR2%3A+A+Retrieve-then-Reason+Framework+for+Long-context+Question+Answering
17. Eliciting In-context Retrieval and Reasoning for Long-context Large Language Models — Yifu Qiu, Varun Embar, Yizhe Zhang, Navdeep Jaitly, Shay B. Cohen, Benjamin Han, 2025
https://scholar.google.com/scholar?q=Eliciting+In-context+Retrieval+and+Reasoning+for+Long-context+Large+Language+Models
18. Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions — Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, Ashish Sabharwal, 2023
https://scholar.google.com/scholar?q=Interleaving+Retrieval+with+Chain-of-Thought+Reasoning+for+Knowledge-Intensive+Multi-Step+Questions
19. Active Prompting with Chain-of-Thought for Large Language Models — Shizhe Diao, Pengcheng Wang, Yong Lin, Rui Pan, Xiang Liu, Tong Zhang, 2023
https://scholar.google.com/scholar?q=Active+Prompting+with+Chain-of-Thought+for+Large+Language+Models
20. Synthetic Prompting: Generating Chain-of-Thought Demonstrations for Large Language Models — Zhihong Shao, Yeyun Gong, Yelong Shen, Minlie Huang, Nan Duan, Weizhu Chen, 2023
https://scholar.google.com/scholar?q=Synthetic+Prompting%3A+Generating+Chain-of-Thought+Demonstrations+for+Large+Language+Models
21. Contrastive Chain-of-Thought Prompting — Yew Ken Chia, Guizhen Chen, Luu Anh Tuan, Soujanya Poria, Lidong Bing, 2023
https://scholar.google.com/scholar?q=Contrastive+Chain-of-Thought+Prompting
22. Let's (not) just put things in Context: Test-Time Training for Long-Context LLMs — Rachit Bansal, Aston Zhang, Rishabh Tiwari, Lovish Madaan, Sai Surya Duvvuri, Devvrit Khatri, David Brandfonbrener, David Alvarez-Melis, Prajjwal Bhargava, Mihir Sanjay Kale, Samy Jelassi, 2025
https://scholar.google.com/scholar?q=Let%27s+%28not%29+just+put+things+in+Context%3A+Test-Time+Training+for+Long-Context+LLMs
23. AI Post Transformers: In-Place Test-Time Training for Transformers — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-09-in-place-test-time-training-for-transfor-d0b976.mp3
24. AI Post Transformers: Training LLMs for Divide-and-Conquer Reasoning — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-04-training-llms-for-divide-and-conquer-rea-ea6e22.mp3
25. AI Post Transformers: Reasoning Theater and Unfaithful Chain-of-Thought — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-reasoning-theater-and-unfaithful-chain-o-a4507e.mp3
26. AI Post Transformers: Test-time Scaling for Multi-Agent Collaborative Reasoning — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-22-test-time-scaling-for-multi-agent-collab-082570.mp3
27. AI Post Transformers: δ-mem and Online Memory for LLMs — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-13-d-mem-and-online-memory-for-llms-6622fa.mp3
Interactive Visualization: When Many-Shot CoT Becomes Test-Time Learning

This episode explores Ministral 3, a family of 3B, 8B, and 14B long-context multimodal models built from a 24B parent through structured pruning and cascade distillation rather than separate full-scale training runs. It explains how the method works step by step, from teacher-student distillation and capacity-gap concerns to the staged pruning pipeline that extends each child model to 256k context windows while preserving useful capabilities. The discussion places the paper in context with earlier distillation and pruning work such as Hinton’s original distillation paper, DistilBERT, teacher-assistant distillation, and NVIDIA’s Minitron, arguing that the contribution is a practical model-family construction recipe rather than a brand-new paradigm. Listeners would find it interesting because it gets at a central 2026 question in AI deployment: whether smaller, cheaper models can stay competitive on long-context and multimodal tasks by amortizing one expensive parent run across several deployable descendants.

Sources:
1. Ministral 3 — Alexander H. Liu, Kartik Khandelwal, Sandeep Subramanian, Victor Jouault, Abhinav Rastogi, Adrien Sadé, Alan Jeffares, Albert Jiang, Alexandre Cahill, Alexandre Gavaudan, Alexandre Sablayrolles, Amélie Héliou, Amos You, Andy Ehrenberg, Andy Lo, Anton Eliseev, Antonia Calvi, Avinash Sooriyarachchi, Baptiste Bout, Baptiste Rozière, Baudouin De Monicault, Clémence Lanfranchi, Corentin Barreau, Cyprien Courtot, Daniele Grattarola, Darius Dabert, Diego de las Casas, Elliot Chane-Sane, Faruk Ahmed, Gabrielle Berrada, Gaëtan Ecrepont, Gauthier Guinet, Georgii Novikov, Guillaume Kunsch, Guillaume Lample, Guillaume Martin, Gunshi Gupta, Jan Ludziejewski, Jason Rute, Joachim Studnia, Jonas Amar, Joséphine Delas, Josselin Somerville Roberts, Karmesh Yadav, Khyathi Chandu, Kush Jain, Laurence Aitchison, Laurent Fainsin, Léonard Blier, Lingxiao Zhao, Louis Martin, Lucile Saulnier, Luyu Gao, Maarten Buyl, Margaret Jennings, Marie Pellat, Mark Prins, Mathieu Poirée, Mathilde Guillaumin, Matthieu Dinot, Matthieu Futeral, Maxime Darrin, Maximilian Augustin, Mia Chiquier, Michel Schimpf, Nathan Grinsztajn, Neha Gupta, Nikhil Raghuraman, Olivier Bousquet, Olivier Duchenne, Patricia Wang, Patrick von Platen, Paul Jacob, Paul Wambergue, Paula Kurylowicz, Pavankumar Reddy Muddireddy, Philomène Chagniot, Pierre Stock, Pravesh Agrawal, Quentin Torroba, Romain Sauvestre, Roman Soletskyi, Rupert Menneer, Sagar Vaze, Samuel Barry, Sanchit Gandhi, Siddhant Waghjale, Siddharth Gandhi, Soham Ghosh, Srijan Mishra, Sumukh Aithal, Szymon Antoniak, Teven Le Scao, Théo Cachet, Theo Simon Sorg, Thibaut Lavril, Thiziri Nait Saada, Thomas Chabal, Thomas Foubert, Thomas Robert, Thomas Wang, Tim Lawson, Tom Bewley, Tom Bewley, Tom Edwards, Umar Jamil, Umberto Tomasini, Valeriia Nemychnikova, Van Phung, Vincent Maladière, Virgile Richard, Wassim Bouaziz, Wen-Ding Li, William Marshall, Xinghui Li, Xinyu Yang, Yassine El Ouahidi, Yihan Wang, Yunhao Tang, Zaccharie Ramzi, 2026
http://arxiv.org/abs/2601.08584
2. Distilling the Knowledge in a Neural Network — Geoffrey Hinton, Oriol Vinyals, Jeff Dean, 2015
https://scholar.google.com/scholar?q=Distilling+the+Knowledge+in+a+Neural+Network
3. Improved Knowledge Distillation via Teacher Assistant — Seyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine, Akihiro Matsukawa, Hassan Ghasemzadeh, 2019
https://scholar.google.com/scholar?q=Improved+Knowledge+Distillation+via+Teacher+Assistant
4. DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter — Victor Sanh, Lysandre Debut, Julien Chaumond, Thomas Wolf, 2019
https://scholar.google.com/scholar?q=DistilBERT%2C+a+distilled+version+of+BERT%3A+smaller%2C+faster%2C+cheaper+and+lighter
5. Compact Language Models via Pruning and Knowledge Distillation — Saurav Muralidharan, Sharath Turuvekere Sreenivas, Raviraj Joshi, Marcin Chochowski, Mostofa Patwary, Mohammad Shoeybi, Bryan Catanzaro, Jan Kautz, Pavlo Molchanov, 2024
https://scholar.google.com/scholar?q=Compact+Language+Models+via+Pruning+and+Knowledge+Distillation
6. LLM Pruning and Distillation in Practice: The Minitron Approach — S. T. Sreenivas, S. Muralidharan, R. Joshi, M. Chochowski, A. S. Mahabaleshwarkar, G. Shen, J. Zeng, Z. Chen, Y. Suhara, S. Diao, C. Yu, W. Chen, H. Ross, O. Olabiyi, A. Aithal, O. Kuchaiev, D. Korzekwa, P. Molchanov, M. Patwary, M. Shoeybi, J. Kautz, and B. Catanzaro, 2024
https://scholar.google.com/scholar?q=LLM+Pruning+and+Distillation+in+Practice%3A+The+Minitron+Approach
7. Distillation Scaling Laws — D. Busbridge, A. Shidani, F. Weers, J. Ramapuram, E. Littwin, and R. Webb, 2025
https://scholar.google.com/scholar?q=Distillation+Scaling+Laws
8. Distilled Pretraining: A Modern Lens of Data, In-Context Learning and Test-Time Scaling — S. Goyal, D. Lopez-Paz, and K. Ahuja, 2025
https://scholar.google.com/scholar?q=Distilled+Pretraining%3A+A+Modern+Lens+of+Data%2C+In-Context+Learning+and+Test-Time+Scaling
9. Pixtral 12B — P. Agrawal, S. Antoniak, E. B. Hanna, B. Bout, D. Chaplot, J. Chudnovsky, D. Costa, B. De Monicault, S. Garg, T. Gervet, et al., 2024
https://scholar.google.com/scholar?q=Pixtral+12B
10. Junk DNA Hypothesis: Pruning Small Pre-Trained Weights Irreversibly and Monotonically Impairs "Difficult" Downstream Tasks in LLMs — Lu Yin et al., 2024
https://scholar.google.com/scholar?q=Junk+DNA+Hypothesis%3A+Pruning+Small+Pre-Trained+Weights+Irreversibly+and+Monotonically+Impairs+%22Difficult%22+Downstream+Tasks+in+LLMs
11. SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot — Elias Frantar and Dan Alistarh, 2023
https://scholar.google.com/scholar?q=SparseGPT%3A+Massive+Language+Models+Can+Be+Accurately+Pruned+in+One-Shot
12. Fast and Effective Weight Update for Pruned Large Language Models — Vladimir Boza, 2024
https://scholar.google.com/scholar?q=Fast+and+Effective+Weight+Update+for+Pruned+Large+Language+Models
13. Exploring Knowledge Purification in Multi-Teacher Knowledge Distillation for LLMs — Ruihan Jin et al., 2026
https://scholar.google.com/scholar?q=Exploring+Knowledge+Purification+in+Multi-Teacher+Knowledge+Distillation+for+LLMs
14. Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models — Siyan Zhao et al., 2026
https://scholar.google.com/scholar?q=Self-Distilled+Reasoner%3A+On-Policy+Self-Distillation+for+Large+Language+Models
15. Data Engineering for Scaling Language Models to 128K Context — Yao Fu et al., 2024
https://scholar.google.com/scholar?q=Data+Engineering+for+Scaling+Language+Models+to+128K+Context
16. How to Train Long-Context Language Models (Effectively) — Tianyu Gao et al., 2025
https://scholar.google.com/scholar?q=How+to+Train+Long-Context+Language+Models+%28Effectively%29
17. Train Small, Infer Large: Memory-Efficient LoRA Training for Large Language Models — Jun Zhang et al., 2025
https://scholar.google.com/scholar?q=Train+Small%2C+Infer+Large%3A+Memory-Efficient+LoRA+Training+for+Large+Language+Models
18. AI Post Transformers: DeepSeek-V4 and Practical Million-Token Context — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-25-deepseek-v4-and-practical-million-token-6f4de1.mp3
19. AI Post Transformers: Muon Is Scalable for LLM Training — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-25-muon-is-scalable-for-llm-training-587ed8.mp3
20. AI Post Transformers: Learning to Reason with 13 Parameters — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-14-learning-to-reason-with-13-parameters-54c87f.mp3
21. AI Post Transformers: AgenticQwen and Small Industrial Tool Agents — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-27-agenticqwen-and-small-industrial-tool-ag-dc676d.mp3
Interactive Visualization: Ministral 3: Cascade Distillation for Long-Context Multimodal Models

This episode explores Causal-JEPA, a world-modeling approach that masks whole object trajectories rather than image patches to force a model to reason about interactions between entities. It explains how the method combines object-centric representations with JEPA-style latent prediction, asking the model to reconstruct hidden objects from scene context and then predict future dynamics, instead of relying on pixel reconstruction or simple autoregressive rollouts. The discussion highlights the paper’s core argument that this training setup makes counterfactual and causal reasoning more necessary by blocking shortcut strategies like temporal interpolation and self-contained single-object motion prediction. Listeners would find it interesting for its sharp comparison between patch-based scaling and object-centric structure, and for its claim that better world models may come from making interaction reasoning unavoidable rather than merely possible.

Sources:
1. Causal-JEPA: Learning World Models through Object-Level Latent Interventions — Heejeong Nam, Quentin Le Lidec, Lucas Maes, Yann LeCun, Randall Balestriero, 2026
http://arxiv.org/abs/2602.11389
2. MONet: Unsupervised Scene Decomposition and Representation — Christopher P. Burgess, Loic Matthey, Nicholas Watters, Rishabh Kabra, Irina Higgins, Matt Botvinick, Alexander Lerchner, 2019
https://scholar.google.com/scholar?q=MONet%3A+Unsupervised+Scene+Decomposition+and+Representation
3. Multi-Object Representation Learning with Iterative Variational Inference — Klaus Greff, Raphael Lopez Kaufman, Rishabh Kabra, Nick Watters, Chris Burgess, Daniel Zoran, Loic Matthey, Matthew Botvinick, Alexander Lerchner, 2019
https://scholar.google.com/scholar?q=Multi-Object+Representation+Learning+with+Iterative+Variational+Inference
4. Object-Centric Learning with Slot Attention — Francesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran, Georg Heigold, Jakob Uszkoreit, Alexey Dosovitskiy, Thomas Kipf, 2020
https://scholar.google.com/scholar?q=Object-Centric+Learning+with+Slot+Attention
5. Bridging the Gap to Real-World Object-Centric Learning — Maximilian Seitzer, Max Horn, Andrii Zadaianchuk, Dominik Zietlow, Tianjun Xiao, Carl-Johann Simon-Gabriel, Tong He, Zheng Zhang, Bernhard Schölkopf, Thomas Brox, Francesco Locatello, 2023
https://scholar.google.com/scholar?q=Bridging+the+Gap+to+Real-World+Object-Centric+Learning
6. A Path Towards Autonomous Machine Intelligence — Yann LeCun, 2022
https://scholar.google.com/scholar?q=A+Path+Towards+Autonomous+Machine+Intelligence
7. Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture — Mahmoud Assran, Quentin Duval, Ishan Misra, Piotr Bojanowski, Pascal Vincent, Michael Rabbat, Yann LeCun, Nicolas Ballas, 2023
https://scholar.google.com/scholar?q=Self-Supervised+Learning+from+Images+with+a+Joint-Embedding+Predictive+Architecture
8. Revisiting Feature Prediction for Learning Visual Representations from Video — Adrien Bardes, Quentin Garrido, Jean Ponce, Xinlei Chen, Michael Rabbat, Yann LeCun, Mahmoud Assran, Nicolas Ballas, 2024
https://scholar.google.com/scholar?q=Revisiting+Feature+Prediction+for+Learning+Visual+Representations+from+Video
9. V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning — Mido Assran, Adrien Bardes, David Fan and many others including Yann LeCun, Michael Rabbat, Nicolas Ballas, 2025
https://scholar.google.com/scholar?q=V-JEPA+2%3A+Self-Supervised+Video+Models+Enable+Understanding%2C+Prediction+and+Planning
10. CLEVRER: CoLlision Events for Video REpresentation and Reasoning — Kexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli, Jiajun Wu, Antonio Torralba, Joshua B. Tenenbaum, 2020
https://scholar.google.com/scholar?q=CLEVRER%3A+CoLlision+Events+for+Video+REpresentation+and+Reasoning
11. Counterfactual VQA: A Cause-Effect Look at Language Bias — Yulei Niu, Kaihua Tang, Hanwang Zhang, Zhiwu Lu, Xian-Sheng Hua, Ji-Rong Wen, 2021
https://scholar.google.com/scholar?q=Counterfactual+VQA%3A+A+Cause-Effect+Look+at+Language+Bias
12. What If the TV Was Off? Examining Counterfactual Reasoning Abilities of Multi-modal Language Models — Letian Zhang, Xiaotong Zhai, Zhongkai Zhao, Yongshuo Zong, Xin Wen, Bingchen Zhao, 2024
https://scholar.google.com/scholar?q=What+If+the+TV+Was+Off%3F+Examining+Counterfactual+Reasoning+Abilities+of+Multi-modal+Language+Models
13. ACQUIRED: A Dataset for Answering Counterfactual Questions In Real-Life Videos — Te-Lin Wu, Zi-Yi Dou, Qingyuan Hu, Yu Hou, Nischal Reddy Chandra, Marjorie Freedman, Ralph M. Weischedel, Nanyun Peng, 2023
https://scholar.google.com/scholar?q=ACQUIRED%3A+A+Dataset+for+Answering+Counterfactual+Questions+In+Real-Life+Videos
14. Towards Causal Representation Learning — Bernhard Schölkopf, Francesco Locatello, Stefan Bauer, Nan Rosemary Ke, Nal Kalchbrenner, Anirudh Goyal, Yoshua Bengio, 2021
https://scholar.google.com/scholar?q=Towards+Causal+Representation+Learning
15. Interventional Causal Representation Learning — Kartik Ahuja, Divyat Mahajan, Yixin Wang, Yoshua Bengio, 2023
https://scholar.google.com/scholar?q=Interventional+Causal+Representation+Learning
16. Desiderata for Representation Learning: A Causal Perspective — Yixin Wang, Michael I. Jordan, 2024
https://scholar.google.com/scholar?q=Desiderata+for+Representation+Learning%3A+A+Causal+Perspective
17. Provably Learning Object-Centric Representations — Stefan Bauer, Bernhard Schölkopf and collaborators, 2023
https://scholar.google.com/scholar?q=Provably+Learning+Object-Centric+Representations
18. SlotFormer: Unsupervised Visual Dynamics Simulation with Object-Centric Models — Yuhang Wu, Yueting Zhuang, Francesco Locatello, et al., 2022
https://scholar.google.com/scholar?q=SlotFormer%3A+Unsupervised+Visual+Dynamics+Simulation+with+Object-Centric+Models
19. Object-Centric Video Prediction via Decoupling of Object Dynamics and Interactions — Angel Villar-Corrales, Ismail Wahdan, Sven Behnke, 2023
https://scholar.google.com/scholar?q=Object-Centric+Video+Prediction+via+Decoupling+of+Object+Dynamics+and+Interactions
20. DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning — Gaoyue Zhou, Hengkai Pan, Yann LeCun, Lerrel Pinto, 2024
https://scholar.google.com/scholar?q=DINO-WM%3A+World+Models+on+Pre-trained+Visual+Features+enable+Zero-shot+Planning
21. Conditional Object-Centric Learning from Video — Thomas Kipf, Gamaleldin F. Elsayed, Aravindh Mahendran, Austin Stone, Sara Sabour, Georg Heigold, Rico Jonschkowski, Alexey Dosovitskiy, Klaus Greff, 2022
https://scholar.google.com/scholar?q=Conditional+Object-Centric+Learning+from+Video
22. Attention over Learned Object Embeddings Enables Complex Visual Reasoning — David Ding, Felix Hill, Adam Santoro, Malcolm Reynolds, Matt Botvinick, 2021
https://scholar.google.com/scholar?q=Attention+over+Learned+Object+Embeddings+Enables+Complex+Visual+Reasoning
23. Dyn-O: Building Structured World Models with Object-Centric Representations — Zizhao Wang, Kaixin Wang, Li Zhao, Peter Stone, Jiang Bian, 2025
https://scholar.google.com/scholar?q=Dyn-O%3A+Building+Structured+World+Models+with+Object-Centric+Representations
24. Learning Interactive World Model for Object-Centric Reinforcement Learning — Fan Feng, Phillip Lippe, Sara Magliacane, 2025
https://scholar.google.com/scholar?q=Learning+Interactive+World+Model+for+Object-Centric+Reinforcement+Learning
25. Object-Centric World Model for Language-Guided Manipulation — Youngjoon Jeong, Junha Chun, Soonwoo Cha, Taesup Kim, 2025
https://scholar.google.com/scholar?q=Object-Centric+World+Model+for+Language-Guided+Manipulation
26. Planning in 8 Tokens: A Compact Discrete Tokenizer for Latent World Model — Dongwon Kim, Gawon Seo, Jinsung Lee, Minsu Cho, Suha Kwak, 2026
https://scholar.google.com/scholar?q=Planning+in+8+Tokens%3A+A+Compact+Discrete+Tokenizer+for+Latent+World+Model
27. Learning nonparametric latent causal graphs with unknown interventions — Yibo Jiang, Bryon Aragam, 2023
https://scholar.google.com/scholar?q=Learning+nonparametric+latent+causal+graphs+with+unknown+interventions
28. Learning Linear Causal Representations from Interventions under General Nonlinear Mixing — Simon Buchholz, Goutham Rajendran, Elan Rosenfeld, Bryon Aragam, Bernhard Schölkopf, Pradeep Ravikumar, 2023
https://scholar.google.com/scholar?q=Learning+Linear+Causal+Representations+from+Interventions+under+General+Nonlinear+Mixing
29. AI Post Transformers: LeWorldModel: Stable Joint-Embedding World Models from Pixels — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-25-leworldmodel-stable-joint-embedding-worl-650f9f.mp3
30. AI Post Transformers: Learning Latent Action World Models from Video — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-09-learning-latent-action-world-models-from-1570a4.mp3
31. AI Post Transformers: DreamerV3 World Models Across 150 Tasks — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-20-dreamerv3-world-models-across-150-tasks-af5edb.mp3
Interactive Visualization: Causal-JEPA for Object-Level World Models

This episode explores whether code models should allocate training data evenly across programming languages or tune the mix based on how each language scales. It walks through the paper’s core experiments on monolingual scaling, bilingual transfer, translation-style code pairing, and multilingual token allocation, with close attention to languages like Python, JavaScript, Rust, and TypeScript. The discussion highlights the paper’s main argument that programming languages differ in scaling behavior and cross-lingual usefulness, which could make non-uniform token budgets more effective than treating all code equally. Listeners would find it interesting for its concrete take on how multilingual software ecosystems challenge simple scaling-law assumptions and for its careful scrutiny of whether the evidence really supports those stronger optimization claims.

Sources:
1. Scaling Laws for Multilingual Code Pretraining
https://arxiv.org/pdf/2512.13472
2. Unsupervised Translation of Programming Languages — Marie-Anne Lachaux, Baptiste Roziere, Lowik Chanussot, Guillaume Lample, 2020
https://scholar.google.com/scholar?q=Unsupervised+Translation+of+Programming+Languages
3. XLCoST: A Benchmark Dataset for Cross-lingual Code Intelligence — Ming Zhu, Aneesh Jain, Karthik Suresh, Roshan Ravindran, Sindhu Tipirneni, Chandan K. Reddy, 2022
https://scholar.google.com/scholar?q=XLCoST%3A+A+Benchmark+Dataset+for+Cross-lingual+Code+Intelligence
4. MultiPL-E: A Scalable and Extensible Approach to Benchmarking Neural Code Generation — Federico Cassano, John Gouwar, Daniel Nguyen, Sydney Nguyen, Luna Phipps-Costin, Donald Pinckney, Ming-Ho Yee, Yangtian Zi, Carolyn Jane Anderson, Molly Q. Feldman, Arjun Guha, Michael Greenberg, Abhinav Jangda, 2022
https://scholar.google.com/scholar?q=MultiPL-E%3A+A+Scalable+and+Extensible+Approach+to+Benchmarking+Neural+Code+Generation
5. CodeGeeX: A Pre-Trained Model for Code Generation with Multilingual Benchmarking on HumanEval-X — Qinkai Zheng, Xiao Xia, Xu Zou, Yuxiao Dong, Shan Wang, Yufei Xue, Zihan Wang, Lei Shen, Andi Wang, Yang Li, Teng Su, Zhilin Yang, Jie Tang, 2023
https://scholar.google.com/scholar?q=CodeGeeX%3A+A+Pre-Trained+Model+for+Code+Generation+with+Multilingual+Benchmarking+on+HumanEval-X
6. Scaling Laws for Code: A More Data-Hungry Regime — Xianzhen Luo, Wenzhen Zheng, Qingfu Zhu, Rongyi Zhang, Houyi Li, Siming Huang, YuanTao Fan, Wanxiang Che, 2025
https://scholar.google.com/scholar?q=Scaling+Laws+for+Code%3A+A+More+Data-Hungry+Regime
7. Cross-lingual Transfer in Programming Languages: An Extensive Empirical Study — Razan Baltaji, Saurabh Pujar, Louis Mandel, Martin Hirzel, Luca Buratti, Lav Varshney, 2025
https://scholar.google.com/scholar?q=Cross-lingual+Transfer+in+Programming+Languages%3A+An+Extensive+Empirical+Study
8. CRUXEval-X: A Benchmark for Multilingual Code Reasoning, Understanding and Execution — Ruiyang Xu, Jialun Cao, Yaojie Lu, Hongyu Lin, Xianpei Han, Ben He, Shing-Chi Cheung, Le Sun, 2024
https://scholar.google.com/scholar?q=CRUXEval-X%3A+A+Benchmark+for+Multilingual+Code+Reasoning%2C+Understanding+and+Execution
9. RepoTransBench: A Real-World Benchmark for Repository-Level Code Translation — Yanli Wang, Yanlin Wang, Suiquan Wang, Daya Guo, Jiachi Chen, John Grundy, Xilin Liu, Yuchi Ma, Mingzhi Mao, Hongyu Zhang, Zibin Zheng, 2024
https://scholar.google.com/scholar?q=RepoTransBench%3A+A+Real-World+Benchmark+for+Repository-Level+Code+Translation
10. Towards Multi-Language Repository-Level Code Generation: From-Scratch to Guided Tasks — authors unclear from snippet, unknown
https://scholar.google.com/scholar?q=Towards+Multi-Language+Repository-Level+Code+Generation%3A+From-Scratch+to+Guided+Tasks
11. Do Not Treat Code as Natural Language: Implications for Repository-Level Code Generation and Beyond — authors unclear from snippet, unknown
https://scholar.google.com/scholar?q=Do+Not+Treat+Code+as+Natural+Language%3A+Implications+for+Repository-Level+Code+Generation+and+Beyond
12. Code-Switching In-Context Learning for Cross-Lingual Transfer of Large Language Models — authors unclear from snippet, unknown
https://scholar.google.com/scholar?q=Code-Switching+In-Context+Learning+for+Cross-Lingual+Transfer+of+Large+Language+Models
13. Bridging the Language Gap: Enhancing Multilingual Prompt-Based Code Generation in LLMs via Zero-Shot Cross-Lingual Transfer — authors unclear from snippet, unknown
https://scholar.google.com/scholar?q=Bridging+the+Language+Gap%3A+Enhancing+Multilingual+Prompt-Based+Code+Generation+in+LLMs+via+Zero-Shot+Cross-Lingual+Transfer
14. Advancing Code Translation With Context-Aware Pre-Training in Data-Scarce Environments — authors unclear from snippet, unknown
https://scholar.google.com/scholar?q=Advancing+Code+Translation+With+Context-Aware+Pre-Training+in+Data-Scarce+Environments
15. Semi-Supervised Code Translation Overcoming the Scarcity of Parallel Code Data — authors unclear from snippet, unknown
https://scholar.google.com/scholar?q=Semi-Supervised+Code+Translation+Overcoming+the+Scarcity+of+Parallel+Code+Data
16. Scaling Laws for Predicting Downstream Performance in LLMs — authors unclear from snippet, unknown
https://scholar.google.com/scholar?q=Scaling+Laws+for+Predicting+Downstream+Performance+in+LLMs
17. AI Post Transformers: CODEGEN: Open Language Model for Code Synthesis — Hal Turing & Dr. Ada Shannon, Fri,
https://podcast.do-not-panic.com/episodes/codegen-open-language-model-for-code-synthesis/
18. AI Post Transformers: Llama 3: Architecture, Capabilities, and Safety — Hal Turing & Dr. Ada Shannon, Sun,
https://podcast.do-not-panic.com/episodes/llama-3-architecture-capabilities-and-safety/
19. AI Post Transformers: Program Synthesis with Large Language Models — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-20-program-synthesis-with-large-language-mo-b962ec.mp3
20. AI Post Transformers: MTEB & MMTEB: The Massive Text Embedding Benchmark — Hal Turing & Dr. Ada Shannon, Fri,
https://podcast.do-not-panic.com/episodes/mteb-mmteb-the-massive-text-embedding-benchmark/
Interactive Visualization: Scaling Laws for Multilingual Code Pretraining

This episode explores JANUS, a systems approach to serving mixture-of-experts transformers efficiently by separating attention layers from expert layers instead of deploying the whole model as a single monolithic unit. It explains why MoE models can still be expensive and latency-prone in practice: even if only a few experts activate per token, the system must still manage large expert memory footprints, skewed expert demand, and strict token-level latency targets such as time per output token. The discussion focuses on JANUS’s core ideas, including separate GPU pools for attention and expert computation, an adaptive two-phase communication scheme that reduces cross-node messaging overhead, and SLO-aware scaling that adjusts attention and expert capacity independently. Listeners would find it interesting because it turns MoE inference from a simple “sparse compute saves money” story into a deeper argument about distributed systems design, load balancing, and the real bottlenecks that determine whether advanced models feel fast in production.

Interactive Visualization: JANUS for Scalable MoE Inference
Sources:
1. Janus: Disaggregating Attention and Experts for Scalable MoE Inference — Zhexiang Zhang, Ye Wang, Yumiao Zhao, Jiayu Xiao, Qianjing Yang, Xiangyu Wang, Jingzhe Jiang, Qizhen Weng, Ruichuan Chen, Shaohuai Shi, Adel N. Toosi, Yin Chen, Minchen Yu, 2025
http://arxiv.org/abs/2512.13525
2. Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity — William Fedus, Barret Zoph, Noam Shazeer, 2021
https://scholar.google.com/scholar?q=Switch+Transformers%3A+Scaling+to+Trillion+Parameter+Models+with+Simple+and+Efficient+Sparsity
3. FastMoE: A Fast Mixture-of-Expert Training System — Jiaao He, Jiezhong Qiu, Aohan Zeng, Zhilin Yang, Jidong Zhai, Jie Tang, 2021
https://scholar.google.com/scholar?q=FastMoE%3A+A+Fast+Mixture-of-Expert+Training+System
4. DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving — Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, Hao Zhang, 2024
https://scholar.google.com/scholar?q=DistServe%3A+Disaggregating+Prefill+and+Decoding+for+Goodput-optimized+Large+Language+Model+Serving
5. MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism — Ruidong Zhu, Ziheng Jiang, Chao Jin, Peng Wu, Cesar A. Stuardo, Dongyang Wang and others, 2025
https://scholar.google.com/scholar?q=MegaScale-Infer%3A+Serving+Mixture-of-Experts+at+Scale+with+Disaggregated+Expert+Parallelism
6. eMoE: Task-aware Memory Efficient Mixture-of-Experts-Based (MoE) Model Inference — Suraiya Tairin, Shohaib Mahmud, Haiying Shen, Anand Iyer, 2025
https://scholar.google.com/scholar?q=eMoE%3A+Task-aware+Memory+Efficient+Mixture-of-Experts-Based+%28MoE%29+Model+Inference
7. SLO-Aware Compute Resource Allocation for Prefill-Decode Disaggregated LLM Inference — Luchang Li, Dongfang Li, Bozhao Gong, Yu Zhang, 2026
https://scholar.google.com/scholar?q=SLO-Aware+Compute+Resource+Allocation+for+Prefill-Decode+Disaggregated+LLM+Inference
8. MoEless: Efficient MoE LLM Serving via Serverless Computing — Hanfei Yu, Bei Ouyang, Shwai He, Ang Li, Hao Wang, 2026
https://scholar.google.com/scholar?q=MoEless%3A+Efficient+MoE+LLM+Serving+via+Serverless+Computing
9. xDeepServe: Model-as-a-Service on Huawei CloudMatrix384 — Ao Xiao et al., 2025
https://scholar.google.com/scholar?q=xDeepServe%3A+Model-as-a-Service+on+Huawei+CloudMatrix384
10. Semantic Parallelism: Redefining Efficient MoE Inference via Model-Data Co-Scheduling — Yan Li, Zhenyu Zhang, Zhengang Wang, Pengfei Chen, Pengfei Zheng, 2025
https://scholar.google.com/scholar?q=Semantic+Parallelism%3A+Redefining+Efficient+MoE+Inference+via+Model-Data+Co-Scheduling
11. GRACE-MoE: Grouping and Replication with Locality-Aware Routing for Efficient Distributed MoE Inference — Yu Han, Lehan Pan, Jie Peng, Ziyang Tao, Hanqi Zhu, Wuyang Zhang, Yanyong Zhang, 2025
https://scholar.google.com/scholar?q=GRACE-MoE%3A+Grouping+and+Replication+with+Locality-Aware+Routing+for+Efficient+Distributed+MoE+Inference
12. BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems — Yuxin Wang, Yuhan Chen, Zeyu Li, Xueze Kang, Yuchu Fang, Yeju Zhou, Yang Zheng, Zhenheng Tang, Xin He, Rui Guo, Xin Wang, Qiang Wang, Amelie Chi Zhou, Xiaowen Chu, 2024
https://scholar.google.com/scholar?q=BurstGPT%3A+A+Real-world+Workload+Dataset+to+Optimize+LLM+Serving+Systems
13. DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model — DeepSeek-AI et al., 2024
https://scholar.google.com/scholar?q=DeepSeek-V2%3A+A+Strong%2C+Economical%2C+and+Efficient+Mixture-of-Experts+Language+Model
14. Pre-gated MoE: An Algorithm-System Co-Design for Fast and Scalable Mixture-of-Expert Inference — Ranggi Hwang et al., 2023/2024
https://arxiv.org/abs/2308.12066
15. HOBBIT: A Mixed Precision Expert Offloading System for Fast MoE Inference — Peng Tang et al., 2024
https://arxiv.org/abs/2411.01433
16. DAOP: Data-Aware Offloading and Predictive Pre-Calculation for Efficient MoE Inference — Yujie Zhang, Shivam Aggarwal, Tulika Mitra, 2025
https://arxiv.org/abs/2501.10375
17. AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding — Zikun Li et al., 2025
https://arxiv.org/abs/2501.12162
18. SLOs-Serve: Optimized Serving of Multi-SLO LLMs — Siyuan Chen et al., 2025
https://arxiv.org/abs/2504.08784
19. Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement — Tian Wu et al., 2025
https://arxiv.org/abs/2508.12851
20. AI Post Transformers: Batch-Aware Expert Routing for Faster MoE Decoding — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-batch-aware-expert-routing-for-faster-mo-683ab6.mp3
21. AI Post Transformers: Splitwise: Phase-Split LLM Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-26-splitwise-phase-split-llm-inference-e8945b.mp3
22. AI Post Transformers: Computation-Bandwidth-Memory Trade-offs for AI Infrastructure — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-09-computation-bandwidth-memory-trade-offs-a83f2b.mp3
23. AI Post Transformers: Why LLM Serving Needs Mathematical Optimization — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-why-llm-serving-needs-mathematical-optim-647fc6.mp3
24. AI Post Transformers: LAPS for Length-Aware LLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-laps-for-length-aware-llm-serving-0c6149.mp3
25. AI Post Transformers: DeepSeek-V4 and Practical Million-Token Context — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-25-deepseek-v4-and-practical-million-token-6f4de1.mp3
26. AI Post Transformers: FengHuang for Rack-Scale LLM Inference Memory — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-12-fenghuang-for-rack-scale-llm-inference-m-62708e.mp3
Interactive Visualization: JANUS for Scalable MoE Inference

This episode explores a systems paper on making reinforcement-learning post-training for large language models practical over ordinary Ethernet and even WAN links, rather than requiring expensive RDMA clusters. It explains why trainer-actor RL creates a synchronization bottleneck, how full policy refreshes can dominate runtime on 1 to 10 gigabit networks, and why that turns bandwidth into a hidden limiter of who can run serious RL workloads. The discussion centers on the paper’s proposed solution: lossless sparse delta checkpoints that send only changed parameters, along with carefully encoded indices, streamed in parallel with rollout generation so actors can reconstruct the exact updated model without quantization or approximation. Listeners would find it interesting because it connects low-level systems design to the economics and accessibility of modern LLM training, asking whether better synchronization methods could open RL post-training to labs and startups outside elite infrastructure environments.

Interactive Visualization: Lossless Sparse Deltas for RL Networks
Sources:
1. RL over Commodity Networks: Overcoming the Bandwidth Barrier with Lossless Sparse Deltas — Chaoyi Ruan, Geng Luo, Xinyi Wan, Long Zhao, Qinghe Wang, Jiaan Zhu, Duling Xu, Guanbin Xu, Dehui Wei, Xiang Liu, Cheng Li, Haifeng Sun, Congcong Miao, Jialin Li, 2026
http://arxiv.org/abs/2602.11456
2. Understanding and Exploiting Weight Update Sparsity for Communication-Efficient Distributed RL — Erfan Miahi, Eugene Belilovsky, 2026
https://scholar.google.com/scholar?q=Understanding+and+Exploiting+Weight+Update+Sparsity+for+Communication-Efficient+Distributed+RL
3. StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation — Yinmin Zhong, Zili Zhang, Xiaoniu Song, Hanpeng Hu, Chao Jin, Bingyang Wu, Nuo Chen, Yukun Chen, Yu Zhou, Changyi Wan, Hongyu Zhou, Yimin Jiang, Yibo Zhu, Daxin Jiang, 2025
https://scholar.google.com/scholar?q=StreamRL%3A+Scalable%2C+Heterogeneous%2C+and+Elastic+RL+for+LLMs+with+Disaggregated+Stream+Generation
4. HybridFlow: A Flexible and Efficient RLHF Framework — Guangming Sheng, Chi Zhang, Zilingfeng Ye, Xibin Wu, Wang Zhang, Ru Zhang, Yanghua Peng, Haibin Lin, Chuan Wu, 2024
https://scholar.google.com/scholar?q=HybridFlow%3A+A+Flexible+and+Efficient+RLHF+Framework
5. OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework — Jian Hu, Xibin Wu, Wei Shen, Jason Klein Liu, Zilin Zhu, Weixun Wang, Songlin Jiang, Haoran Wang, Hao Chen, Bin Chen, Weikai Fang, Xianyu, Yu Cao, Haotian Xu, Yiming Liu, 2024
https://scholar.google.com/scholar?q=OpenRLHF%3A+An+Easy-to-use%2C+Scalable+and+High-performance+RLHF+Framework
6. How Can We Train Deep Learning Models Across Clouds and Continents? An Experimental Study — Alexander Erben, Ruben Mayer, Hans-Arno Jacobsen, 2023
https://scholar.google.com/scholar?q=How+Can+We+Train+Deep+Learning+Models+Across+Clouds+and+Continents%3F+An+Experimental+Study
7. Efficient Memory Management for Large Language Model Serving with PagedAttention — Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, Ion Stoica, 2023
https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention
8. Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models — authors not confirmed from snippet, recent, unconfirmed from snippet
https://scholar.google.com/scholar?q=Asynchronous+RLHF%3A+Faster+and+More+Efficient+Off-Policy+RL+for+Language+Models
9. Faster, More Efficient RLHF through Off-Policy Asynchronous Learning — authors not confirmed from snippet, recent, unconfirmed from snippet
https://scholar.google.com/scholar?q=Faster%2C+More+Efficient+RLHF+through+Off-Policy+Asynchronous+Learning
10. Stable Asynchrony: Variance-Controlled Off-Policy RL for LLMs — authors not confirmed from snippet, recent, unconfirmed from snippet
https://scholar.google.com/scholar?q=Stable+Asynchrony%3A+Variance-Controlled+Off-Policy+RL+for+LLMs
11. Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance — authors not confirmed from snippet, recent, unconfirmed from snippet
https://scholar.google.com/scholar?q=Efficient+Online+RFT+with+Plug-and-Play+LLM+Judges%3A+Unlocking+State-of-the-Art+Performance
12. Accelerating RL Post-Training Rollouts via System-Integrated Speculative Decoding — authors not confirmed from snippet, recent, unconfirmed from snippet
https://scholar.google.com/scholar?q=Accelerating+RL+Post-Training+Rollouts+via+System-Integrated+Speculative+Decoding
13. Beat the Long Tail: Distribution-Aware Speculative Decoding for RL Training — authors not confirmed from snippet, recent, unconfirmed from snippet
https://scholar.google.com/scholar?q=Beat+the+Long+Tail%3A+Distribution-Aware+Speculative+Decoding+for+RL+Training
14. AI Post Transformers: HALoS: Hierarchical Asynchronous LLM Training over Slow Networks — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/halos-hierarchical-asynchronous-llm-training-over-slow-networks/
15. AI Post Transformers: TensorFlow for Distributed Machine Learning Systems — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-06-tensorflow-for-distributed-machine-learn-b7fa52.mp3
16. AI Post Transformers: AgenticQwen and Small Industrial Tool Agents — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-27-agenticqwen-and-small-industrial-tool-ag-dc676d.mp3
Interactive Visualization: Lossless Sparse Deltas for RL Networks

This episode explores how the FlashFuser paper uses Hopper GPU inter-core communication to push kernel fusion beyond the usual single-SM memory limits, especially for transformer feed-forward networks and gated FFNs. It explains why this matters now: H100-class GPUs have gained compute far faster than memory bandwidth, making activation spills to HBM an increasingly painful bottleneck for workloads that can consume 40 to 60 percent of inference time. The discussion walks through Hopper’s distributed shared memory model and FlashFuser’s core idea of coordinating reduce, shuffle, and multiply patterns across SM clusters so large intermediate activations can stay on chip longer. Listeners would find it interesting because it connects compiler techniques, GPU architecture, and real transformer inference bottlenecks into a concrete argument about when newer hardware may finally make more aggressive fusion worthwhile.

Sources:
1. FlashFuser: Expanding the Scale of Kernel Fusion for Compute-Intensive Operators via Inter-Core Connection — Ziyu Huang, Yangjie Zhou, Zihan Liu, Xinhao Luo, Yijia Diao, Minyi Guo, Jidong Zhai, Yu Feng, Chen Zhang, Anbang Wu, Jingwen Leng, 2025
http://arxiv.org/abs/2512.12949
2. TVM: An Automated End-to-End Optimizing Compiler for Deep Learning — Tianqi Chen, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Yan, Meghan Cowan, Haichen Shen, Leyuan Wang, Yuwei Hu, Luis Ceze, Carlos Guestrin, Arvind Krishnamurthy, 2018
https://scholar.google.com/scholar?q=TVM%3A+An+Automated+End-to-End+Optimizing+Compiler+for+Deep+Learning
3. FusionStitching: Boosting Memory Intensive Computations for Deep Learning Workloads — Zhen Zheng, Pengzhan Zhao, Guoping Long, Feiwen Zhu, Kai Zhu, Wenyi Zhao, Lansong Diao, Jun Yang, Wei Lin, 2020
https://scholar.google.com/scholar?q=FusionStitching%3A+Boosting+Memory+Intensive+Computations+for+Deep+Learning+Workloads
4. Operator Fusion in XLA: Analysis and Evaluation — Daniel Snider, Ruofan Liang, 2023
https://scholar.google.com/scholar?q=Operator+Fusion+in+XLA%3A+Analysis+and+Evaluation
5. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness — Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, Christopher Re, 2022
https://scholar.google.com/scholar?q=FlashAttention%3A+Fast+and+Memory-Efficient+Exact+Attention+with+IO-Awareness
6. Benchmarking and Dissecting the Nvidia Hopper GPU Architecture — Weile Luo, Ruibo Fan, Zeyu Li, Dayou Du, Qiang Wang, Xiaowen Chu, 2024
https://scholar.google.com/scholar?q=Benchmarking+and+Dissecting+the+Nvidia+Hopper+GPU+Architecture
7. A Case Study in CUDA Kernel Fusion: Implementing FlashAttention-2 on NVIDIA Hopper Architecture using the CUTLASS Library — Ganesh Bikshandi, Jay Shah, 2023
https://scholar.google.com/scholar?q=A+Case+Study+in+CUDA+Kernel+Fusion%3A+Implementing+FlashAttention-2+on+NVIDIA+Hopper+Architecture+using+the+CUTLASS+Library
8. Scaling Deep Learning Computation over the Inter-Core Connected Intelligence Processor with T10 — Yiqi Liu, Yuqi Xue, Yu Cheng, Lingxiao Ma, Ziming Miao, Jilong Xue, Jian Huang, 2024
https://scholar.google.com/scholar?q=Scaling+Deep+Learning+Computation+over+the+Inter-Core+Connected+Intelligence+Processor+with+T10
9. FlashFuser: Expanding the Scale of Kernel Fusion for Compute-Intensive Operators via Inter-Core Connection — Ziyu Huang, Yangjie Zhou, Zihan Liu, Xinhao Luo, Yijia Diao, Minyi Guo, Jidong Zhai, Yu Feng, Chen Zhang, Anbang Wu, Jingwen Leng, 2025
https://scholar.google.com/scholar?q=FlashFuser%3A+Expanding+the+Scale+of+Kernel+Fusion+for+Compute-Intensive+Operators+via+Inter-Core+Connection
10. Chimera: An Analytical Optimizing Framework for Effective Compute-intensive Operators Fusion — Size Zheng, Siyuan Chen, Peidi Song, Renze Chen, Xiuhong Li, Shengen Yan, Dahua Lin, Jingwen Leng, Yun Liang, 2023
https://scholar.google.com/scholar?q=Chimera%3A+An+Analytical+Optimizing+Framework+for+Effective+Compute-intensive+Operators+Fusion
11. BOLT: Bridging the Gap between Auto-tuners and Hardware-native Performance — Jiarong Xing, Leyuan Wang, Shang Zhang, Jack Chen, Ang Chen, Yibo Zhu, 2022
https://scholar.google.com/scholar?q=BOLT%3A+Bridging+the+Gap+between+Auto-tuners+and+Hardware-native+Performance
12. MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators — Zheng Zhang, Donglin Yang, Xiaobo Zhou, Dazhao Cheng, 2024
https://scholar.google.com/scholar?q=MCFuser%3A+High-Performance+and+Rapid+Fusion+of+Memory-Bound+Compute-Intensive+Operators
13. Deep Kernel Fusion for Transformers — Zixi Zhang, Zhiwen Mo, Yiren Zhao, Robert Mullins, 2026
https://scholar.google.com/scholar?q=Deep+Kernel+Fusion+for+Transformers
14. Benchmarking thread block cluster — approximate; unclear from snippet, 2023-2026
https://scholar.google.com/scholar?q=Benchmarking+thread+block+cluster
15. ClusterSim: modeling thread block clusters in Hopper GPUs — approximate; unclear from snippet, 2023-2026
https://scholar.google.com/scholar?q=ClusterSim%3A+modeling+thread+block+clusters+in+Hopper+GPUs
16. Analysing and Reducing Costs of Deep Learning Compiler Auto-tuning — approximate; unclear from snippet, 2023-2026
https://scholar.google.com/scholar?q=Analysing+and+Reducing+Costs+of+Deep+Learning+Compiler+Auto-tuning
17. AI Post Transformers: Why LLM Serving Needs Mathematical Optimization — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-why-llm-serving-needs-mathematical-optim-647fc6.mp3
18. AI Post Transformers: Splitwise: Phase-Split LLM Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-26-splitwise-phase-split-llm-inference-e8945b.mp3
19. AI Post Transformers: LAPS for Length-Aware LLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-laps-for-length-aware-llm-serving-0c6149.mp3
20. AI Post Transformers: Speculative Decoding in Real vLLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-speculative-decoding-in-real-vllm-servin-6f4e2b.mp3
21. AI Post Transformers: Caffeine: A Unified FPGA for CNNs — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-06-caffeine-a-unified-fpga-for-cnns-e8acbe.mp3
22. AI Post Transformers: Gated Linear Attention for Efficient Long Sequences — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-18-gated-linear-attention-for-efficient-lon-c858ab.mp3
Interactive Visualization: FlashFuser and Hopper-Era FFN Kernel Fusion

This episode explores a systems paper on speeding up Transformer decoding by tightly fusing the SwiGLU MLP path, rather than focusing only on attention or long-context tricks. It explains why long output generation becomes memory-bandwidth bound, clarifying concepts like kernel fusion, HBM traffic, prefill versus autoregressive decode, and why repeated token-by-token inference exposes the MLP as a real bottleneck. The discussion walks through the paper’s main design choice: a disciplined fusion of the up-projection, gate projection, SiLU activation, and elementwise multiply into a single decode-stage kernel, while leaving the down projection separate to avoid worse scheduling and register-pressure tradeoffs. It also highlights the paper’s practical argument for profiler-driven runtime scheduling across row-major and column-major kernel variants, making the result interesting to listeners who care about how large-model serving performance is won through careful hardware-aware engineering rather than headline-grabbing algorithm changes.

Sources:
1. Deep Kernel Fusion for Transformer Decoding
https://arxiv.org/pdf/2602.11808
2. GLU Variants Improve Transformer — Noam Shazeer, 2020
https://scholar.google.com/scholar?q=GLU+Variants+Improve+Transformer
3. PaLM: Scaling Language Modeling with Pathways — Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Noam Shazeer and many others, 2022
https://scholar.google.com/scholar?q=PaLM%3A+Scaling+Language+Modeling+with+Pathways
4. DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale — Reza Yazdani Aminabadi, Samyam Rajbhandari, Minjia Zhang, Ammar Ahmad Awan, Cheng Li and others, 2022
https://scholar.google.com/scholar?q=DeepSpeed+Inference%3A+Enabling+Efficient+Inference+of+Transformer+Models+at+Unprecedented+Scale
5. FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning — Tri Dao, 2023
https://scholar.google.com/scholar?q=FlashAttention-2%3A+Faster+Attention+with+Better+Parallelism+and+Work+Partitioning
6. SGLang: Efficient Execution of Structured Language Model Programs — Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun, Jeff Huang, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonzalez, Clark Barrett, and Ying Sheng, 2024
https://scholar.google.com/scholar?q=SGLang%3A+Efficient+Execution+of+Structured+Language+Model+Programs
7. Efficient Memory Management for Large Language Model Serving with PagedAttention — Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica, 2023
https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention
8. Welder: Scheduling Deep Learning Memory Access via Tile-Graph — Yining Shi, Zhi Yang, Jilong Xue, Lingxiao Ma, Yuqing Xia, Ziming Miao, Yuxiao Guo, Fan Yang, and Lidong Zhou, 2023
https://scholar.google.com/scholar?q=Welder%3A+Scheduling+Deep+Learning+Memory+Access+via+Tile-Graph
9. FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving — Zihao Ye, Lequn Chen, Ruihang Lai, Wuwei Lin, Yineng Zhang, Stephanie Wang, Tianqi Chen, Baris Kasikci, Vinod Grover, Arvind Krishnamurthy, and Luis Ceze, 2025
https://scholar.google.com/scholar?q=FlashInfer%3A+Efficient+and+Customizable+Attention+Engine+for+LLM+Inference+Serving
10. Masked Gated Linear Unit — unknown from snippet, likely 2024 or 2025
https://scholar.google.com/scholar?q=Masked+Gated+Linear+Unit
11. SCBench: A KV Cache-Centric Analysis of Long-Context Methods — unknown from snippet, likely 2025
https://scholar.google.com/scholar?q=SCBench%3A+A+KV+Cache-Centric+Analysis+of+Long-Context+Methods
12. Model Tells You Where to Merge: Adaptive KV Cache Merging for LLMs on Long-Context Tasks — unknown from snippet, likely 2025
https://scholar.google.com/scholar?q=Model+Tells+You+Where+to+Merge%3A+Adaptive+KV+Cache+Merging+for+LLMs+on+Long-Context+Tasks
13. MEDA: Dynamic KV Cache Allocation for Efficient Multimodal Long-Context Inference — unknown from snippet, likely 2025
https://scholar.google.com/scholar?q=MEDA%3A+Dynamic+KV+Cache+Allocation+for+Efficient+Multimodal+Long-Context+Inference
14. Efficient LLM Inference Using Dynamic Input Pruning and Cache-Aware Masking — unknown from snippet, likely 2025
https://scholar.google.com/scholar?q=Efficient+LLM+Inference+Using+Dynamic+Input+Pruning+and+Cache-Aware+Masking
15. Enhancing Transformer Performance and Portability Through Auto-Tuning Frameworks — P. Siwinska et al. (approx.), unknown, likely recent
https://scholar.google.com/scholar?q=Enhancing+Transformer+Performance+and+Portability+Through+Auto-Tuning+Frameworks
16. AI Post Transformers: Splitwise: Phase-Split LLM Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-26-splitwise-phase-split-llm-inference-e8945b.mp3
17. AI Post Transformers: Why LLM Serving Needs Mathematical Optimization — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-why-llm-serving-needs-mathematical-optim-647fc6.mp3
18. AI Post Transformers: PackKV Lossy Compression for KV Caches — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-04-packkv-lossy-compression-for-kv-caches-b37bce.mp3
19. AI Post Transformers: SGLang for Faster Structured LLM Programs — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-06-sglang-for-faster-structured-llm-program-c59f1c.mp3
20. AI Post Transformers: Speculative Decoding in Real vLLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-speculative-decoding-in-real-vllm-servin-6f4e2b.mp3
21. AI Post Transformers: LAPS for Length-Aware LLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-laps-for-length-aware-llm-serving-0c6149.mp3
22. AI Post Transformers: CacheFlow and 3D-Parallel KV Cache Restoration — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-01-cacheflow-and-3d-parallel-kv-cache-resto-8db883.mp3
23. AI Post Transformers: FengHuang for Rack-Scale LLM Inference Memory — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-12-fenghuang-for-rack-scale-llm-inference-m-62708e.mp3
Interactive Visualization: Deep Kernel Fusion for Transformer Decoding

This episode explores a paper that argues AI can help mathematics most by orchestrating the full research workflow rather than acting as a one-shot chatbot. It discusses why real mathematical work depends on durable memory, branching hypotheses, literature search, proof attempts, computation, and recorded failures, and contrasts that with both ordinary chat interfaces and formal theorem provers such as Lean or Coq. The conversation details the paper’s multi-agent design, where a coordinator delegates parallel tasks like literature review, coding, proof exploration, and claim checking into a living draft document with provenance and uncertainty markers. It also highlights reported results on 100 research-level problems, where the full system outperformed strong single-shot models by using tactics such as SAT reduction, theorem retrieval, and coordinated theory-computation pipelines, making the episode interesting for listeners curious about how AI might become a practical research collaborator instead of just a clever text generator.

Sources:
1. AI Co-Mathematician for Mathematical Research
https://arxiv.org/pdf/2605.06651
2. A Survey on Deep Learning for Theorem Proving — Zhaoyu Li, Jialiang Sun, Logan Murphy, Qidong Su, Zenan Li, Xian Zhang, Kaiyu Yang, Xujie Si, 2024
https://arxiv.org/abs/2404.09939
3. GPT-f: Generative Language Modeling for Automated Theorem Proving — Stanislas Polu, Ilya Sutskever, 2020
https://openai.com/index/generative-language-modeling-for-automated-theorem-proving//
4. Draft, Sketch, and Prove: Guiding Formal Theorem Provers with Informal Proofs — Albert Q. Jiang, Sean Welleck, Jin Peng Zhou, Wenda Li, Jiacheng Liu, Mateja Jamnik, Timothee Lacroix, Yuhuai Wu, Guillaume Lample, 2023
https://arxiv.org/abs/2210.12283
5. Solving olympiad geometry without human demonstrations — Trieu H. Trinh, Yuhuai Wu, Quoc V. Le, He He, Thang Luong, 2024
https://www.nature.com/articles/s41586-023-06747-5
6. Exploration and Explanation in Computational Notebooks — Adam Rule, Aurelien Tabard, James D. Hollan, 2018
https://adamrule.com/files/papers/chi_2018_computational_notebooks_camera_ready.pdf
7. What’s Wrong with Computational Notebooks? Pain Points, Needs, and Design Opportunities — Souti Chattopadhyay, Ishita Prasad, Austin Z. Henley, Anita Sarma, Titus Barik, 2020
https://www.microsoft.com/en-us/research/publication/whats-wrong-with-computational-notebooks/
8. Principles for data analysis workflows — Sara Stoudt, Valeri N. Vasquez, Ciera C. Martinez, 2021
https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1008770
9. Towards an AI co-scientist — Juraj Gottweis, Wei-Hung Weng, Alexander Daryin, Tao Tu, Anil Palepu, Petar Sirkovic and many others, 2025
https://arxiv.org/abs/2502.18864
10. Towards Autonomous Mathematics Research — T. Feng, T. H. Trinh, G. Bingham, D. Hwang, Y. Chervonyi, J. Jung, J. Lee, C. Pagano, S.-h. Kim, F. Pasqualotto, S. Gukov, J. N. Lee, J. Kim, K. Hou, G. Ghiasi, Y. Tay, Y. Li, C. Kuang, Y. Liu, H. Lin, E. Z. Liu, N. Nayakanti, X. Yang, H.-t. Cheng, D. Hassabis, K. Kavukcuoglu, Q. V. Le, and T. Luong, 2026
https://scholar.google.com/scholar?q=Towards+Autonomous+Mathematics+Research
11. Olympiad-level formal mathematical reasoning with reinforcement learning — T. Hubert, R. S. Mehta, L. Sartran, M. Z. Horvath, G. Zuzic, E. Wieser, A. Huang, J. Schrittwieser, Y. Schroecker, H. Masoom, O. Bertolli, T. Zahavy, A. Mandhane, J. Yung, I. Beloshapka, B. Ibarz, V. Veeriah, L. Yu, O. Nash, P. Lezeau, S. Mercuri, C. Sonne, B. Mehta, A. Davies, D. Zheng, F. Pedregosa, Y. Li, I. von Glehn, M. Rowland, S. Albanie, A. Velingker, S. Schmitt, E. Lockhart, E. Hughes, H. Michalewski, N. Sonnerat, D. Hassabis, P. Kohli, and D. Silver, 2025
https://scholar.google.com/scholar?q=Olympiad-level+formal+mathematical+reasoning+with+reinforcement+learning
12. FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI — E. Glazer, E. Erdil, T. Besiroglu, D. Chicharro, E. Chen, A. Gunning, C. F. Olsson, J.-S. Denain, A. Ho, E. de Oliveira Santos, O. Jarviniemi, M. Barnett, R. Sandler, M. Vrzala, J. Sevilla, Q. Ren, E. Pratt, L. Levine, G. Barkley, N. Stewart, B. Grechuk, T. Grechuk, S. V. Enugandla, and M. Wildon, 2024
https://scholar.google.com/scholar?q=FrontierMath%3A+A+Benchmark+for+Evaluating+Advanced+Mathematical+Reasoning+in+AI
13. Hard2Verify: A Step-Level Verification Benchmark for Open-Ended Frontier Math — S. Pandit, A. Xu, X.-P. Nguyen, Y. Ming, C. Xiong, and S. Joty, 2025
https://scholar.google.com/scholar?q=Hard2Verify%3A+A+Step-Level+Verification+Benchmark+for+Open-Ended+Frontier+Math
14. Longer Context, Deeper Thinking: Uncovering the Role of Long-Context Ability in Reasoning — Wang Yang et al., 2025
https://scholar.google.com/scholar?q=Longer+Context%2C+Deeper+Thinking%3A+Uncovering+the+Role+of+Long-Context+Ability+in+Reasoning
15. InftyThink: Breaking the Length Limits of Long-Context Reasoning in Large Language Models — Yuchen Yan et al., 2025
https://scholar.google.com/scholar?q=InftyThink%3A+Breaking+the+Length+Limits+of+Long-Context+Reasoning+in+Large+Language+Models
16. Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models — Andy Zhou et al., 2023
https://scholar.google.com/scholar?q=Language+Agent+Tree+Search+Unifies+Reasoning+Acting+and+Planning+in+Language+Models
17. Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training — Xidong Feng et al., 2023
https://scholar.google.com/scholar?q=Alphazero-like+Tree-Search+can+Guide+Large+Language+Model+Decoding+and+Training
18. ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving — Zhibin Gou et al., 2023
https://scholar.google.com/scholar?q=ToRA%3A+A+Tool-Integrated+Reasoning+Agent+for+Mathematical+Problem+Solving
19. Efficient Tool Use with Chain-of-Abstraction Reasoning — Silin Gao et al., 2024
https://scholar.google.com/scholar?q=Efficient+Tool+Use+with+Chain-of-Abstraction+Reasoning
20. MuMath-Code: Combining Tool-Use Large Language Models with Multi-perspective Data Augmentation for Mathematical Reasoning — Shuo Yin et al., 2024
https://scholar.google.com/scholar?q=MuMath-Code%3A+Combining+Tool-Use+Large+Language+Models+with+Multi-perspective+Data+Augmentation+for+Mathematical+Reasoning
21. HybridProver: Augmenting Theorem Proving with LLM-Driven Proof Synthesis and Refinement — Jilin Hu et al., 2025
https://scholar.google.com/scholar?q=HybridProver%3A+Augmenting+Theorem+Proving+with+LLM-Driven+Proof+Synthesis+and+Refinement
22. A Minimal Agent for Automated Theorem Proving — Borja Requena et al., 2026
https://scholar.google.com/scholar?q=A+Minimal+Agent+for+Automated+Theorem+Proving
23. Tree Search for Language Model Agents — Jing Yu Koh et al., 2024
https://scholar.google.com/scholar?q=Tree+Search+for+Language+Model+Agents
24. AI Post Transformers: Test-time Scaling for Multi-Agent Collaborative Reasoning — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-22-test-time-scaling-for-multi-agent-collab-082570.mp3
25. AI Post Transformers: Agentic AI and the Next Intelligence Explosion — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-28-agentic-ai-and-the-next-intelligence-exp-d06561.mp3
26. AI Post Transformers: Kimi K2.5 and Visual Agent Swarms — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-24-kimi-k25-and-visual-agent-swarms-7d04d7.mp3
Interactive Visualization: AI Co-Mathematician for Mathematical Research

This episode explores TMAS, a framework for scaling test-time reasoning by coordinating multiple specialized agents instead of simply letting a single model think longer. It explains how the system combines proposal, verification, refinement, and shared hierarchical memory so that useful intermediate results and higher-level strategy guidance can be reused across parallel reasoning attempts. The discussion highlights the paper’s central argument that better orchestration, careful memory design, and reinforcement learning objectives for exploration and productive memory use can turn extra inference compute into genuine reasoning gains rather than redundant or noisy work. Listeners would find it interesting for its clear comparison to self-consistency, Tree of Thoughts, and newer coordinated-reasoning methods, along with its emphasis on compute-matched evaluation, reproducibility, and the practical challenge of making multi-agent “synergy” real instead of just expensive parallelism.

Sources:
1. TMAS: Scaling Test-Time Compute via Multi-Agent Synergy — George Wu, Nan Jing, Qing Yi, Chuan Hao, Ming Yang, Feng Chang, Yuan Wei, Jian Yang, Ran Tao, Bryan Dai, 2026
http://arxiv.org/abs/2605.10344
2. A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well? — Qiyuan Zhang, Fuyuan Lyu, Zexu Sun, Lei Wang, Weixu Zhang, Wenyue Hua, Haolun Wu, Zhihan Guo, Yufei Wang, Niklas Muennighoff, Irwin King, Xue Liu, Chen Ma, 2025
https://scholar.google.com/scholar?q=A+Survey+on+Test-Time+Scaling+in+Large+Language+Models%3A+What%2C+How%2C+Where%2C+and+How+Well%3F
3. Self-Consistency Improves Chain of Thought Reasoning in Language Models — Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, Denny Zhou, 2023
https://scholar.google.com/scholar?q=Self-Consistency+Improves+Chain+of+Thought+Reasoning+in+Language+Models
4. Tree of Thoughts: Deliberate Problem Solving with Large Language Models — Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, Karthik Narasimhan, 2023
https://scholar.google.com/scholar?q=Tree+of+Thoughts%3A+Deliberate+Problem+Solving+with+Large+Language+Models
5. PaCoRe: Learning to Scale Test-Time Compute with Parallel Coordinated Reasoning — Jingcheng Hu, Yinmin Zhang, Shijie Shang, Xiaobo Yang, Yue Peng, Zhewei Huang, Hebin Zhou, Xin Wu, Jie Cheng, Fanqi Wan, Xiangwen Kong, Chengyuan Yao, Kaiwen Yan, Ailin Huang, Hongyu Zhou, Qi Han, Zheng Ge, Daxin Jiang, Xiangyu Zhang, Heung-Yeung Shum, 2026
https://scholar.google.com/scholar?q=PaCoRe%3A+Learning+to+Scale+Test-Time+Compute+with+Parallel+Coordinated+Reasoning
6. Do Not Waste Your Rollouts: Recycling Search Experience for Efficient Test-Time Scaling — Xinglin Wang, Jiayi Shi, Shaoxiong Feng, Peiwen Yuan, Yiwei Li, Yueqi Zhang, Chuyi Tan, Ji Zhang, Boyuan Pan, Yao Hu, Kan Li, 2026
https://scholar.google.com/scholar?q=Do+Not+Waste+Your+Rollouts%3A+Recycling+Search+Experience+for+Efficient+Test-Time+Scaling
7. Multi-Agent Verification: Scaling Test-Time Compute with Multiple Verifiers — Shalev Lifshitz, Sheila A. McIlraith, Yilun Du, 2025
https://scholar.google.com/scholar?q=Multi-Agent+Verification%3A+Scaling+Test-Time+Compute+with+Multiple+Verifiers
8. Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters — Charlie Snell, Jaehoon Lee, Kelvin Xu, Aviral Kumar, 2024
https://scholar.google.com/scholar?q=Scaling+LLM+Test-Time+Compute+Optimally+can+be+More+Effective+than+Scaling+Model+Parameters
9. Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation — Zhuolin Yang, Zihan Liu, Yang Chen, Wenliang Dai, Boxin Wang, Sheng-Chieh Lin, Chankyu Lee, Yangyi Chen, Dongfu Jiang, Jiafan He, Renjie Pi, Grace Lam, Nayeon Lee, Alexander Bukharin, Mohammad Shoeybi, Bryan Catanzaro, Wei Ping, 2026
https://scholar.google.com/scholar?q=Nemotron-Cascade+2%3A+Post-Training+LLMs+with+Cascade+RL+and+Multi-Domain+On-Policy+Distillation
10. ArcMemo: Abstract Reasoning Composition with Lifelong LLM Memory — Matthew Ho et al., 2025
https://scholar.google.com/scholar?q=ArcMemo%3A+Abstract+Reasoning+Composition+with+Lifelong+LLM+Memory
11. Metacognitive Reuse: Turning Recurring LLM Reasoning Into Concise Behaviors — Aniket Didolkar, Nicolas Ballas, Sanjeev Arora, Anirudh Goyal, 2025
https://scholar.google.com/scholar?q=Metacognitive+Reuse%3A+Turning+Recurring+LLM+Reasoning+Into+Concise+Behaviors
12. Reuse, Don't Recompute: Efficient Large Reasoning Model Inference via Memory Orchestration — Daivik Patel, Shrenik Patel, 2025
https://scholar.google.com/scholar?q=Reuse%2C+Don%27t+Recompute%3A+Efficient+Large+Reasoning+Model+Inference+via+Memory+Orchestration
13. Diversity of Thought Improves Reasoning Abilities of Large Language Models — Ranjita Naik et al., 2023
https://scholar.google.com/scholar?q=Diversity+of+Thought+Improves+Reasoning+Abilities+of+Large+Language+Models
14. Diversity-Enhanced Reasoning for Subjective Questions — Yumeng Wang et al., 2025
https://scholar.google.com/scholar?q=Diversity-Enhanced+Reasoning+for+Subjective+Questions
15. Scaling Large Language Model-based Multi-Agent Collaboration — Chen Qian et al., 2024
https://scholar.google.com/scholar?q=Scaling+Large+Language+Model-based+Multi-Agent+Collaboration
16. Multi-Agent Sampling: Scaling Inference Compute for Data Synthesis with Tree Search-Based Agentic Collaboration — Hai Ye, Mingbao Lin, Hwee Tou Ng, Shuicheng Yan, 2024
https://scholar.google.com/scholar?q=Multi-Agent+Sampling%3A+Scaling+Inference+Compute+for+Data+Synthesis+with+Tree+Search-Based+Agentic+Collaboration
17. Reasoning Relay: Evaluating Stability and Interchangeability of Large Language Models in Mathematical Reasoning — Leo Lu et al., 2025
https://scholar.google.com/scholar?q=Reasoning+Relay%3A+Evaluating+Stability+and+Interchangeability+of+Large+Language+Models+in+Mathematical+Reasoning
18. AI Post Transformers: Test-time Scaling for Multi-Agent Collaborative Reasoning — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-22-test-time-scaling-for-multi-agent-collab-082570.mp3
19. AI Post Transformers: TUMIX Multi-Agent Test-Time Scaling with Tools — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-22-tumix-multi-agent-test-time-scaling-with-40671c.mp3
20. AI Post Transformers: DeepVerifier: Self-Evolving Research Agents via Rubric-Guided Verification — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/deepverifier-self-evolving-research-agents-via-rubric-guided-verification/
21. AI Post Transformers: MEMSEARCHER: Reinforcement Learning for LLM Memory Management — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-memsearcher-reinforcement-learning-for-l-e9ad84.mp3
22. AI Post Transformers: IMO-Bench for Robust Mathematical Reasoning — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-imo-bench-for-robust-mathematical-reason-143489.mp3
23. AI Post Transformers: Breaking the Prefix Barrier with Shared KV Cache — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-24-breaking-the-prefix-barrier-with-shared-a5e5a6.mp3
Interactive Visualization: TMAS: Scaling Test-Time Compute with Multi-Agent Synergy

This episode explores MELT, a looped language model architecture that aims to preserve latent reasoning benefits while preventing KV-cache memory from growing with every reasoning pass. It explains how the paper reframes the problem as a systems and architecture challenge, replacing per-loop cached attention state with a single shared, gated cache per layer inspired by recurrent models like LSTMs and Universal Transformers. The discussion weighs whether this is a genuine shift in reasoning architecture or a narrower engineering improvement, ultimately arguing that the paper’s real contribution is efficient cache management rather than a wholly new paradigm. Listeners would find it interesting for its clear breakdown of inference-time compute scaling, latent reasoning, and why memory bottlenecks could shape the future of practical reasoning models.

Interactive Visualization: MELT: Decoupling Compute From Memory
Sources:
1. MELT: Decoupling Compute From Memory
https://arxiv.org/pdf/2605.07721
2. Scaling Latent Reasoning via Looped Language Models — Rui-Jie Zhu et al., 2025
https://scholar.google.com/scholar?q=Scaling+Latent+Reasoning+via+Looped+Language+Models
3. Universal Transformers — Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, Lukasz Kaiser, 2018
https://scholar.google.com/scholar?q=Universal+Transformers
4. Reducing Transformer Key-Value Cache Size with Cross-Layer Attention — William Brandon, Mayank Mishra, Aniruddha Nrusimha, Rameswar Panda, Jonathan Ragan Kelly, 2024
https://scholar.google.com/scholar?q=Reducing+Transformer+Key-Value+Cache+Size+with+Cross-Layer+Attention
5. GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints — Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebron, Sumit Sanghai, 2023
https://scholar.google.com/scholar?q=GQA%3A+Training+Generalized+Multi-Query+Transformer+Models+from+Multi-Head+Checkpoints
6. Fast Transformer Decoding: One Write-Head is All You Need — Noam Shazeer, 2019
https://scholar.google.com/scholar?q=Fast+Transformer+Decoding%3A+One+Write-Head+is+All+You+Need
7. Beyond Speedup -- Utilizing KV Cache for Sampling and Reasoning — Zeyu Xing, Xing Li, Hui-Ling Zhen, Mingxuan Yuan, Sinno Jialin Pan, 2026
https://scholar.google.com/scholar?q=Beyond+Speedup+--+Utilizing+KV+Cache+for+Sampling+and+Reasoning
8. MemShare: Memory Efficient Inference for Large Reasoning Models through KV Cache Reuse — Kaiwen Chen, Xin Tan, Minchen Yu, Hong Xu, 2025
https://scholar.google.com/scholar?q=MemShare%3A+Memory+Efficient+Inference+for+Large+Reasoning+Models+through+KV+Cache+Reuse
9. Beyond Attention: Breaking the Limits of Transformer Context Length with Recurrent Memory — Aydar Bulatov, Yurii Kuratov, Yermek Kapushev, Mikhail Burtsev, 2024
https://scholar.google.com/scholar?q=Beyond+Attention%3A+Breaking+the+Limits+of+Transformer+Context+Length+with+Recurrent+Memory
10. Towards Understanding Distilled Reasoning Models: A Representational Approach — David D. Baek, Max Tegmark, 2025
https://scholar.google.com/scholar?q=Towards+Understanding+Distilled+Reasoning+Models%3A+A+Representational+Approach
11. AI Post Transformers: Gated Linear Attention for Efficient Long Sequences — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-18-gated-linear-attention-for-efficient-lon-c858ab.mp3
12. AI Post Transformers: PackKV Lossy Compression for KV Caches — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-04-packkv-lossy-compression-for-kv-caches-b37bce.mp3
13. AI Post Transformers: RetrievalAttention for Long-Context LLM Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-17-retrievalattention-for-long-context-llm-ddf566.mp3
14. AI Post Transformers: Parallelizing DeltaNet Linear Transformers over Sequence Length — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-18-parallelizing-deltanet-linear-transforme-2d0377.mp3
15. AI Post Transformers: Mamba-3 for Efficient Sequence Modeling — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-16-mamba-3-for-efficient-sequence-modeling-97a22a.mp3
16. AI Post Transformers: Jet-Nemotron and PostNAS for Faster Long Context — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-24-jet-nemotron-and-postnas-for-faster-long-436381.mp3
17. AI Post Transformers: In-Place Test-Time Training for Transformers — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-09-in-place-test-time-training-for-transfor-d0b976.mp3
18. AI Post Transformers: How Induction Heads Emerge in Transformers — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-03-how-induction-heads-emerge-in-transforme-a7bfcb.mp3
Interactive Visualization: MELT: Decoupling Compute From Memory

This episode explores the Qwen-Image-2.0 technical report and its central claim that a single unified multimodal diffusion model can handle high-quality image generation, precise image editing, multilingual text rendering, ultra-long text, and complex instruction following in one system. It traces the technical background from vision transformers and latent diffusion through newer editing and text-rendering methods, explaining why image editing is fundamentally harder than generation because users expect strict preservation of identity, layout, and other details. The discussion emphasizes that text in images remains a stubborn problem because models must treat letters as exact symbols rather than visual texture, especially for posters, ads, slides, comics, and UI mockups. Listeners would find it interesting because it connects benchmark claims to real product failures and asks whether this model finally reduces the messy patchwork of specialized tools into a single backbone that can actually obey, spell, preserve, and compose well.

Sources:
1. Qwen-Image-2.0 Technical Report — Bing Zhao, Chenfei Wu, Deqing Li, Hao Meng, Jiahao Li, Jie Zhang, Jingren Zhou, Junyang Lin, Kaiyuan Gao, Kuan Cao, Kun Yan, Liang Peng, Lihan Jiang, Niantong Li, Ningyuan Tang, Shengming Yin, Tianhe Wu, Xiao Xu, Xiaoyue Chen, Xihua Wang, Yan Shu, Yanran Zhang, Yi Wang, Yilei Chen, Ying Ba, Yixian Xu, Yujia Wu, Yuxiang Chen, Zecheng Tang, Zekai Zhang, Zhendong Wang, Zihao Liu, Zikai Zhou, An Yang, Chen Cheng, Chenxu Lv, Dayiheng Liu, Fan Zhou, Hantian Xiong, Hongzhu Shi, Hu Wei, Huihong Zhao, Ivy Liu, Jianwei Zhang, Jiawei Zhang, Kai Chen, Kang He, Levon Xue, Lin Qu, Linhan Tang, Luwen Feng, Minggang Wu, Minmin Sun, Na Ni, Rui Men, Shuai Bai, Sishou Zheng, Tao Lan, Tianqi Zhang, Tingkun Wen, Wei Wang, Weixu Qiao, Weiyi Lu, Wenmeng Zhou, Xiaodong Deng, Xiaoxiao Xu, Xinlei Fang, Xionghui Chen, Yanan Wang, Yang Fan, Yichang Zhang, Yixuan Xu, Yu Wu, Zhiyuan Ma, Zhizhi Cai, 2026
http://arxiv.org/abs/2605.10730
2. SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations — Chenlin Meng, Yutong He, Yang Song, Jiaming Song, Jiajun Wu, Jun-Yan Zhu, Stefano Ermon, 2021
https://scholar.google.com/scholar?q=SDEdit%3A+Guided+Image+Synthesis+and+Editing+with+Stochastic+Differential+Equations
3. Prompt-to-Prompt Image Editing with Cross-Attention Control — Amir Hertz, Ron Mokady, Jay Tenenbaum, Kfir Aberman, Yael Pritch, Daniel Cohen-Or, 2022
https://scholar.google.com/scholar?q=Prompt-to-Prompt+Image+Editing+with+Cross-Attention+Control
4. InstructPix2Pix: Learning to Follow Image Editing Instructions — Tim Brooks, Aleksander Holynski, Alexei A. Efros, 2022
https://scholar.google.com/scholar?q=InstructPix2Pix%3A+Learning+to+Follow+Image+Editing+Instructions
5. Emu Edit: Precise Image Editing via Recognition and Generation Tasks — Shelly Sheynin, Adam Polyak, Uriel Singer, Yuval Kirstain, Amit Zohar, Oron Ashual, Devi Parikh, Yaniv Taigman, 2023
https://scholar.google.com/scholar?q=Emu+Edit%3A+Precise+Image+Editing+via+Recognition+and+Generation+Tasks
6. GlyphDraw: Seamlessly Rendering Text with Intricate Spatial Structures in Text-to-Image Generation — Jian Ma, Mingjun Zhao, Chen Chen, Ruichen Wang, Di Niu, Haonan Lu, Xiaodong Lin, 2023
https://scholar.google.com/scholar?q=GlyphDraw%3A+Seamlessly+Rendering+Text+with+Intricate+Spatial+Structures+in+Text-to-Image+Generation
7. TextDiffuser: Diffusion Models as Text Painters — Jingye Chen, Yupan Huang, Tengchao Lv, Lei Cui, Qifeng Chen, Furu Wei, 2023
https://scholar.google.com/scholar?q=TextDiffuser%3A+Diffusion+Models+as+Text+Painters
8. AnyText: Multilingual Visual Text Generation and Editing — Yuxiang Tuo, Wangmeng Xiang, Jun-Yan He, Yifeng Geng, Xuansong Xie, 2023
https://scholar.google.com/scholar?q=AnyText%3A+Multilingual+Visual+Text+Generation+and+Editing
9. EasyText: Controllable Diffusion Transformer for Multilingual Text Rendering — Runnan Lu, Yuxuan Zhang, Jailing Liu, Haifa Wang, Yiren Song, 2025
https://scholar.google.com/scholar?q=EasyText%3A+Controllable+Diffusion+Transformer+for+Multilingual+Text+Rendering
10. Improving Image Generation with Better Captions — James Betker, Gabriel Goh, Li Jing, Tim Brooks, Jianfeng Wang, Linjie Li, Long Ouyang, Juntang Zhuang, Joyce Lee, Yufei Guo, Wesam Manassra, Prafulla Dhariwal, Casey Chu, Yunxin Jiao, Aditya Ramesh, 2023
https://scholar.google.com/scholar?q=Improving+Image+Generation+with+Better+Captions
11. T2I-CompBench: A Comprehensive Benchmark for Open-world Compositional Text-to-image Generation — Kaiyi Huang, Kaiyue Sun, Enze Xie, Zhenguo Li, Xihui Liu, 2023
https://scholar.google.com/scholar?q=T2I-CompBench%3A+A+Comprehensive+Benchmark+for+Open-world+Compositional+Text-to-image+Generation
12. MagicBrush: A Manually Annotated Dataset for Instruction-Guided Image Editing — Kai Zhang, Lingbo Mo, Wenhu Chen, Huan Sun, Yu Su, 2023
https://scholar.google.com/scholar?q=MagicBrush%3A+A+Manually+Annotated+Dataset+for+Instruction-Guided+Image+Editing
13. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale — Alexey Dosovitskiy et al., 2020
https://scholar.google.com/scholar?q=An+Image+is+Worth+16x16+Words%3A+Transformers+for+Image+Recognition+at+Scale
14. High-Resolution Image Synthesis with Latent Diffusion Models — Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, Bjorn Ommer, 2022
https://scholar.google.com/scholar?q=High-Resolution+Image+Synthesis+with+Latent+Diffusion+Models
15. Scalable Diffusion Models with Transformers — William Peebles, Saining Xie, 2023
https://scholar.google.com/scholar?q=Scalable+Diffusion+Models+with+Transformers
16. PixArt-α: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis — Junsong Chen et al., 2024
https://scholar.google.com/scholar?q=PixArt-%CE%B1%3A+Fast+Training+of+Diffusion+Transformer+for+Photorealistic+Text-to-Image+Synthesis
17. FLUX — Black Forest Labs, 2024
https://scholar.google.com/scholar?q=FLUX
18. Qwen3-VL — Qwen Team, 2025
https://scholar.google.com/scholar?q=Qwen3-VL
19. OpenAI Image Generation System Card / product report — OpenAI, 2025
https://scholar.google.com/scholar?q=OpenAI+Image+Generation+System+Card+%2F+product+report
20. Google image generation product/report cited in the introduction — Google, 2025
https://scholar.google.com/scholar?q=Google+image+generation+product%2Freport+cited+in+the+introduction
21. Long-Text-to-Image Generation via Compositional Prompt Decomposition — Jen-Yuan Huang, Tong Lin, Yilun Du, 2026
https://scholar.google.com/scholar?q=Long-Text-to-Image+Generation+via+Compositional+Prompt+Decomposition
22. TIT-Score: Evaluating Long-Prompt Based Text-to-Image Alignment via Text-to-Image-to-Text Consistency — Juntong Wang et al., 2025
https://scholar.google.com/scholar?q=TIT-Score%3A+Evaluating+Long-Prompt+Based+Text-to-Image+Alignment+via+Text-to-Image-to-Text+Consistency
23. ImgEdit: A Unified Image Editing Dataset and Benchmark — Yang Ye, Xianyi He, Zongjian Li, Bin Lin, Shenghai Yuan, Zhiyuan Yan, Bohan Hou, Li Yuan, 2025
https://scholar.google.com/scholar?q=ImgEdit%3A+A+Unified+Image+Editing+Dataset+and+Benchmark
24. UniVG: A Generalist Diffusion Model for Unified Image Generation and Editing — Tsu-Jui Fu, Yusu Qian, Chen Chen, Wenze Hu, Zhe Gan, Yinfei Yang, 2025
https://scholar.google.com/scholar?q=UniVG%3A+A+Generalist+Diffusion+Model+for+Unified+Image+Generation+and+Editing
25. DreamOmni: Unified Image Generation and Editing — Bin Xia, Yuechen Zhang, Jingyao Li, Chengyao Wang, Yitong Wang, Xinglong Wu, Bei Yu, Jiaya Jia, 2024
https://scholar.google.com/scholar?q=DreamOmni%3A+Unified+Image+Generation+and+Editing
26. Taming Rectified Flow for Inversion and Editing — Jiangshan Wang, Junfu Pu, Zhongang Qi, Jiayi Guo, et al., 2024
https://scholar.google.com/scholar?q=Taming+Rectified+Flow+for+Inversion+and+Editing
27. InstantEdit: Text-Guided Few-Step Image Editing with Piecewise Rectified Flow — Yiming Gong, Zhen Zhu, Minjia Zhang, 2025
https://scholar.google.com/scholar?q=InstantEdit%3A+Text-Guided+Few-Step+Image+Editing+with+Piecewise+Rectified+Flow
28. DS-VLM: Diffusion Supervision Vision Language Model — Zhen Sun, Yunhang Shen, Jie Li, Xing Sun, Pingyang Dai, Liujuan Cao, Rongrong Ji, 2025
https://scholar.google.com/scholar?q=DS-VLM%3A+Diffusion+Supervision+Vision+Language+Model
29. AI Post Transformers: VL-JEPA for Vision-Language Semantic Prediction — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-12-vl-jepa-for-vision-language-semantic-pre-69c9f4.mp3
Interactive Visualization: Qwen-Image-2.0 for Unified Generation and Editing

This episode explores the paper δ-mem, which argues that long context windows are not the same as true memory and proposes a compact online memory module for frozen language models. It explains how the method uses a tiny mutable state matrix, updated with a delta rule, to store residual errors over time and feed that state back into generation as a low-rank attention correction rather than replaying full conversation history. The discussion also examines why benchmarks like LoCoMo and MemoryAgentBench matter more than generic reasoning tests for evaluating memory, because they probe persistence, conflict resolution, and incremental updating across turns. Listeners would find it interesting because the episode connects an unusual architectural idea to concrete empirical gains, including stronger results on memory-heavy tasks despite using an extremely small memory state.

Interactive Visualization: δ-mem and Online Memory for LLMs
Sources:
1. $δ$-mem: Efficient Online Memory for Large Language Models — Jingdi Lei, Di Zhang, Junxian Li, Weida Wang, Kaixuan Fan, Xiang Liu, Qihan Liu, Xiaoteng Ma, Baian Chen, Soujanya Poria, 2026
http://arxiv.org/abs/2605.12357
2. Adaptive Switching Circuits — Bernard Widrow, Marcian E. Hoff, 1960
https://isl.stanford.edu/~widrow/papers/c1960adaptiveswitching.pdf
3. Online Learning and Online Convex Optimization — Shai Shalev-Shwartz, 2012
https://cir.nii.ac.jp/crid/1363388845866612864
4. Test-Time Training with Self-Supervision for Generalization under Distribution Shifts — Yu Sun, Xiaolong Wang, Zhuang Liu, John Miller, Alexei A. Efros, Moritz Hardt, 2020
https://proceedings.mlr.press/v119/sun20b.html
5. TiC-CLIP: Continual Training of CLIP Models — Saurabh Garg, Mehrdad Farajtabar, Hadi Pouransari, Sachin Mehta, Raviteja Vemulapalli, Oncel Tuzel, Vaishaal Shankar, Fartash Faghri, 2024
https://machinelearning.apple.com/research/tic-clip-v2
6. Evaluating Very Long-Term Conversational Memory of LLM Agents — Adyasha Maharana, Dong-Ho Lee, Sergey Tulyakov, Mohit Bansal, Francesco Barbieri, Yuwei Fang, 2024
https://scholar.google.com/scholar?q=Evaluating+Very+Long-Term+Conversational+Memory+of+LLM+Agents
7. Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions — Yuanzhe Hu, Yu Wang, Julian McAuley, 2025
https://scholar.google.com/scholar?q=Evaluating+Memory+in+LLM+Agents+via+Incremental+Multi-Turn+Interactions
8. Titans: Learning to Memorize at Test Time — Ali Behrouz, Peilin Zhong, Vahab Mirrokni, 2025
https://scholar.google.com/scholar?q=Titans%3A+Learning+to+Memorize+at+Test+Time
9. Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory — Prateek Chhikara, Dev Khant, Saket Aryan, Taranjeet Singh, Deshraj Yadav, 2025
https://scholar.google.com/scholar?q=Mem0%3A+Building+Production-Ready+AI+Agents+with+Scalable+Long-Term+Memory
10. E-mem: Multi-agent based Episodic Context Reconstruction for LLM Agent Memory — Kaixiang Wang, Yidan Lin, Jiong Lou, Zhaojiacheng Zhou, Bunyod Suvonov, Jie Li, 2026
https://scholar.google.com/scholar?q=E-mem%3A+Multi-agent+based+Episodic+Context+Reconstruction+for+LLM+Agent+Memory
11. HiMem: Hierarchical Long-Term Memory for LLM Long-Horizon Agents — Ningning Zhang, Xingxing Yang, Zhizhong Tan, Weiping Deng, Wenyong Wang, 2026
https://scholar.google.com/scholar?q=HiMem%3A+Hierarchical+Long-Term+Memory+for+LLM+Long-Horizon+Agents
12. Continuum Memory Architectures for Long-Horizon LLM Agents — Joe Logan, 2026
https://scholar.google.com/scholar?q=Continuum+Memory+Architectures+for+Long-Horizon+LLM+Agents
13. Gated Delta Networks: Improving Mamba2 with Delta Rule — Songlin Yang, Jan Kautz, Ali Hatamizadeh, 2024/2025
https://scholar.google.com/scholar?q=Gated+Delta+Networks%3A+Improving+Mamba2+with+Delta+Rule
14. Parallelizing Linear Transformers with the Delta Rule over Sequence Length — Songlin Yang, Bailin Wang, Yu Zhang, Yikang Shen, Yoon Kim, 2024
https://scholar.google.com/scholar?q=Parallelizing+Linear+Transformers+with+the+Delta+Rule+over+Sequence+Length
15. Efficiently Learning at Test-Time: Active Fine-Tuning of LLMs — Jonas Hübotter, Sascha Bongni, Ido Hakimi, Andreas Krause, 2024
https://scholar.google.com/scholar?q=Efficiently+Learning+at+Test-Time%3A+Active+Fine-Tuning+of+LLMs
16. Test-Time Learning for Large Language Models — Jinwu Hu, Zhitian Zhang, Guohao Chen, Xutao Wen, Chao Shuai, Wei Luo, Bin Xiao, Yuanqing Li, Mingkui Tan, 2025
https://scholar.google.com/scholar?q=Test-Time+Learning+for+Large+Language+Models
17. Efficient Low Rank Attention for Long-Context Inference in Large Language Models — Tenghui Li, Guoxu Zhou, Xuyang Zhao, Yuning Qiu, Qibin Zhao, 2025
https://scholar.google.com/scholar?q=Efficient+Low+Rank+Attention+for+Long-Context+Inference+in+Large+Language+Models
18. AI Post Transformers: Explicit Information Transmission for Context Compression — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-explicit-information-transmission-for-co-24e3c2.mp3
19. AI Post Transformers: Parallelizing DeltaNet Linear Transformers over Sequence Length — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-18-parallelizing-deltanet-linear-transforme-2d0377.mp3
20. AI Post Transformers: Doc-to-LoRA: Internalizing Context as LoRA — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-29-doc-to-lora-internalizing-context-as-lor-8dd5ec.mp3
21. AI Post Transformers: DeepSeek-V4 and Practical Million-Token Context — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-25-deepseek-v4-and-practical-million-token-6f4de1.mp3
22. AI Post Transformers: PackKV Lossy Compression for KV Caches — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-04-packkv-lossy-compression-for-kv-caches-b37bce.mp3
Interactive Visualization: δ-mem and Online Memory for LLMs

This episode explores a May 7, 2026 arXiv paper on Lighthouse Attention and asks whether long-context language models can be pretrained cheaply with hierarchical sparse attention, then switched back to standard dense attention late in training without losing dense-model quality. It explains why long-context training is so expensive even with FlashAttention, contrasting dense quadratic attention with sparse and hierarchical schemes that try to narrow which tokens interact. The discussion walks through the paper’s core design: building multi-level pooled query/key/value pyramids, using a gradient-free top-K selector to choose relevant causal subsequences, running ordinary FlashAttention on that smaller set, and scattering the results back. Listeners would find it interesting because it frames the method as a practical systems bet with potentially major implications for 128K- to million-token pretraining, while also stressing that the evidence is still preliminary and far from proving it works at frontier scale.

Sources:
1. Long Context Pre-Training with Lighthouse Attention
https://arxiv.org/pdf/2605.06554
2. H-Transformer-1D: Fast One-Dimensional Hierarchical Attention for Sequences — Zhenhai Zhu, Radu Soricut, 2021
https://scholar.google.com/scholar?q=H-Transformer-1D%3A+Fast+One-Dimensional+Hierarchical+Attention+for+Sequences
3. LongT5: Efficient Text-To-Text Transformer for Long Sequences — Mandy Guo, Joshua Ainslie, David Uthus, Santiago Ontanon, Jianmo Ni, Yun-Hsuan Sung, Yinfei Yang, 2021
https://scholar.google.com/scholar?q=LongT5%3A+Efficient+Text-To-Text+Transformer+for+Long+Sequences
4. HISA: Efficient Hierarchical Indexing for Fine-Grained Sparse Attention — Yufei Xu, Fanxu Meng, Fan Jiang, Yuxuan Wang, Ruijie Zhou, Jiexi Wu, Zhixin Pan, Zhaohui Wang, Xiaojuan Tang, Wenjie Pei, Tongxuan Liu, Di Yin, Xing Sun, Muhan Zhang, 2026
https://scholar.google.com/scholar?q=HISA%3A+Efficient+Hierarchical+Indexing+for+Fine-Grained+Sparse+Attention
5. Long Context Pre-Training with Lighthouse Attention — Bowen Peng, Subho Ghosh, Jeffrey Quesnelle, 2026
https://scholar.google.com/scholar?q=Long+Context+Pre-Training+with+Lighthouse+Attention
6. Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention — Jingyang Yuan, Huazuo Gao, Damai Dai, Junyu Luo, Liang Zhao, Zhengyan Zhang, Zhenda Xie, Y. X. Wei, Lean Wang, Zhiping Xiao, Yuqing Wang, Chong Ruan, Ming Zhang, Wenfeng Liang, Wangding Zeng, 2025
https://scholar.google.com/scholar?q=Native+Sparse+Attention%3A+Hardware-Aligned+and+Natively+Trainable+Sparse+Attention
7. MoBA: Mixture of Block Attention for Long-Context LLMs — Enzhe Lu, Zhejun Jiang, Jingyuan Liu, Yulun Du, Tao Jiang, Chao Hong, Shaowei Liu, Weiran He, Enming Yuan, Yuzhi Wang, Zhiqi Huang, Huan Yuan, Suting Xu, Xinran Xu, Guokun Lai, Yanru Chen, Huabin Zheng, Junjie Yan, Jianlin Su, Yuxin Wu, Neo Y. Zhang, Zhilin Yang, Xinyu Zhou, Mingxing Zhang, Jiezhong Qiu, 2025
https://scholar.google.com/scholar?q=MoBA%3A+Mixture+of+Block+Attention+for+Long-Context+LLMs
8. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness — Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, Christopher Re, 2022
https://scholar.google.com/scholar?q=FlashAttention%3A+Fast+and+Memory-Efficient+Exact+Attention+with+IO-Awareness
9. Every Token Counts: Generalizing 16M Ultra-Long Context in Large Language Models — Xiang Hu, Zhanchao Zhou, Ruiqi Liang, Zehuan Li, Wei Wu, Jianguo Li, 2025
https://scholar.google.com/scholar?q=Every+Token+Counts%3A+Generalizing+16M+Ultra-Long+Context+in+Large+Language+Models
10. Hyperattention: Long-context attention in near-linear time — not verified from snippet, recent (not verified from snippet)
https://scholar.google.com/scholar?q=Hyperattention%3A+Long-context+attention+in+near-linear+time
11. When Does Content-Based Routing Work? Representation Requirements for Selective Attention in Hybrid Sequence Models — not verified from snippet, recent (not verified from snippet)
https://scholar.google.com/scholar?q=When+Does+Content-Based+Routing+Work%3F+Representation+Requirements+for+Selective+Attention+in+Hybrid+Sequence+Models
12. Delta attention: Fast and accurate sparse attention inference by delta correction — not verified from snippet, recent (not verified from snippet)
https://scholar.google.com/scholar?q=Delta+attention%3A+Fast+and+accurate+sparse+attention+inference+by+delta+correction
13. Spargeattention: Accurate and training-free sparse attention accelerating any model inference — not verified from snippet, recent (not verified from snippet)
https://scholar.google.com/scholar?q=Spargeattention%3A+Accurate+and+training-free+sparse+attention+accelerating+any+model+inference
14. Post-training sparse attention with double sparsity — not verified from snippet, recent (not verified from snippet)
https://scholar.google.com/scholar?q=Post-training+sparse+attention+with+double+sparsity
15. AI Post Transformers: Memory Sparse Attention for 100M-Token Scaling — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-07-memory-sparse-attention-for-100m-token-s-377cff.mp3
16. AI Post Transformers: RetrievalAttention for Long-Context LLM Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-17-retrievalattention-for-long-context-llm-ddf566.mp3
17. AI Post Transformers: Gated Linear Attention for Efficient Long Sequences — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-18-gated-linear-attention-for-efficient-lon-c858ab.mp3
18. AI Post Transformers: Mamba-3 for Efficient Sequence Modeling — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-16-mamba-3-for-efficient-sequence-modeling-97a22a.mp3
19. AI Post Transformers: Explicit Information Transmission for Context Compression — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-explicit-information-transmission-for-co-24e3c2.mp3
20. AI Post Transformers: How Induction Heads Emerge in Transformers — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-03-how-induction-heads-emerge-in-transforme-a7bfcb.mp3
Interactive Visualization: Long Context Pre-Training with Lighthouse Attention

This episode explores MiA-Signature, a long-context reasoning method that argues models should approximate a broad, query-driven activation pattern over memory rather than rely on narrow top-k retrieval. It explains how the paper builds a two-stage pipeline: first retrieving a wide pool of potentially relevant context, then compressing that pool into a compact signature of high-level concepts chosen with submodular optimization to maximize coverage and reduce redundancy. The discussion digs into the paper’s central claim that this behaves more like a planning or memory-compression layer than classic RAG, while also questioning whether the real gains come from the signature itself or from the broader retrieval and refinement machinery around it. Listeners would find it interesting because it connects long-context failures, agent memory design, and distributed evidence tracking into a concrete systems debate about what better memory access in LLMs should look like.

Sources:
1. MiA-Signature: Approximating Global Activation for Long-Context Understanding — Yuqing Li, Jiangnan Li, Mo Yu, Zheng Lin, Weiping Wang, Jie Zhou, 2026
http://arxiv.org/abs/2605.06416
2. An Analysis of Approximations for Maximizing Submodular Set Functions—I — George L. Nemhauser, Laurence A. Wolsey, Marshall L. Fisher, 1978
https://scholar.google.com/scholar?q=An+Analysis+of+Approximations+for+Maximizing+Submodular+Set+Functions%E2%80%94I
3. Submodular Function Maximization — Andreas Krause, Daniel Golovin, 2014
https://scholar.google.com/scholar?q=Submodular+Function+Maximization
4. Near-Optimal Sensor Placements in Gaussian Processes: Theory, Efficient Algorithms and Empirical Studies — Andreas Krause, Ajit Singh, Carlos Guestrin, 2008
https://scholar.google.com/scholar?q=Near-Optimal+Sensor+Placements+in+Gaussian+Processes%3A+Theory%2C+Efficient+Algorithms+and+Empirical+Studies
5. A Class of Submodular Functions for Document Summarization — Hui Lin, Jeff Bilmes, 2011
https://scholar.google.com/scholar?q=A+Class+of+Submodular+Functions+for+Document+Summarization
6. Working Memory — Alan D. Baddeley, Graham Hitch, 1974
https://scholar.google.com/scholar?q=Working+Memory
7. The Episodic Buffer: A New Component of Working Memory? — Alan Baddeley, 2000
https://scholar.google.com/scholar?q=The+Episodic+Buffer%3A+A+New+Component+of+Working+Memory%3F
8. Hybrid Computing Using a Neural Network with Dynamic External Memory — Alex Graves, Greg Wayne, Malcolm Reynolds and colleagues, 2016
https://scholar.google.com/scholar?q=Hybrid+Computing+Using+a+Neural+Network+with+Dynamic+External+Memory
9. Transformer-XL: Attentive Language Models beyond a Fixed-Length Context — Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc Le, Ruslan Salakhutdinov, 2019
https://scholar.google.com/scholar?q=Transformer-XL%3A+Attentive+Language+Models+beyond+a+Fixed-Length+Context
10. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks — Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Kuttler, Mike Lewis, Wen-tau Yih, Tim Rocktaschel, Sebastian Riedel, Douwe Kiela, 2020
https://scholar.google.com/scholar?q=Retrieval-Augmented+Generation+for+Knowledge-Intensive+NLP+Tasks
11. What is consciousness, and could machines have it? — Stanislas Dehaene, Hakwan Lau, Sid Kouider, 2017
https://scholar.google.com/scholar?q=What+is+consciousness%2C+and+could+machines+have+it%3F
12. DetectiveQA: Evaluating Long-Context Reasoning on Detective Novels — Zhe Xu, Jiasheng Ye, Xiangyang Liu, Tianxiang Sun, Xiaoran Liu, Qipeng Guo, Linlin Li, Qun Liu, Xuanjing Huang, Xipeng Qiu, 2024
https://scholar.google.com/scholar?q=DetectiveQA%3A+Evaluating+Long-Context+Reasoning+on+Detective+Novels
13. One Thousand and One Pairs: A "novel" challenge for long-context language models — Marzena Karpinska, Katherine Thai, Kyle Lo, Tanya Goyal, Mohit Iyyer, 2024
https://scholar.google.com/scholar?q=One+Thousand+and+One+Pairs%3A+A+%22novel%22+challenge+for+long-context+language+models
14. Stateful Evidence-Driven Retrieval-Augmented Generation with Iterative Reasoning — Qi Dong, Ziheng Lin, Ning Ding, 2026
https://scholar.google.com/scholar?q=Stateful+Evidence-Driven+Retrieval-Augmented+Generation+with+Iterative+Reasoning
15. LongRAG: Enhancing Retrieval-Augmented Generation with Long-Context LLMs — approx. Wang et al., 2024/2025
https://scholar.google.com/scholar?q=LongRAG%3A+Enhancing+Retrieval-Augmented+Generation+with+Long-Context+LLMs
16. Retrieval Augmented Generation or Long-Context LLMs? A Study and Hybrid Approach — approx. Li et al., 2024/2025
https://scholar.google.com/scholar?q=Retrieval+Augmented+Generation+or+Long-Context+LLMs%3F+A+Study+and+Hybrid+Approach
17. Long Context Compression with Activation Beacon — approx. Liu et al., 2024
https://scholar.google.com/scholar?q=Long+Context+Compression+with+Activation+Beacon
18. UniGist: Towards General and Hardware-Aligned Sequence-Level Long Context Compression — approx. authors unclear from snippet, 2024/2025
https://scholar.google.com/scholar?q=UniGist%3A+Towards+General+and+Hardware-Aligned+Sequence-Level+Long+Context+Compression
19. Retaining Key Information under High Compression Ratios: Query-Guided Compressor for LLMs — approx. authors unclear from snippet, 2024/2025
https://scholar.google.com/scholar?q=Retaining+Key+Information+under+High+Compression+Ratios%3A+Query-Guided+Compressor+for+LLMs
20. QEC-LLM: Training Query-Focused Extractive Compression Model with LLM-Guiding for Open Domain QA — approx. authors unclear from snippet, 2024/2025
https://scholar.google.com/scholar?q=QEC-LLM%3A+Training+Query-Focused+Extractive+Compression+Model+with+LLM-Guiding+for+Open+Domain+QA
21. HMT: Hierarchical Memory Transformer for Efficient Long Context Language Processing — approx. authors unclear from snippet, 2024/2025
https://scholar.google.com/scholar?q=HMT%3A+Hierarchical+Memory+Transformer+for+Efficient+Long+Context+Language+Processing
22. UniMem: Towards a Unified View of Long-Context Large Language Models — approx. authors unclear from snippet, 2025
https://scholar.google.com/scholar?q=UniMem%3A+Towards+a+Unified+View+of+Long-Context+Large+Language+Models
23. AI Post Transformers: Explicit Information Transmission for Context Compression — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-explicit-information-transmission-for-co-24e3c2.mp3
24. AI Post Transformers: MEMSEARCHER: Reinforcement Learning for LLM Memory Management — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-memsearcher-reinforcement-learning-for-l-e9ad84.mp3
25. AI Post Transformers: Memory Intelligence Agents for Deep Research — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-09-memory-intelligence-agents-for-deep-rese-cd39e3.mp3
Interactive Visualization: MiA-Signature and Global Activation for Long Context

This episode explores ForkKV, a systems paper on serving multiple LoRA-based agents from one base language model without duplicating massive KV caches for shared context. It explains why ordinary prefix caching breaks once different LoRA adapters change the activations, then walks through the paper’s core idea: split cache state into a large shared base and a small adapter-specific residual, using an operating-system-style copy-on-write model for agent branches. The discussion connects that design to prior work on LoRA, prefix caching, PagedAttention, and disaggregated memory, making the argument that the real win is practical GPU memory efficiency for coding assistants and tool-using agent workflows. Listeners would find it interesting because it frames transformer serving as a memory-management problem and shows how borrowing ideas from Unix process forking could make multi-agent LLM systems far more scalable.

Interactive Visualization: ForkKV for Multi-LoRA Agent Serving
Sources:
1. ForkKV: Scaling Multi-LoRA Agent Serving via Copy-on-Write Disaggregated KV Cache — Shao Wang, Rui Ren, Lin Gui, 2026
http://arxiv.org/abs/2604.06370
2. The UNIX Time-Sharing System — Dennis M. Ritchie and Ken Thompson, 1974
https://scholar.google.com/scholar?q=The+UNIX+Time-Sharing+System
3. Efficient Memory Management for Large Language Model Serving with PagedAttention — Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica, 2023
https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention
4. vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention — Ramya Prabhu, Ajay Nayak, Jayashree Mohan, Ramachandran Ramjee, and Ashish Panwar, 2024
https://scholar.google.com/scholar?q=vAttention%3A+Dynamic+Memory+Management+for+Serving+LLMs+without+PagedAttention
5. ForkKV: Scaling Multi-LoRA Agent Serving via Copy-on-Write Disaggregated KV Cache — Shao Wang, Rui Ren, and Lin Gui, 2026
https://scholar.google.com/scholar?q=ForkKV%3A+Scaling+Multi-LoRA+Agent+Serving+via+Copy-on-Write+Disaggregated+KV+Cache
6. Efficiently Programming Large Language Models using SGLang — Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Jeff Huang, Chuyue Sun, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonzalez, Clark Barrett, and Ying Sheng, 2023
https://scholar.google.com/scholar?q=Efficiently+Programming+Large+Language+Models+using+SGLang
7. DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving — Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, and Hao Zhang, 2024
https://scholar.google.com/scholar?q=DistServe%3A+Disaggregating+Prefill+and+Decoding+for+Goodput-optimized+Large+Language+Model+Serving
8. MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool — Cunchen Hu, Heyang Huang, Junhao Hu, Jiang Xu, Xusheng Chen, Tao Xie, Chenxi Wang, Sa Wang, Yungang Bao, Ninghui Sun, and Yizhou Shan, 2024
https://scholar.google.com/scholar?q=MemServe%3A+Context+Caching+for+Disaggregated+LLM+Serving+with+Elastic+Memory+Pool
9. Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving — Ruoyu Qin, Zheming Li, Weiran He, Mingxing Zhang, Yongwei Wu, Weimin Zheng, and Xinran Xu, 2024
https://scholar.google.com/scholar?q=Mooncake%3A+A+KVCache-centric+Disaggregated+Architecture+for+LLM+Serving
10. LRAgent: efficient kv cache sharing for multi-lora llm agents — H. Jeon, H. Ha, and J. Kim, 2026
https://scholar.google.com/scholar?q=LRAgent%3A+efficient+kv+cache+sharing+for+multi-lora+llm+agents
11. Punica: multi-tenant lora serving — L. Chen, Z. Ye, Y. Wu, D. Zhuo, L. Ceze, and A. Krishnamurthy, 2024
https://scholar.google.com/scholar?q=Punica%3A+multi-tenant+lora+serving
12. S-LoRA: serving thousands of concurrent lora adapters — Y. Sheng, S. Cao, D. Li, C. Hooper, N. Lee, S. Yang, C. Chou, B. Zhu, L. Zheng, K. Keutzer, J. E. Gonzalez, and I. Stoica, 2023
https://scholar.google.com/scholar?q=S-LoRA%3A+serving+thousands+of+concurrent+lora+adapters
13. DLoRA: dynamically orchestrating requests and adapters for LoRA LLM serving — B. Wu, R. Zhu, Z. Zhang, P. Sun, X. Liu, and X. Jin, 2024
https://scholar.google.com/scholar?q=DLoRA%3A+dynamically+orchestrating+requests+and+adapters+for+LoRA+LLM+serving
14. Tokencake: A KV-Cache-centric Serving Framework for LLM-based Multi-Agent Applications — Z. Bian, F. Wu, T. Ma, and Y. Zhuo, 2025
https://scholar.google.com/scholar?q=Tokencake%3A+A+KV-Cache-centric+Serving+Framework+for+LLM-based+Multi-Agent+Applications
15. KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows — Z. Pan, A. Patel, Z. Hu, Y. Shen, Y. Guan, W. Li, L. Qin, Y. Wang, and Y. Ding, 2025
https://scholar.google.com/scholar?q=KVFlow%3A+Efficient+Prefix+Caching+for+Accelerating+LLM-Based+Multi-Agent+Workflows
16. MobiLoRA: Accelerating LoRA-based LLM Inference on Mobile Devices via Context-aware KV Cache Optimization — Borui Li et al., 2025
https://scholar.google.com/scholar?q=MobiLoRA%3A+Accelerating+LoRA-based+LLM+Inference+on+Mobile+Devices+via+Context-aware+KV+Cache+Optimization
17. Efficient Multi-Adapter LLM Serving via Cross-Model KV-Cache Reuse with Activated LoRA — Allison Li, Kristjan Greenewald, Thomas Parnell, Navid Azizan, 2025
https://scholar.google.com/scholar?q=Efficient+Multi-Adapter+LLM+Serving+via+Cross-Model+KV-Cache+Reuse+with+Activated+LoRA
18. KVLink: Accelerating Large Language Models via Efficient KV Cache Reuse — Jingbo Yang et al., 2025
https://scholar.google.com/scholar?q=KVLink%3A+Accelerating+Large+Language+Models+via+Efficient+KV+Cache+Reuse
19. AdaFuse: Accelerating Dynamic Adapter Inference via Token-Level Pre-Gating and Fused Kernel Optimization — Qiyang Li et al., 2026
https://scholar.google.com/scholar?q=AdaFuse%3A+Accelerating+Dynamic+Adapter+Inference+via+Token-Level+Pre-Gating+and+Fused+Kernel+Optimization
20. ChunkKV: Semantic-Preserving KV Cache Compression for Efficient Long-Context LLM Inference — Xiang Liu et al., 2025
https://scholar.google.com/scholar?q=ChunkKV%3A+Semantic-Preserving+KV+Cache+Compression+for+Efficient+Long-Context+LLM+Inference
21. SCBench: A KV Cache-Centric Analysis of Long-Context Methods — Yucheng Li et al., 2025
https://scholar.google.com/scholar?q=SCBench%3A+A+KV+Cache-Centric+Analysis+of+Long-Context+Methods
22. ELORA: Efficient LoRA and KV Cache Management for Multi-LoRA LLM Serving — Jiuchen Shi et al., 2026
https://scholar.google.com/scholar?q=ELORA%3A+Efficient+LoRA+and+KV+Cache+Management+for+Multi-LoRA+LLM+Serving
23. AI Post Transformers: Efficient KV Cache Sharing for Multi-LoRA Agents — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-22-efficient-kv-cache-sharing-for-multi-lor-afda05.mp3
24. AI Post Transformers: TokenDance for Multi-Agent KV Cache Sharing — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-22-tokendance-for-multi-agent-kv-cache-shar-aa9b99.mp3
25. AI Post Transformers: ContiguousKV for Faster LLM Prefill KV Reuse — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-20-contiguouskv-for-faster-llm-prefill-kv-r-59f545.mp3
26. AI Post Transformers: Speculative Decoding in Real vLLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-speculative-decoding-in-real-vllm-servin-6f4e2b.mp3
27. AI Post Transformers: Why LLM Serving Needs Mathematical Optimization — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-why-llm-serving-needs-mathematical-optim-647fc6.mp3
Interactive Visualization: ForkKV for Multi-LoRA Agent Serving

This episode explores ELF: Embedded Language Flows, a continuous-time diffusion language model that stays in embedding space until the final decoding step instead of repeatedly snapping back to discrete tokens during generation. It explains how that design lets the model borrow flow-matching and guidance techniques from image diffusion, while arguing that earlier continuous text models may have underperformed because of token-level constraints rather than any fundamental weakness. The discussion highlights reported results on OpenWebText, where a 105M-parameter ELF model achieves better generative perplexity than 170M baselines with far fewer training tokens and fewer sampling steps, while also extending to translation and summarization. It also digs into the main caveat: whether the gains really come from late discretization and continuous-time modeling, or from a bundle of confounded training and inference choices, making the episode interesting both as a technical walkthrough and as a skeptical evaluation of a bold research claim.

Interactive Visualization: ELF and Continuous Language Diffusion
Sources:
1. ELF and Continuous Language Diffusion
https://arxiv.org/pdf/2605.10938
2. Structured Denoising Diffusion Models in Discrete State-Spaces — Jacob Austin, Daniel D. Johnson, Jonathan Ho, Daniel Tarlow, Rianne van den Berg, 2021
https://scholar.google.com/scholar?q=Structured+Denoising+Diffusion+Models+in+Discrete+State-Spaces
3. Diffusion-LM Improves Controllable Text Generation — Xiang Lisa Li, John Thickstun, Ishaan Gulrajani, Percy Liang, Tatsunori B. Hashimoto, 2022
https://scholar.google.com/scholar?q=Diffusion-LM+Improves+Controllable+Text+Generation
4. Simple and Effective Masked Diffusion Language Models — Subham Sekhar Sahoo, Marianne Arriola, Yair Schiff, Aaron Gokaslan, Edgar Marroquin, Justin T. Chiu, Alexander M. Rush, Volodymyr Kuleshov, 2024
https://scholar.google.com/scholar?q=Simple+and+Effective+Masked+Diffusion+Language+Models
5. Large Language Diffusion Models — Shen Nie, Fengqi Zhu, Zebin You, Xiaolu Zhang, Jingyang Ou, Jun Hu, Jun Zhou, Yankai Lin, Ji-Rong Wen, Chongxuan Li, 2025
https://scholar.google.com/scholar?q=Large+Language+Diffusion+Models
6. Flow Matching for Generative Modeling — Yaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel, Matthew Le, 2023
https://scholar.google.com/scholar?q=Flow+Matching+for+Generative+Modeling
7. Improving and Generalizing Flow-Based Generative Models with Minibatch Optimal Transport — Alexander Tong, Kilian Fatras, Nikolay Malkin, Guillaume Huguet, Yanlei Zhang, Jarrid Rector-Brooks, Guy Wolf, Yoshua Bengio, 2024
https://scholar.google.com/scholar?q=Improving+and+Generalizing+Flow-Based+Generative+Models+with+Minibatch+Optimal+Transport
8. Scaling Rectified Flow Transformers for High-Resolution Image Synthesis — Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Muller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, Dustin Podell, Tim Dockhorn, Zion English, Kyle Lacey, Alex Goodwin, Yannik Marek, Robin Rombach, 2024
https://scholar.google.com/scholar?q=Scaling+Rectified+Flow+Transformers+for+High-Resolution+Image+Synthesis
9. Discrete Flow Matching — Itai Gat, Tal Remez, Neta Shaul, Felix Kreuk, Ricky T. Q. Chen, Gabriel Synnaeve, Yossi Adi, Yaron Lipman, 2024
https://scholar.google.com/scholar?q=Discrete+Flow+Matching
10. Self-conditioned Embedding Diffusion for Text Generation — Robin Strudel, Corentin Tallec, Florent Altche, Yilun Du, Yaroslav Ganin, Arthur Mensch, Will Grathwohl, Nikolay Savinov, Sander Dieleman, Laurent Sifre, Remi Leblond, 2022
https://scholar.google.com/scholar?q=Self-conditioned+Embedding+Diffusion+for+Text+Generation
11. Difformer: Empowering Diffusion Models on the Embedding Space for Text Generation — Zhujin Gao, Junliang Guo, Xu Tan, Yongxin Zhu, Fang Zhang, Jiang Bian, Linli Xu, 2022
https://scholar.google.com/scholar?q=Difformer%3A+Empowering+Diffusion+Models+on+the+Embedding+Space+for+Text+Generation
12. LangFlow: Continuous Diffusion Rivals Discrete in Language Modeling — Yuxin Chen, Chumeng Liang, Hangke Sui, Ruihan Guo, Chaoran Cheng, Jiaxuan You, Ge Liu, 2026
https://scholar.google.com/scholar?q=LangFlow%3A+Continuous+Diffusion+Rivals+Discrete+in+Language+Modeling
13. Classifier-Free Diffusion Guidance — Jonathan Ho, Tim Salimans, 2021
https://scholar.google.com/scholar?q=Classifier-Free+Diffusion+Guidance
14. GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models — Alex Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, Mark Chen, 2021
https://scholar.google.com/scholar?q=GLIDE%3A+Towards+Photorealistic+Image+Generation+and+Editing+with+Text-Guided+Diffusion+Models
15. Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding — Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S. Sara Mahdavi, Rapha Gontijo Lopes, Tim Salimans, Jonathan Ho, David J. Fleet, Mohammad Norouzi, 2022
https://scholar.google.com/scholar?q=Photorealistic+Text-to-Image+Diffusion+Models+with+Deep+Language+Understanding
16. High-Resolution Image Synthesis with Latent Diffusion Models — Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, Bjorn Ommer, 2021
https://scholar.google.com/scholar?q=High-Resolution+Image+Synthesis+with+Latent+Diffusion+Models
17. MDLM: Masked Diffusion Language Models — likely the MDLM authors cited as [56] in the paper, 2024
https://scholar.google.com/scholar?q=MDLM%3A+Masked+Diffusion+Language+Models
18. Duo — likely the Duo authors cited as [57] in the paper, 2025
https://scholar.google.com/scholar?q=Duo
19. Latent Diffusion Models — Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, Bjorn Ommer, 2022
https://scholar.google.com/scholar?q=Latent+Diffusion+Models
20. LangFlow — the LangFlow authors cited as [10] in the paper, 2026
https://scholar.google.com/scholar?q=LangFlow
21. FLM — the FLM authors cited as [30] in the paper, 2026
https://scholar.google.com/scholar?q=FLM
22. Discrete Diffusion Modeling by Estimating the Ratios of the Data Distribution — Aaron Lou, Chenlin Meng, Stefano Ermon, 2024
https://scholar.google.com/scholar?q=Discrete+Diffusion+Modeling+by+Estimating+the+Ratios+of+the+Data+Distribution
23. Scaling Behavior of Discrete Diffusion Language Models — Dimitri von Rutte, Janis Fluri, Omead Pooladzandi, Bernhard Scholkopf, Thomas Hofmann, Antonio Orvieto, 2025
https://scholar.google.com/scholar?q=Scaling+Behavior+of+Discrete+Diffusion+Language+Models
24. Coevolutionary Continuous Discrete Diffusion: Make Your Diffusion Language Model a Latent Reasoner — Cai Zhou, Chenxiao Yang, Yi Hu, Chenyu Wang, Chubin Zhang, Muhan Zhang, Lester Mackey, Tommi Jaakkola, Stephen Bates, Dinghuai Zhang, 2025
https://scholar.google.com/scholar?q=Coevolutionary+Continuous+Discrete+Diffusion%3A+Make+Your+Diffusion+Language+Model+a+Latent+Reasoner
25. Stay on Topic with Classifier-Free Guidance — Guillaume Sanchez, Honglu Fan, Alexander Spangher, Elad Levi, Pawan Sasanka Ammanamanchi, Stella Biderman, 2023
https://scholar.google.com/scholar?q=Stay+on+Topic+with+Classifier-Free+Guidance
26. Adaptive Classifier-Free Guidance via Dynamic Low-Confidence Masking — Pengxiang Li, Shilin Yan, Joey Tsai, Renrui Zhang, Ruichuan An, Ziyu Guo, Xiaowei Gao, 2025
https://scholar.google.com/scholar?q=Adaptive+Classifier-Free+Guidance+via+Dynamic+Low-Confidence+Masking
27. Studying Classifier(-Free) Guidance From a Classifier-Centric Perspective — Xiaoming Zhao, Alexander G. Schwing, 2025
https://scholar.google.com/scholar?q=Studying+Classifier%28-Free%29+Guidance+From+a+Classifier-Centric+Perspective
28. DEPT: Decoupled Embeddings for Pre-training Language Models — Alex Iacob, Lorenzo Sani, Meghdad Kurmanji, William F. Shen, Xinchi Qiu, Dongqi Cai, Yan Gao, Nicholas D. Lane, 2024
https://scholar.google.com/scholar?q=DEPT%3A+Decoupled+Embeddings+for+Pre-training+Language+Models
29. AI Post Transformers: Generative Modeling via Drifting in One Step — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-generative-modeling-via-drifting-in-one-671da0.mp3
30. AI Post Transformers: VL-JEPA for Vision-Language Semantic Prediction — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-12-vl-jepa-for-vision-language-semantic-pre-69c9f4.mp3
31. AI Post Transformers: Why Transformers Fail at Counting — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-08-why-transformers-fail-at-counting-137924.mp3
32. AI Post Transformers: Mamba-3 for Efficient Sequence Modeling — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-16-mamba-3-for-efficient-sequence-modeling-97a22a.mp3
Interactive Visualization: ELF and Continuous Language Diffusion

This episode explores a paper on test-time scaling that asks whether an LLM agent can automatically discover better inference-time control policies than the hand-built heuristics researchers usually rely on. It explains the core search framework in concrete terms: a controller decides when to branch, continue, probe, prune, or stop, balancing reasoning depth, breadth, and compute budget rather than simply generating more tokens. The discussion highlights the paper’s main technical argument that offline replay over logged reasoning traces, combined with a compact controller parameterization and detailed execution feedback, makes policy discovery cheap enough to be practical. Listeners would find it interesting because it connects abstract ideas about reasoning agents to a very specific claim: smarter inference may come not from larger models, but from better learned strategies for spending compute.

Interactive Visualization: Agentic Discovery for Test-Time Scaling
Sources:
1. LLMs Improving LLMs: Agentic Discovery for Test-Time Scaling — Tong Zheng, Haolin Liu, Chengsong Huang, Huiwen Bao, Sheng Zhang, Rui Liu, Runpeng Dai, Ruibo Chen, Chenxi Liu, Tianyi Xiong, Xidong Wu, Hongming Zhang, Heng Huang, 2026
http://arxiv.org/abs/2605.08083
2. Supervisory Control of a Class of Discrete Event Systems — Peter J. Ramadge, Walter M. Wonham, 1987
https://scholar.google.com/scholar?q=Supervisory+Control+of+a+Class+of+Discrete+Event+Systems
3. On the Synthesis of a Reactive Module — Amir Pnueli, Roni Rosner, 1989
https://scholar.google.com/scholar?q=On+the+Synthesis+of+a+Reactive+Module
4. Formal Methods for Control Synthesis: An Optimization Perspective — Calin Belta, Sadra Sadraddini, 2019
https://scholar.google.com/scholar?q=Formal+Methods+for+Control+Synthesis%3A+An+Optimization+Perspective
5. Formal Synthesis of Controllers for Safety-Critical Autonomous Systems: Developments and Challenges — Xiang Yin, Bingzhao Gao, Xiao Yu, 2024
https://scholar.google.com/scholar?q=Formal+Synthesis+of+Controllers+for+Safety-Critical+Autonomous+Systems%3A+Developments+and+Challenges
6. Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems — Sergey Levine, Aviral Kumar, George Tucker, Justin Fu, 2020
https://scholar.google.com/scholar?q=Offline+Reinforcement+Learning%3A+Tutorial%2C+Review%2C+and+Perspectives+on+Open+Problems
7. Off-Policy Deep Reinforcement Learning without Exploration — Scott Fujimoto, David Meger, Doina Precup, 2019
https://scholar.google.com/scholar?q=Off-Policy+Deep+Reinforcement+Learning+without+Exploration
8. Conservative Q-Learning for Offline Reinforcement Learning — Aviral Kumar, Aurick Zhou, George Tucker, Sergey Levine, 2020
https://scholar.google.com/scholar?q=Conservative+Q-Learning+for+Offline+Reinforcement+Learning
9. D4RL: Datasets for Deep Data-Driven Reinforcement Learning — Justin Fu, Aviral Kumar, Ofir Nachum, George Tucker, Sergey Levine, 2020
https://scholar.google.com/scholar?q=D4RL%3A+Datasets+for+Deep+Data-Driven+Reinforcement+Learning
10. Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters — Charlie Snell, Jaehoon Lee, Kelvin Xu, Aviral Kumar, 2024
https://scholar.google.com/scholar?q=Scaling+LLM+Test-Time+Compute+Optimally+can+be+More+Effective+than+Scaling+Model+Parameters
11. Parallel-Probe: Towards Efficient Parallel Thinking via 2D Probing — Tong Zheng, Chengsong Huang, Runpeng Dai, Yun He, Rui Liu, Xin Ni, Huiwen Bao, Kaishen Wang, Hongtu Zhu, Jiaxin Huang, Furong Huang, Heng Huang, 2026
https://scholar.google.com/scholar?q=Parallel-Probe%3A+Towards+Efficient+Parallel+Thinking+via+2D+Probing
12. Scaling Test-time Compute for LLM Agents — King Zhu, Hanhao Li, Siwei Wu, Tianshun Xing, Dehua Ma, Xiangru Tang, Minghao Liu, Jian Yang, Jiaheng Liu, Yuchen Eleanor Jiang, Changwang Zhang, Chenghua Lin, Jun Wang, Ge Zhang, Wangchunshu Zhou, 2025
https://scholar.google.com/scholar?q=Scaling+Test-time+Compute+for+LLM+Agents
13. Timely Machine: Awareness of Time Makes Test-Time Scaling Agentic — Yichuan Ma, Linyang Li, Yongkang Chen, Peiji Li, Xiaozhe Li, Qipeng Guo, Dahua Lin, Kai Chen, 2026
https://scholar.google.com/scholar?q=Timely+Machine%3A+Awareness+of+Time+Makes+Test-Time+Scaling+Agentic
14. Predicting and improving test-time scaling laws via reward tail-guided search — Muheng Li, Jian Qian, Wenlong Mou, 2026
https://scholar.google.com/scholar?q=Predicting+and+improving+test-time+scaling+laws+via+reward+tail-guided+search
15. TTRL: Test-Time Reinforcement Learning — Yuxin Zuo, Kaiyan Zhang, Shang Qu, Li Sheng, et al., 2025
https://scholar.google.com/scholar?q=TTRL%3A+Test-Time+Reinforcement+Learning
16. CTRLS: Chain-of-Thought Reasoning via Latent State-Transition — Junda Wu, Yuxin Xiong, Xintong Li, Zhengmian Hu, Tong Yu, Rui Wang, Xiang Chen, Jingbo Shang, Julian McAuley, 2025
https://scholar.google.com/scholar?q=CTRLS%3A+Chain-of-Thought+Reasoning+via+Latent+State-Transition
17. Reasoning Models Know When They're Right: Probing Hidden States for Self-Verification — Anqi Zhang, Yulin Chen, Jane Pan, Chen Zhao, Aurojit Panda, Jinyang Li, He He, 2025
https://scholar.google.com/scholar?q=Reasoning+Models+Know+When+They%27re+Right%3A+Probing+Hidden+States+for+Self-Verification
18. Calibrated Reasoning: An Explanatory Verifier for Dynamic and Efficient Problem-Solving — Anisha Garg, Engin Tekin, Yash More, David Bick, Nishit Neema, Ganesh Venkatesh, 2025
https://scholar.google.com/scholar?q=Calibrated+Reasoning%3A+An+Explanatory+Verifier+for+Dynamic+and+Efficient+Problem-Solving
19. Contextual Drag: How Errors in the Context Affect LLM Reasoning — Yun Cheng, Xingyu Zhu, Haoyu Zhao, Sanjeev Arora, 2026
https://scholar.google.com/scholar?q=Contextual+Drag%3A+How+Errors+in+the+Context+Affect+LLM+Reasoning
20. Reinforcement Learning Teachers of Test Time Scaling — Edoardo Cetin, Tianyu Zhao, Yujin Tang, 2025
https://scholar.google.com/scholar?q=Reinforcement+Learning+Teachers+of+Test+Time+Scaling
21. AI Post Transformers: Test-time Scaling for Multi-Agent Collaborative Reasoning — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-22-test-time-scaling-for-multi-agent-collab-082570.mp3
22. AI Post Transformers: ASI-Evolve for Data, Architectures, and RL — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-05-asi-evolve-for-data-architectures-and-rl-197b2b.mp3
23. AI Post Transformers: Generalist Reward Modeling with Inference-Time Scaling — Hal Turing & Dr. Ada Shannon, Tue,
https://podcast.do-not-panic.com/episodes/generalist-reward-modeling-with-inference-time-scaling/
24. AI Post Transformers: Reasoning Theater and Unfaithful Chain-of-Thought — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-reasoning-theater-and-unfaithful-chain-o-a4507e.mp3
25. AI Post Transformers: SGLang for Faster Structured LLM Programs — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-06-sglang-for-faster-structured-llm-program-c59f1c.mp3
Interactive Visualization: Agentic Discovery for Test-Time Scaling

This episode explores TIDE, a transformer variant that lets every layer re-access the original token identity instead of relying entirely on contextual hidden states to preserve that information. It explains how this design targets rare-token failures and “contextual collapse,” where tokens appearing in similar contexts can become too hard for the model to distinguish, especially in scientific, biomedical, or code-heavy text. The discussion walks through TIDE’s mechanism of token-indexed memory tables and layer-wise routing, framing it as a lightweight side channel rather than retrieval or mixture-of-experts. Listeners would find it interesting because it gets at a basic but rarely questioned assumption in modern transformers and asks whether a small architectural change could improve how models handle the long tail of language.

Interactive Visualization: TIDE and the Rare Token Problem
Sources:
1. TIDE: Every Layer Knows the Token Beneath the Context — Ajay Jaiswal, Lauren Hannah, Han-Byul Kim, Duc Hoang, Mehrdad Farajtabar, Minsik Cho, 2026
http://arxiv.org/abs/2605.06216
2. Attention Is All You Need — Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones and others, 2017
https://scholar.google.com/scholar?q=Attention+Is+All+You+Need
3. Transformer Feed-Forward Layers Are Key-Value Memories — Mor Geva, Roei Schuster, Jonathan Berant, Omer Levy, 2021
https://scholar.google.com/scholar?q=Transformer+Feed-Forward+Layers+Are+Key-Value+Memories
4. Locating and Editing Factual Associations in GPT — Kevin Meng, David Bau, Alex Andonian, Yonatan Belinkov, 2022
https://scholar.google.com/scholar?q=Locating+and+Editing+Factual+Associations+in+GPT
5. TIDE: Every Layer Knows the Token Beneath the Context — Ajay Jaiswal, Lauren Hannah, Han-Byul Kim, Duc Hoang, Mehrdad Farajtabar, Minsik Cho, 2026
https://scholar.google.com/scholar?q=TIDE%3A+Every+Layer+Knows+the+Token+Beneath+the+Context
6. Adaptive Input Representations for Neural Language Modeling — Alexei Baevski, Michael Auli, 2018
https://scholar.google.com/scholar?q=Adaptive+Input+Representations+for+Neural+Language+Modeling
7. CANINE: Pre-training an Efficient Tokenization-Free Encoder for Language Representation — Jonathan H. Clark, Dan Garrette, Iulia Turc, John Wieting, 2021
https://scholar.google.com/scholar?q=CANINE%3A+Pre-training+an+Efficient+Tokenization-Free+Encoder+for+Language+Representation
8. ByT5: Towards a token-free future with pre-trained byte-to-byte models — Linting Xue, Aditya Barua, Noah Constant, Rami Al-Rfou, Sharan Narang, Mihir Kale, Adam Roberts, Colin Raffel, 2021
https://scholar.google.com/scholar?q=ByT5%3A+Towards+a+token-free+future+with+pre-trained+byte-to-byte+models
9. XLM-V: Overcoming the Vocabulary Bottleneck in Multilingual Masked Language Models — Davis Liang, Hila Gonen, Yuning Mao, Rui Hou, Naman Goyal, Marjan Ghazvininejad, Luke Zettlemoyer, Madian Khabsa, 2023
https://scholar.google.com/scholar?q=XLM-V%3A+Overcoming+the+Vocabulary+Bottleneck+in+Multilingual+Masked+Language+Models
10. Knowledge Neurons in Pretrained Transformers — Damai Dai, Li Dong, Yaru Hao, Zhifang Sui, Baobao Chang, Furu Wei, 2022
https://scholar.google.com/scholar?q=Knowledge+Neurons+in+Pretrained+Transformers
11. Improving Language Models by Retrieving from Trillions of Tokens — Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford and colleagues, 2022
https://scholar.google.com/scholar?q=Improving+Language+Models+by+Retrieving+from+Trillions+of+Tokens
12. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks — Patrick Lewis, Ethan Perez, Aleksandara Piktus, Fabio Petroni, Vladimir Karpukhin and colleagues, 2020
https://scholar.google.com/scholar?q=Retrieval-Augmented+Generation+for+Knowledge-Intensive+NLP+Tasks
13. Reconsidering Degeneration of Token Embeddings with Definitions for Encoder-Based Pre-Trained Language Models — approx. recent NLP authors; exact list unclear from snippet, recent, likely 2020s
https://scholar.google.com/scholar?q=Reconsidering+Degeneration+of+Token+Embeddings+with+Definitions+for+Encoder-Based+Pre-Trained+Language+Models
14. MemoryLLM: Plug-n-Play Interpretable Feed-Forward Memory for Transformers — approx. recent LLM/memory-systems authors; exact list unclear from snippet, recent, likely 2020s
https://scholar.google.com/scholar?q=MemoryLLM%3A+Plug-n-Play+Interpretable+Feed-Forward+Memory+for+Transformers
15. The FFN as a Key-Value Memory: Functional Specialization in Transformer Computation — approx. recent mechanistic-interpretability authors; exact list unclear from snippet, recent, likely 2020s
https://scholar.google.com/scholar?q=The+FFN+as+a+Key-Value+Memory%3A+Functional+Specialization+in+Transformer+Computation
16. MrT5: Dynamic Token Merging for Efficient Byte-Level Language Models — approx. recent efficiency/LM authors; exact list unclear from snippet, recent, likely 2020s
https://scholar.google.com/scholar?q=MrT5%3A+Dynamic+Token+Merging+for+Efficient+Byte-Level+Language+Models
17. AI Post Transformers: Muon Is Scalable for LLM Training — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-25-muon-is-scalable-for-llm-training-587ed8.mp3
18. AI Post Transformers: Nemotron 3 Super Hybrid Mamba-Transformer MoE — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-19-nemotron-3-super-hybrid-mamba-transforme-31ac75.mp3
19. AI Post Transformers: Gated Linear Attention for Efficient Long Sequences — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-18-gated-linear-attention-for-efficient-lon-c858ab.mp3
20. AI Post Transformers: Parallelizing DeltaNet Linear Transformers over Sequence Length — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-18-parallelizing-deltanet-linear-transforme-2d0377.mp3
21. AI Post Transformers: Long Short-Term Memory and Vanishing Gradients — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-19-long-short-term-memory-and-vanishing-gra-72448c.mp3
Interactive Visualization: TIDE and the Rare Token Problem

This episode explores a 2026 paper arguing that the real trust problem in language models is not error alone, but confident error, and that improving trust may depend more on metacognition than on simply scaling up knowledge. It unpacks key distinctions such as knowledge boundaries, calibration, discrimination, and the gap between intrinsic uncertainty and the uncertainty a model expresses in words, using factoid question answering as a clean test bed where correctness is measurable. The discussion also situates the paper within prior work on self-knowledge, verbalized uncertainty, and self-correction, while stressing that many apparent factuality gains may come from expanded knowledge or external tools rather than genuine awareness of limits. A listener would find it interesting because it reframes hallucinations as a trust and decision-making problem, and offers a sharper way to judge whether AI systems actually know when they should hedge, abstain, or seek evidence.

Sources:
1. Hallucinations Undermine Trust; Metacognition is a Way Forward — Gal Yona, Mor Geva, Yossi Matias, 2026
http://arxiv.org/abs/2605.01428
2. Language Models (Mostly) Know What They Know — Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Ethan Perez, Deep Ganguli, Dario Amodei, Jack Clark, Jared Kaplan and collaborators, 2022
https://scholar.google.com/scholar?q=Language+Models+%28Mostly%29+Know+What+They+Know
3. Teaching Models to Express Their Uncertainty in Words — Stephanie Lin, Jacob Hilton, Owain Evans, 2022
https://scholar.google.com/scholar?q=Teaching+Models+to+Express+Their+Uncertainty+in+Words
4. What Large Language Models Know and What People Think They Know — Mark Steyvers, Heliodoro Tejeda, Aakriti Kumar, Catarina Belem, Sheer Karny, Xinyue Hu, Lukas W. Mayer, Padhraic Smyth, 2025
https://scholar.google.com/scholar?q=What+Large+Language+Models+Know+and+What+People+Think+They+Know
5. When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs — Ryo Kamoi, Yusen Zhang, Nan Zhang, Jiawei Han, Rui Zhang, 2024
https://scholar.google.com/scholar?q=When+Can+LLMs+Actually+Correct+Their+Own+Mistakes%3F+A+Critical+Survey+of+Self-Correction+of+LLMs
6. Can LLMs Express Their Uncertainty in Their Generated Responses? — Gal Yona, Roee Aharoni, Mor Geva, Yossi Matias, 2024
https://scholar.google.com/scholar?q=Can+LLMs+Express+Their+Uncertainty+in+Their+Generated+Responses%3F
7. Faithful or Fluent? Evaluating Natural Language Explanations of Uncertainty — Ghafouri et al., 2024
https://scholar.google.com/scholar?q=Faithful+or+Fluent%3F+Evaluating+Natural+Language+Explanations+of+Uncertainty
8. TruthfulQA: Measuring How Models Mimic Human Falsehoods — Stephanie Lin, Jacob Hilton, Owain Evans, 2022
https://scholar.google.com/scholar?q=TruthfulQA%3A+Measuring+How+Models+Mimic+Human+Falsehoods
9. Survey of Hallucination in Natural Language Generation — Ziwei Ji, et al., 2023
https://scholar.google.com/scholar?q=Survey+of+Hallucination+in+Natural+Language+Generation
10. The Geometry of Truth: Emergent Linear Structure in LLM Representations of Factuality — Marks and Tegmark, 2023
https://scholar.google.com/scholar?q=The+Geometry+of+Truth%3A+Emergent+Linear+Structure+in+LLM+Representations+of+Factuality
11. The Internal State of an LLM Knows When It's Lying — Levinstein and Herrmann, 2023
https://scholar.google.com/scholar?q=The+Internal+State+of+an+LLM+Knows+When+It%27s+Lying
12. Calibrating Verbal Uncertainty as a Linear Feature to Reduce Hallucinations — Ziwei Ji, Lei Yu, Yeskendir Koishekenov, Yejin Bang, Anthony Hartshorn, Alan Schelten, Cheng Zhang, Pascale Fung, Nicola Cancedda, 2025
https://scholar.google.com/scholar?q=Calibrating+Verbal+Uncertainty+as+a+Linear+Feature+to+Reduce+Hallucinations
13. Calibrating the Voice of Doubt: How LLMs Diverge from Humans in Verbal Uncertainty — Jinhao Duan, Zicheng Liu, Zijie Liu, Kaidi Xu, Tianlong Chen, 2025
https://scholar.google.com/scholar?q=Calibrating+the+Voice+of+Doubt%3A+How+LLMs+Diverge+from+Humans+in+Verbal+Uncertainty
14. More Is Not Better: Visual Uncertainty Cues and the Fragility of Trust Calibration in LLM-Assisted Decision Making — authors not recovered from snippet, 2026
https://scholar.google.com/scholar?q=More+Is+Not+Better%3A+Visual+Uncertainty+Cues+and+the+Fragility+of+Trust+Calibration+in+LLM-Assisted+Decision+Making
15. Do Large Language Models Know What They Don't Know? — Zhangyue Yin, Qiushi Sun, Qipeng Guo, Jiawen Wu, Xipeng Qiu, Xuanjing Huang, 2023
https://scholar.google.com/scholar?q=Do+Large+Language+Models+Know+What+They+Don%27t+Know%3F
16. KnowRL: Teaching Language Models to Know What They Know — Sahil Kale, Devendra Singh Dhami, 2025
https://scholar.google.com/scholar?q=KnowRL%3A+Teaching+Language+Models+to+Know+What+They+Know
17. What Models Know, How Well They Know It: Knowledge-Weighted Fine-Tuning for Learning When to Say "I Don't Know" — Joosung Lee, Hwiyeol Jo, Donghyeon Ko, Kyubyung Chae, Cheonbok Park, Jeonghoon Kim, 2026
https://scholar.google.com/scholar?q=What+Models+Know%2C+How+Well+They+Know+It%3A+Knowledge-Weighted+Fine-Tuning+for+Learning+When+to+Say+%22I+Don%27t+Know%22
18. Selective-LAMA: Selective Prediction for Confidence-Aware Evaluation of Language Models — Hiyori Yoshikawa, Naoaki Okazaki, 2023
https://scholar.google.com/scholar?q=Selective-LAMA%3A+Selective+Prediction+for+Confidence-Aware+Evaluation+of+Language+Models
19. Selective Generation for Controllable Language Models — Minjae Lee, Kyungmin Kim, Taesoo Kim, Sangdon Park, 2024
https://scholar.google.com/scholar?q=Selective+Generation+for+Controllable+Language+Models
20. Dynamic Uncertainty Ranking: Enhancing Retrieval-Augmented In-Context Learning for Long-Tail Knowledge in LLMs — Shuyang Yu, Runxue Bao, Parminder Bhatia, Taha Kass-Hout, Jiayu Zhou, Cao Xiao, 2025
https://scholar.google.com/scholar?q=Dynamic+Uncertainty+Ranking%3A+Enhancing+Retrieval-Augmented+In-Context+Learning+for+Long-Tail+Knowledge+in+LLMs
21. UncertaintyRAG: Span-Level Uncertainty Enhanced Long-Context Modeling for Retrieval-Augmented Generation — Zixuan Li, Jing Xiong, Fanghua Ye, Chuanyang Zheng, Xun Wu, Jianqiao Lu, Zhongwei Wan, Xiaodan Liang, Chengming Li, Zhenan Sun, Lingpeng Kong, Ngai Wong, 2024
https://scholar.google.com/scholar?q=UncertaintyRAG%3A+Span-Level+Uncertainty+Enhanced+Long-Context+Modeling+for+Retrieval-Augmented+Generation
22. AI Post Transformers: Can LLMs Judge Their Own Capabilities? — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-01-can-llms-judge-their-own-capabilities-d78fed.mp3
23. AI Post Transformers: Do Language Models Know Their Limits — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-01-do-language-models-know-their-limits-48e444.mp3
24. AI Post Transformers: Teaching Language Models to Verbalize Uncertainty — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-01-teaching-language-models-to-verbalize-un-a1d774.mp3
25. AI Post Transformers: Experimental Comparison of Agentic and Enhanced RAG — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-14-experimental-comparison-of-agentic-and-e-37d8bc.mp3
Interactive Visualization: Metacognition Against Confident Hallucinations

This episode explores whether CXL memory expansion has finally become practical for hyperscale production, using the 2025 Vistara system as a case study. It explains the core ideas behind CXL, tiered memory, and memory disaggregation, then argues that the real comparison is not against swap but against transparent page placement that keeps hot data in local DRAM and colder pages in a slower expanded tier. The discussion highlights Vistara’s full-stack design, from a custom low-latency ASIC and Linux support to workload-specific tuning on a production server with 768 GB of local DDR5 and 256 GB of CXL-attached DDR4. Listeners would find it interesting because the episode moves past industry hype and examines the concrete tradeoffs around latency, bandwidth, operational complexity, and whether memory can finally be managed as a flexible datacenter resource rather than a fixed property of a single machine.

Interactive Visualization: Vistara Brings CXL Memory to Hyperscale
Sources:
1. Vistara Brings CXL Memory to Hyperscale
https://aisystemcodesign.github.io/papers/isca26/vistara_camera_ready.pdf
2. Software-Defined Far Memory in Warehouse-Scale Computers — H. Andres Lagar-Cavilla, Junwhan Ahn, Suleiman Souhlal, Neha Agarwal, Junaid Shahid, Greg Thelen, Parthasarathy Ranganathan, and others, 2019
https://scholar.google.com/scholar?q=Software-Defined+Far+Memory+in+Warehouse-Scale+Computers
3. TMO: Transparent Memory Offloading in Datacenters — Johannes Weiner, Niket Agarwal, Dan Schatzberg, Leon Yang, Hao Wang, Blaise Sanouillet, Bikash Sharma, Tejun Heo, Mayank Jain, Chunqiang Tang, Dimitrios Skarlatos, 2022
https://scholar.google.com/scholar?q=TMO%3A+Transparent+Memory+Offloading+in+Datacenters
4. Pond: CXL-Based Memory Pooling Systems for Cloud Platforms — Huaicheng Li, Daniel S. Berger, Stanko Novakovic, Lisa Hsu, Dan Ernst, Pantea Zardoshti, Monish Shah, Samir Rajadnya, Scott Lee, Ishwar Agarwal, Mark D. Hill, Marcus Fontoura, Ricardo Bianchini, 2023
https://scholar.google.com/scholar?q=Pond%3A+CXL-Based+Memory+Pooling+Systems+for+Cloud+Platforms
5. TPP: Transparent Page Placement for CXL-Enabled Tiered-Memory — Hasan Al Maruf, Hao Wang, Abhishek Dhanotia, Johannes Weiner, Niket Agarwal, Pallab Bhattacharya, Chris Petersen, Mosharaf Chowdhury, Shobhit Kanaujia, Prakash Chauhan, 2023
https://scholar.google.com/scholar?q=TPP%3A+Transparent+Page+Placement+for+CXL-Enabled+Tiered-Memory
6. Managing Memory Tiers with CXL in Virtualized Environments — Yuhong Zhong, Daniel S. Berger, Carl Waldspurger, Richard Wee, Ishan Agarwal, Raghav Agarwal, Fred Hady, K. Kumar, Mark D. Hill, Mosharaf Chowdhury, Ahmed Cidon, 2024
https://scholar.google.com/scholar?q=Managing+Memory+Tiers+with+CXL+in+Virtualized+Environments
7. Demystifying CXL Memory with Genuine CXL-Ready Systems and Devices — Yongjun Sun, Ye Yuan, Ziming Yu, Ryan Kuper, Chao Song, Jiyong Huang, Honggyu Ji, Saurabh Agarwal, Jingwen Lou, Inhwan Jeong, Rui Wang, Joonho H. Ahn, Tianyin Xu, Nam Sung Kim, 2023
https://scholar.google.com/scholar?q=Demystifying+CXL+Memory+with+Genuine+CXL-Ready+Systems+and+Devices
8. M5: Mastering Page Migration and Memory Management for CXL-based Tiered Memory Systems — Yongjun Sun, Jiyoon Kim, Ziming Yu, Joonwon Zhang, Sangho Chai, Minjae J. Kim, Hyeonsu Nam, Jihye Park, Euna Na, Ye Yuan, Rui Wang, Joonho H. Ahn, Tianyin Xu, Nam Sung Kim, 2025
https://scholar.google.com/scholar?q=M5%3A+Mastering+Page+Migration+and+Memory+Management+for+CXL-based+Tiered+Memory+Systems
9. Dissecting CXL Memory Performance at Scale: Analysis, Modeling, and Optimization — approx. recent systems authors, 2024/2025
https://scholar.google.com/scholar?q=Dissecting+CXL+Memory+Performance+at+Scale%3A+Analysis%2C+Modeling%2C+and+Optimization
10. Improving Key-Value Cache Performance with Heterogeneous Memory Tiering: A Case Study of CXL-Based Memory Expansion — approx. recent cache/tiering authors, 2024/2025
https://scholar.google.com/scholar?q=Improving+Key-Value+Cache+Performance+with+Heterogeneous+Memory+Tiering%3A+A+Case+Study+of+CXL-Based+Memory+Expansion
11. Tolerate It if You Cannot Reduce It: Handling Latency in Tiered Memory — approx. recent tiered-memory authors, 2024/2025
https://scholar.google.com/scholar?q=Tolerate+It+if+You+Cannot+Reduce+It%3A+Handling+Latency+in+Tiered+Memory
12. Can Hardware Outsmart Software in Tiered Memory Management? A CMM-H Case Study — approx. recent tiered-memory authors, 2024/2025
https://scholar.google.com/scholar?q=Can+Hardware+Outsmart+Software+in+Tiered+Memory+Management%3F+A+CMM-H+Case+Study
13. NeoMem: Hardware/Software Co-Design for CXL-Native Memory Tiering — approx. recent CXL co-design authors, 2024/2025
https://scholar.google.com/scholar?q=NeoMem%3A+Hardware%2FSoftware+Co-Design+for+CXL-Native+Memory+Tiering
14. Survey of Disaggregated Memory: Cross-Layer Technique Insights for Next-Generation Datacenters — approx. survey authors, 2024/2025
https://scholar.google.com/scholar?q=Survey+of+Disaggregated+Memory%3A+Cross-Layer+Technique+Insights+for+Next-Generation+Datacenters
15. Disaggregated Memory in the Datacenter: A Survey — approx. survey authors, 2024/2025
https://scholar.google.com/scholar?q=Disaggregated+Memory+in+the+Datacenter%3A+A+Survey
16. AI Post Transformers: CXL Computational Memory Offloading for Lower Runtime — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-cxl-computational-memory-offloading-for-3b2124.mp3
17. AI Post Transformers: FengHuang for Rack-Scale LLM Inference Memory — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-12-fenghuang-for-rack-scale-llm-inference-m-62708e.mp3
18. AI Post Transformers: Splitwise: Phase-Split LLM Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-26-splitwise-phase-split-llm-inference-e8945b.mp3
Interactive Visualization: Vistara Brings CXL Memory to Hyperscale

This episode explores a historical argument that simple Hebbian learning rules are too weak to explain real intelligence, because they mainly capture local correlations rather than solving multivariate credit-assignment problems. It examines how that critique points toward global optimization methods such as backpropagation and, in some settings, reinforcement learning, while contrasting their engineering success with the biological appeal of local synaptic updates. The discussion uses examples like XOR and later Hebbian variants such as Oja and BCM to show that the real issue is not whether Hebbian ideas are useless, but what kind of optimization principle is powerful enough to support complex learning. A listener would find it interesting for its mix of AI history, mathematical intuition, and an early attempt to connect learning theory to broader questions about consciousness.

Sources:
1. Optimization, Credit Assignment, and Consciousness
https://gwern.net/doc/ai/nn/rnn/1998-werbos.pdf
2. The Organization of Behavior: A Neuropsychological Theory — Donald O. Hebb, 1949
https://scholar.google.com/scholar?q=The+Organization+of+Behavior%3A+A+Neuropsychological+Theory
3. A Simplified Neuron Model as a Principal Component Analyzer — Erkki Oja, 1982
https://scholar.google.com/scholar?q=A+Simplified+Neuron+Model+as+a+Principal+Component+Analyzer
4. Theory for the Development of Neuron Selectivity: Orientation Specificity and Binocular Interaction in Visual Cortex — Elie L. Bienenstock, Leon N. Cooper, Paul W. Munro, 1982
https://scholar.google.com/scholar?q=Theory+for+the+Development+of+Neuron+Selectivity%3A+Orientation+Specificity+and+Binocular+Interaction+in+Visual+Cortex
5. Equivalence of Backpropagation and Contrastive Hebbian Learning in a Layered Network — Xiaohui Xie, H. Sebastian Seung, 2003
https://scholar.google.com/scholar?q=Equivalence+of+Backpropagation+and+Contrastive+Hebbian+Learning+in+a+Layered+Network
6. The Organization of Behavior — Donald O. Hebb, 1949
https://scholar.google.com/scholar?q=The+Organization+of+Behavior
7. Optimization: A Foundation for Understanding Consciousness — Paul J. Werbos, 1996
https://scholar.google.com/scholar?q=Optimization%3A+A+Foundation+for+Understanding+Consciousness
8. Learning Representations by Back-Propagating Errors — David E. Rumelhart, Geoffrey E. Hinton, Ronald J. Williams, 1986
https://scholar.google.com/scholar?q=Learning+Representations+by+Back-Propagating+Errors
9. A Theory of Cerebral Neocortex — Elie L. Bienenstock, Leon N. Cooper, Paul W. Munro, 1982
https://scholar.google.com/scholar?q=A+Theory+of+Cerebral+Neocortex
10. Reinforcement Learning: An Introduction — Richard S. Sutton, Andrew G. Barto, 1998
https://scholar.google.com/scholar?q=Reinforcement+Learning%3A+An+Introduction
11. Backpropagation-free spiking neural networks with the forward-forward algorithm — authors not identifiable from the snippet, recent; exact year not identifiable from the snippet
https://scholar.google.com/scholar?q=Backpropagation-free+spiking+neural+networks+with+the+forward-forward+algorithm
12. AI Post Transformers: Reverse-Mode Differentiation Across AD and Neural Nets — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-29-reverse-mode-differentiation-across-ad-a-5c1f77.mp3
13. AI Post Transformers: Backpropagation Through Time Explained — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-08-backpropagation-through-time-explained-eea44a.mp3
14. AI Post Transformers: Reinforcement Learning in 2025: An Overview — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-04-reinforcement-learning-in-2025-an-overvi-e7a4ce.mp3
15. AI Post Transformers: Long Short-Term Memory and Vanishing Gradients — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-19-long-short-term-memory-and-vanishing-gra-72448c.mp3
16. AI Post Transformers: Deep Learning in Spiking Neural Networks — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-24-deep-learning-in-spiking-neural-networks-bf558f.mp3
Interactive Visualization: Optimization, Credit Assignment, and Consciousness

This episode explores a mechanistic interpretability paper arguing that transformers often fail at counting not because they lack an internal notion of quantity, but because the pathway that converts that latent count into digit tokens is poorly aligned. It explains key ideas like linear probes, the logit lens, attention, LoRA, constrained next-token evaluation, and autoregressive generation to show how the authors separate “the model knows” from “the model can say.” The discussion highlights striking evidence that intermediate hidden states can encode counts almost perfectly while the corresponding digit readout directions remain nearly orthogonal, creating a readout bottleneck. Listeners would find it interesting because it reframes a familiar model weakness into a precise geometric and causal diagnosis, with implications for how to fix generation failures in modern model families like Pythia, Qwen3, and Mistral.

Interactive Visualization: Why Transformers Fail at Counting
Sources:
1. Why Transformers Fail at Counting
https://arxiv.org/pdf/2605.03258
2. Teaching Arithmetic to Small Transformers — Andrew McLeish, David Irving, Simon Sokota, Max Black, Berlin Chen, et al., 2024
https://scholar.google.com/scholar?q=Teaching+Arithmetic+to+Small+Transformers
3. Language Models Use Trigonometry to Do Addition — Stephen McLeish, et al., 2024
https://scholar.google.com/scholar?q=Language+Models+Use+Trigonometry+to+Do+Addition
4. Faithfulness of Linear Probes in Transformers — Various probe-critique literature; a representative reference should be cited explicitly by the author, 2019-2024
https://scholar.google.com/scholar?q=Faithfulness+of+Linear+Probes+in+Transformers
5. ROME: Locating and Editing Factual Associations in GPT — Kevin Meng, David Bau, Alex Andonian, Yonatan Belinkov, 2022
https://scholar.google.com/scholar?q=ROME%3A+Locating+and+Editing+Factual+Associations+in+GPT
6. The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets — Wes Gurnee, et al., 2023
https://scholar.google.com/scholar?q=The+Geometry+of+Truth%3A+Emergent+Linear+Structure+in+Large+Language+Model+Representations+of+True%2FFalse+Datasets
7. A Mathematical Framework for Transformer Circuits — Nelson Elhage, et al., 2021
https://scholar.google.com/scholar?q=A+Mathematical+Framework+for+Transformer+Circuits
8. Finding Transformer Circuits with Edge-Level Attribution Patching — Neel Nanda, et al., 2023
https://scholar.google.com/scholar?q=Finding+Transformer+Circuits+with+Edge-Level+Attribution+Patching
9. Tokenization counts: the impact of tokenization on arithmetic in frontier LLMs — Aaditya K. Singh, DJ Strouse, 2024
https://scholar.google.com/scholar?q=Tokenization+counts%3A+the+impact+of+tokenization+on+arithmetic+in+frontier+LLMs
10. Efficient numeracy in language models through single-token number embeddings — Linus Kreitner, Paul Hager, Jonathan Mengedoht, Georgios Kaissis, Daniel Rueckert, Martin J. Menten, 2025
https://scholar.google.com/scholar?q=Efficient+numeracy+in+language+models+through+single-token+number+embeddings
11. Arithmetic-Based Pretraining Improving Numeracy of Pretrained Language Models — Dominic Petrak, Nafise Sadat Moosavi, Iryna Gurevych, 2023
https://scholar.google.com/scholar?q=Arithmetic-Based+Pretraining+Improving+Numeracy+of+Pretrained+Language+Models
12. Rethinking Weight Tying: Pseudo-Inverse Tying for Stable LM Training and Updates — Jian Gu, Aldeida Aleti, Chunyang Chen, Hongyu Zhang, 2026
https://scholar.google.com/scholar?q=Rethinking+Weight+Tying%3A+Pseudo-Inverse+Tying+for+Stable+LM+Training+and+Updates
13. Latent Causal Probing: A Formal Perspective on Probing with Causal Models of Data — Charles Jin, 2024
https://scholar.google.com/scholar?q=Latent+Causal+Probing%3A+A+Formal+Perspective+on+Probing+with+Causal+Models+of+Data
14. AI Post Transformers: Linear Classifier Probes for Intermediate Layers — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-16-linear-classifier-probes-for-intermediat-927ae3.mp3
15. AI Post Transformers: Neural Chameleons and Evading Activation Monitors — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-14-neural-chameleons-and-evading-activation-bc470e.mp3
16. AI Post Transformers: How Induction Heads Emerge in Transformers — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-03-how-induction-heads-emerge-in-transforme-a7bfcb.mp3
17. AI Post Transformers: Language Models are Injective and Hence Invertible — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-21-language-models-are-injective-an-7545e0.mp3
18. AI Post Transformers: Latent Space as a New Computational Paradigm — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-05-latent-space-as-a-new-computational-para-810f39.mp3
Interactive Visualization: Why Transformers Fail at Counting

This episode explores a 2026 paper on “split personality training,” a method for attaching an internal reviewer to a language model that can reveal what the model knows about its own deceptive or reward-hacking behavior without changing the answer shown to the user. It situates the work in the broader lineage of latent knowledge elicitation, alignment faking, and mechanistic interpretability, explaining why a model’s hidden state may contain more honest information than its final text output. The discussion focuses on the paper’s use of a LoRA-based “honest persona” that activates only after the main response, and on benchmark setups like Anthropic’s auditing game that test whether internal representations expose hidden objectives that outside observers cannot infer. Listeners would find it interesting because it tackles a central safety problem: whether models can be audited for strategic deception using their own internal signals rather than their polished outward behavior.

Sources:
1. Split Personality Training Reveals Latent Knowledge
https://arxiv.org/pdf/2602.05532
2. Discovering Latent Knowledge in Language Models Without Supervision — Collin Burns, Haotian Ye, Dan Klein, Jacob Steinhardt, 2023
https://scholar.google.com/scholar?q=Discovering+Latent+Knowledge+in+Language+Models+Without+Supervision
3. Eliciting Latent Knowledge from Quirky Language Models — Alex Mallen, Nora Belrose, 2023
https://scholar.google.com/scholar?q=Eliciting+Latent+Knowledge+from+Quirky+Language+Models
4. Challenges with Unsupervised LLM Knowledge Discovery — Sebastian Farquhar, Vikrant Varma, Zachary Kenton, Johannes Gasteiger, Vladimir Mikulik, Rohin Shah, 2023
https://scholar.google.com/scholar?q=Challenges+with+Unsupervised+LLM+Knowledge+Discovery
5. LatentQA: Teaching LLMs to Decode Activations Into Natural Language — Alexander Pan, Lijie Chen, Jacob Steinhardt, 2024
https://scholar.google.com/scholar?q=LatentQA%3A+Teaching+LLMs+to+Decode+Activations+Into+Natural+Language
6. Language Models Mostly Know What They Know — Collin Burns, Haotian Ye, Dan Klein, Jacob Steinhardt, 2023
https://scholar.google.com/scholar?q=Language+Models+Mostly+Know+What+They+Know
7. ELK Report: Eliciting Latent Knowledge — Paul Christiano, Ajeya Cotra, Mark Xu, et al., 2021
https://scholar.google.com/scholar?q=ELK+Report%3A+Eliciting+Latent+Knowledge
8. The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets — Samuel Marks, Max Tegmark, 2024
https://scholar.google.com/scholar?q=The+Geometry+of+Truth%3A+Emergent+Linear+Structure+in+Large+Language+Model+Representations+of+True%2FFalse+Datasets
9. LatentQA: Teaching Language Models to Decode Activations into Natural Language — Yida Pan et al., 2024
https://scholar.google.com/scholar?q=LatentQA%3A+Teaching+Language+Models+to+Decode+Activations+into+Natural+Language
10. Activation Oracles — Jesse Karvonen et al., 2026
https://scholar.google.com/scholar?q=Activation+Oracles
11. Confessions — Nikhil Joglekar et al., 2025
https://scholar.google.com/scholar?q=Confessions
12. Self-Report Fine-Tuning — Li et al., 2025
https://scholar.google.com/scholar?q=Self-Report+Fine-Tuning
13. Auditing Language Models for Hidden Objectives — Anthropic, 2025
https://scholar.google.com/scholar?q=Auditing+Language+Models+for+Hidden+Objectives
14. Alignment Faking in Large Language Models — Ryan Greenblatt et al., 2024
https://scholar.google.com/scholar?q=Alignment+Faking+in+Large+Language+Models
15. Towards Eliciting Latent Knowledge from LLMs with Mechanistic Interpretability — Bartosz Cywinski, Emil Ryd, Senthooran Rajamanoharan, Neel Nanda, 2025
https://scholar.google.com/scholar?q=Towards+Eliciting+Latent+Knowledge+from+LLMs+with+Mechanistic+Interpretability
16. Quantifying Elicitation of Latent Capabilities in Language Models — Elizabeth Donoway et al., 2025
https://scholar.google.com/scholar?q=Quantifying+Elicitation+of+Latent+Capabilities+in+Language+Models
17. BEEAR: Embedding-based Adversarial Removal of Safety Backdoors in Instruction-tuned Language Models — Yi Zeng, Weiyu Sun, Tran Huynh, Dawn Song, Bo Li, Ruoxi Jia, 2024
https://scholar.google.com/scholar?q=BEEAR%3A+Embedding-based+Adversarial+Removal+of+Safety+Backdoors+in+Instruction-tuned+Language+Models
18. Investigating Adversarial Trigger Transfer in Large Language Models — Nicholas Meade, Arkil Patel, Siva Reddy, 2024
https://scholar.google.com/scholar?q=Investigating+Adversarial+Trigger+Transfer+in+Large+Language+Models
19. When Thinking LLMs Lie: Unveiling the Strategic Deception in Representations of Reasoning Models — Kai Wang, Yihao Zhang, Meng Sun, 2025
https://scholar.google.com/scholar?q=When+Thinking+LLMs+Lie%3A+Unveiling+the+Strategic+Deception+in+Representations+of+Reasoning+Models
20. Is It Thinking or Cheating? Detecting Implicit Reward Hacking by Measuring Reasoning Effort — Xinpeng Wang, Nitish Joshi, Barbara Plank, Rico Angell, He He, 2025
https://scholar.google.com/scholar?q=Is+It+Thinking+or+Cheating%3F+Detecting+Implicit+Reward+Hacking+by+Measuring+Reasoning+Effort
21. AI Post Transformers: Linear Classifier Probes for Intermediate Layers — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-16-linear-classifier-probes-for-intermediat-927ae3.mp3
22. AI Post Transformers: CLUE: Hidden-State Clustering for Non-parametric Verification — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/clue-hidden-state-clustering-for-non-parametric-verification/
23. AI Post Transformers: Latent Space as a New Computational Paradigm — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-05-latent-space-as-a-new-computational-para-810f39.mp3
24. AI Post Transformers: Language Models are Injective and Hence Invertible — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-21-language-models-are-injective-an-7545e0.mp3
Interactive Visualization: Split Personality Training Reveals Latent Knowledge

This episode explores RAPTOR, a method for extracting concept directions from language model hidden states using ridge-regularized logistic probes, with the goal of making those directions accurate enough for interpretation and stable enough for activation steering. It explains the core probe-then-steer workflow, why linear probes can reveal what a model has encoded, and why good classification accuracy does not necessarily produce a reliable control vector. The discussion situates the paper within broader debates in mechanistic interpretability, including concerns about brittle probes, distribution shift, and whether a single direction can really capture a concept like sentiment, refusal, or honesty. A listener would find it interesting because the episode turns an abstract interpretability question into a concrete engineering tradeoff about robustness, causal usefulness, and whether cheap white-box methods could become practical tools for controlling large models.

Sources:
1. RAPTOR: Ridge-Adaptive Logistic Probes — Ziqi Gao, Yaotian Zhu, Qingcheng Zeng, Xu Zhao, Ziqing Wang, Feng Ruan, Kaize Ding, 2026
http://arxiv.org/abs/2602.00158
2. Plug and Play Language Models: a Simple Approach to Controlled Text Generation — Sumanth Dathathri, Andrea Madotto, Janice Lan, Jason Yosinski, Rosanne Liu, et al., 2019
https://arxiv.org/abs/1912.02164
3. Activation Addition: Steering Language Models Without Optimization — Alexander Matt Turner, Lisa Thiergart, Gavin Leech, David Udell, Ulisse Mini, Monte MacDiarmid, 2023
https://arxiv.org/abs/2308.10248
4. Representation Engineering: A Top-Down Approach to AI Transparency — Andy Zou, Long Phan, Sarah Chen, James Campbell, Dan Hendrycks, J. Zico Kolter, et al., 2023
https://arxiv.org/abs/2310.01405
5. Refusal in Language Models Is Mediated by a Single Direction — Andy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka, Nina Panickssery, Wes Gurnee, Neel Nanda, 2024
https://arxiv.org/abs/2406.11717
6. Understanding Intermediate Layers Using Linear Classifier Probes — Guillaume Alain, Yoshua Bengio, 2016
https://openreview.net/forum?id=HJ4-rAVtl
7. Designing and Interpreting Probes with Control Tasks — John Hewitt, Percy Liang, 2019
https://aclanthology.org/D19-1275/
8. Information-Theoretic Probing with Minimum Description Length — Elena Voita, Ivan Titov, 2020
https://aclanthology.org/2020.emnlp-main.14/
9. RAPTOR: Ridge-Adaptive Logistic Probes — Ziqi Gao, Yaotian Zhu, Qingcheng Zeng, Xu Zhao, Ziqing Wang, Feng Ruan, Kaize Ding, 2026
https://arxiv.org/abs/2602.00158
10. Steering Llama 2 via Contrastive Activation Addition — Nina Rimsky, Nick Gabrieli, Julian Schulz, Meg Tong, Evan Hubinger, Alexander Turner, 2024
https://scholar.google.com/scholar?q=Steering+Llama+2+via+Contrastive+Activation+Addition
11. Analysing the Generalisation and Reliability of Steering Vectors — Daniel Tan, David Chanin, Aengus Lynch, Brooks Paige, Dimitrios Kanoulas, Adrià Garriga-Alonso, Robert Kirk, 2024
https://scholar.google.com/scholar?q=Analysing+the+Generalisation+and+Reliability+of+Steering+Vectors
12. Beyond Single Concept Vector: Modeling Concept Subspace in LLMs with Gaussian Distribution — Haiyan Zhao, Heng Zhao, Bo Shen, Ali Payani, Fan Yang, Mengnan Du, 2025
https://scholar.google.com/scholar?q=Beyond+Single+Concept+Vector%3A+Modeling+Concept+Subspace+in+LLMs+with+Gaussian+Distribution
13. Controlling Large Language Models Through Concept Activation Vectors — Hanyu Zhang, Xiting Wang, Chengao Li, Xiang Ao, Qing He, 2025
https://scholar.google.com/scholar?q=Controlling+Large+Language+Models+Through+Concept+Activation+Vectors
14. Token prepending: A training-free approach for eliciting better sentence embeddings from llms — authors not confirmed from provided snippet, recent, unverified
https://scholar.google.com/scholar?q=Token+prepending%3A+A+training-free+approach+for+eliciting+better+sentence+embeddings+from+llms
15. Rep2Text: Decoding Full Text from a Single LLM Token Representation — authors not confirmed from provided snippet, recent, unverified
https://scholar.google.com/scholar?q=Rep2Text%3A+Decoding+Full+Text+from+a+Single+LLM+Token+Representation
16. Context Matters: Analyzing the Generalizability of Linear Probing and Steering Across Diverse Scenarios — authors not confirmed from provided snippet, recent, unverified
https://scholar.google.com/scholar?q=Context+Matters%3A+Analyzing+the+Generalizability+of+Linear+Probing+and+Steering+Across+Diverse+Scenarios
17. Angular steering: Behavior control via rotation in activation space — authors not confirmed from provided snippet, recent, unverified
https://scholar.google.com/scholar?q=Angular+steering%3A+Behavior+control+via+rotation+in+activation+space
18. Global Evolutionary Steering: Refining Activation Steering Control via Cross-Layer Consistency — authors not confirmed from provided snippet, recent, unverified
https://scholar.google.com/scholar?q=Global+Evolutionary+Steering%3A+Refining+Activation+Steering+Control+via+Cross-Layer+Consistency
19. Fine-Grained Activation Steering: Steering Less, Achieving More — authors not confirmed from provided snippet, recent, unverified
https://scholar.google.com/scholar?q=Fine-Grained+Activation+Steering%3A+Steering+Less%2C+Achieving+More
20. AI Post Transformers: Language Models are Injective and Hence Invertible — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-21-language-models-are-injective-an-7545e0.mp3
21. AI Post Transformers: Learning to Reason with 13 Parameters — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-14-learning-to-reason-with-13-parameters-54c87f.mp3
22. AI Post Transformers: When Spectral Gradient Updates Help Deep Learning — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-when-spectral-gradient-updates-help-deep-9c8441.mp3
Interactive Visualization: RAPTOR: Stable Concept Directions From Logistic Probes

This episode explores a paper arguing that a model’s internal answer belief can form well before its visible chain-of-thought reveals it, raising doubts about whether reasoning traces are true explanations or polished post hoc narratives. It explains core ideas such as chain-of-thought faithfulness, activation monitoring, mechanistic interpretability, confidence calibration, and dynamic inference, framing the broader safety question of whether text reasoning can really serve as an audit trail. The discussion focuses on the paper’s method of comparing internal activation probes, forced early answers, and monitors of partial written reasoning to test whether models “know” the answer before their text shows it. Listeners would find it interesting because it connects interpretability research to practical concerns about oversight, trust, and compute efficiency, while contrasting easy recall-heavy benchmarks with harder multistep science questions where genuine belief updates may still happen during inference.

Sources:
1. Reasoning Theater: Disentangling Model Beliefs from Chain-of-Thought — Siddharth Boppana, Annabel Ma, Max Loeffler, Raphael Sarfati, Eric Bigelow, Atticus Geiger, Owen Lewis, Jack Merullo, 2026
http://arxiv.org/abs/2603.05488
2. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models — Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, Denny Zhou, 2022
https://scholar.google.com/scholar?q=Chain-of-Thought+Prompting+Elicits+Reasoning+in+Large+Language+Models
3. Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting — Miles Turpin, Julian Michael, Ethan Perez, Samuel R. Bowman, 2023
https://scholar.google.com/scholar?q=Language+Models+Don%27t+Always+Say+What+They+Think%3A+Unfaithful+Explanations+in+Chain-of-Thought+Prompting
4. Measuring Faithfulness in Chain-of-Thought Reasoning — Tamera Lanham, Anna Chen, Ansh Radhakrishnan, Benoit Steiner, Carson Denison, Danny Hernandez, Dustin Li, Esin Durmus, Evan Hubinger, Jackson Kernion and others, 2023
https://scholar.google.com/scholar?q=Measuring+Faithfulness+in+Chain-of-Thought+Reasoning
5. Reasoning Models Don't Always Say What They Think — Yanda Chen, Joe Benton, Ansh Radhakrishnan, Jonathan Uesato, Carson Denison, John Schulman, Arushi Somani, Peter Hase, Misha Wagner, Fabien Roger, Vlad Mikulik, Samuel R. Bowman, Jan Leike, Jared Kaplan, Ethan Perez, 2025
https://scholar.google.com/scholar?q=Reasoning+Models+Don%27t+Always+Say+What+They+Think
6. Performative Thinking? The Brittle Correlation between CoT Length and Problem Complexity — Vivek Palod, Karthik Valmeekam, Kyle Stechly, Subbarao Kambhampati, 2025
https://scholar.google.com/scholar?q=Performative+Thinking%3F+The+Brittle+Correlation+between+CoT+Length+and+Problem+Complexity
7. Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety — Tomek Korbak, Mikita Balesni, Elizabeth Barnes, Yoshua Bengio, et al., 2025
https://scholar.google.com/scholar?q=Chain+of+Thought+Monitorability%3A+A+New+and+Fragile+Opportunity+for+AI+Safety
8. Reasoning Models Know When They're Right: Probing Hidden States for Self-Verification — Anqi Zhang, Yulin Chen, Jane Pan, Chen Zhao, Aurojit Panda, Jinyang Li, He He, 2025
https://scholar.google.com/scholar?q=Reasoning+Models+Know+When+They%27re+Right%3A+Probing+Hidden+States+for+Self-Verification
9. A Positive Case for Faithfulness: LLM Self-Explanations Help Predict Model Behavior — Harry Mayne, Justin Singh Kang, Dewi Gould, Kannan Ramchandran, Adam Mahdi, Noah Y. Siegel, 2026
https://scholar.google.com/scholar?q=A+Positive+Case+for+Faithfulness%3A+LLM+Self-Explanations+Help+Predict+Model+Behavior
10. Base Models Know How to Reason, Thinking Models Learn When — Constantin Venhoff, Ivan Arcuschin, Philip Torr, Arthur Conmy, Neel Nanda, 2025
https://scholar.google.com/scholar?q=Base+Models+Know+How+to+Reason%2C+Thinking+Models+Learn+When
11. Measuring Chain of Thought Faithfulness by Unlearning Reasoning Steps — Martin Tutek, Fateme Hashemi Chaleshtori, Ana Marasovic, Yonatan Belinkov, 2025
https://scholar.google.com/scholar?q=Measuring+Chain+of+Thought+Faithfulness+by+Unlearning+Reasoning+Steps
12. Faithful Chain-of-Thought Reasoning — Qing Lyu et al., 2023
https://scholar.google.com/scholar?q=Faithful+Chain-of-Thought+Reasoning
13. Towards Thinking-Optimal Scaling of Test-Time Compute for LLM Reasoning — Wenkai Yang, Shuming Ma, Yankai Lin, Furu Wei, 2025
https://scholar.google.com/scholar?q=Towards+Thinking-Optimal+Scaling+of+Test-Time+Compute+for+LLM+Reasoning
14. AI Post Transformers: Do Language Models Know Their Limits — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-01-do-language-models-know-their-limits-48e444.mp3
15. AI Post Transformers: Selective Classification with Deep Neural Networks — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-01-selective-classification-with-deep-neura-bed8cb.mp3
16. AI Post Transformers: Stabilizing Efficient Reasoning with Step-Level Advantage Selection — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-29-stabilizing-efficient-reasoning-with-ste-1e589d.mp3
17. AI Post Transformers: Test-time Scaling for Multi-Agent Collaborative Reasoning — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-22-test-time-scaling-for-multi-agent-collab-082570.mp3
Interactive Visualization: Reasoning Theater and Unfaithful Chain-of-Thought

This episode explores a paper on long-context compression that argues standard “soft compression” methods, which rely on learned memory or gist tokens, lose information because those tokens get overwritten across layers and fail to coordinate what each slot should retain. It explains the paper’s alternative design, which keeps the language model backbone frozen and instead explicitly transmits information from hidden states into a small set of latent slots through a two-stage process: selecting useful signals across layers, then globally allocating token information to slots with a transport-based assignment. The discussion highlights why this matters for deployment, where long contexts and growing KV caches make inference expensive, while also noting the risks of latent compression for exact recall, citations, and fine-grained factual detail. Listeners would find it interesting for both the strong benchmark results, where the method substantially outperforms prior compressors on several QA datasets, and the debate over whether those gains on a 512-token testbed really translate to the much larger context problems practitioners care about.

Sources:
1. Context Compression via Explicit Information Transmission — Jiangnan Ye, Hanqi Yan, Zhenyi Shen, Heng Chang, Ye Mao, Yulan He, 2026
http://arxiv.org/abs/2602.03784
2. Compressive Transformers for Long-Range Sequence Modelling — Jack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, Timothy P. Lillicrap, 2019
https://scholar.google.com/scholar?q=Compressive+Transformers+for+Long-Range+Sequence+Modelling
3. Learning to Compress Prompts with Gist Tokens — Jesse Mu, Xiang Lisa Li, Noah D. Goodman, 2023
https://scholar.google.com/scholar?q=Learning+to+Compress+Prompts+with+Gist+Tokens
4. Adapting Language Models to Compress Contexts — Alexis Chevalier, Alexander Wettig, Anirudh Ajith, Danqi Chen, 2023
https://scholar.google.com/scholar?q=Adapting+Language+Models+to+Compress+Contexts
5. A Silver Bullet or a Compromise for Full Attention? A Comprehensive Study of Gist Token-based Context Compression — Chenlong Deng, Zhisong Zhang, Kelong Mao, Shuaiyi Li, Xinting Huang, Dong Yu, Zhicheng Dou, 2024
https://scholar.google.com/scholar?q=A+Silver+Bullet+or+a+Compromise+for+Full+Attention%3F+A+Comprehensive+Study+of+Gist+Token-based+Context+Compression
6. In-context Autoencoder for Context Compression in a Large Language Model — Tao Ge, Jing Hu, Haixun Wang, Si-Qing Chen, Furu Wei, 2024
https://scholar.google.com/scholar?q=In-context+Autoencoder+for+Context+Compression+in+a+Large+Language+Model
7. 500xCompressor: Generalized Prompt Compression for Large Language Models — Zongqian Li, Yixuan Su, Nigel Collier, 2025
https://scholar.google.com/scholar?q=500xCompressor%3A+Generalized+Prompt+Compression+for+Large+Language+Models
8. Long Context Compression with Activation Beacon — Peitian Zhang, Zheng Liu, Shitao Xiao, Ninglu Shao, Qiwei Ye, Zhicheng Dou, 2025
https://scholar.google.com/scholar?q=Long+Context+Compression+with+Activation+Beacon
9. RazorAttention: Efficient KV Cache Compression through Retrieval Heads — approx. anonymous/unknown from snippet, 2025/2026
https://scholar.google.com/scholar?q=RazorAttention%3A+Efficient+KV+Cache+Compression+through+Retrieval+Heads
10. Not All Heads Matter: A Head-Level KV Cache Compression Method with Integrated Retrieval and Reasoning — approx. anonymous/unknown from snippet, 2025/2026
https://scholar.google.com/scholar?q=Not+All+Heads+Matter%3A+A+Head-Level+KV+Cache+Compression+Method+with+Integrated+Retrieval+and+Reasoning
11. Efficient Context Selection for Long-Context QA: No Tuning, No Iteration, Just Adaptive-k — approx. anonymous/unknown from snippet, 2025/2026
https://scholar.google.com/scholar?q=Efficient+Context+Selection+for+Long-Context+QA%3A+No+Tuning%2C+No+Iteration%2C+Just+Adaptive-k
12. TokenSelect: Efficient Long-Context Inference and Length Extrapolation for LLMs via Dynamic Token-Level KV Cache Selection — approx. anonymous/unknown from snippet, 2025/2026
https://scholar.google.com/scholar?q=TokenSelect%3A+Efficient+Long-Context+Inference+and+Length+Extrapolation+for+LLMs+via+Dynamic+Token-Level+KV+Cache+Selection
13. Generative Adapter: Contextualizing Language Models in Parameters with a Single Forward Pass — approx. anonymous/unknown from snippet, 2025/2026
https://scholar.google.com/scholar?q=Generative+Adapter%3A+Contextualizing+Language+Models+in+Parameters+with+a+Single+Forward+Pass
14. Demystifying the Roles of LLM Layers in Retrieval, Knowledge, and Reasoning — approx. anonymous/unknown from snippet, 2025/2026
https://scholar.google.com/scholar?q=Demystifying+the+Roles+of+LLM+Layers+in+Retrieval%2C+Knowledge%2C+and+Reasoning
15. AI Post Transformers: DeepSeek-V4 and Practical Million-Token Context — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-25-deepseek-v4-and-practical-million-token-6f4de1.mp3
16. AI Post Transformers: Gated Linear Attention for Efficient Long Sequences — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-18-gated-linear-attention-for-efficient-lon-c858ab.mp3
17. AI Post Transformers: Doc-to-LoRA: Internalizing Context as LoRA — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-29-doc-to-lora-internalizing-context-as-lor-8dd5ec.mp3
18. AI Post Transformers: Memory Sparse Attention for 100M-Token Scaling — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-07-memory-sparse-attention-for-100m-token-s-377cff.mp3
19. AI Post Transformers: Recursive Language Models for Arbitrarily Long Prompts — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-recursive-language-models-for-arbitraril-fbcd1c.mp3
20. AI Post Transformers: How Induction Heads Emerge in Transformers — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-03-how-induction-heads-emerge-in-transforme-a7bfcb.mp3
Interactive Visualization: Explicit Information Transmission for Context Compression

This episode explores a position paper arguing that modern LLM serving has outgrown simple heuristics like FIFO, shortest-queue routing, and LRU eviction. It explains why transformer inference creates harder control problems than standard inference, focusing on continuous batching, KV-cache growth, and the tension between compute-heavy prefill and memory-bound decode phases. The discussion highlights the paper’s central claim that serving systems need explicit objective-driven optimization for routing, admission control, scheduling, and cache management, while also questioning where formal methods would truly outperform today’s stronger heuristic baselines such as vLLM and PagedAttention-inspired designs. Listeners would find it interesting because it connects low-level serving mechanics to real product tradeoffs like latency, throughput, and cache churn, showing why infrastructure choices increasingly shape LLM performance.

Sources:
1. Position: LLM Serving Needs Mathematical Optimization and Algorithmic Foundations, Not Just Heuristics — Zijie Zhou, 2026
http://arxiv.org/abs/2605.01280
2. PREBLE: Efficient Distributed Prompt Scheduling for LLM Serving — Vikranth Srivatsa, Zijian He, Reyna Abhyankar, Dongming Li, Yiying Zhang, 2024
https://scholar.google.com/scholar?q=PREBLE%3A+Efficient+Distributed+Prompt+Scheduling+for+LLM+Serving
3. DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving — Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, Hao Zhang, 2024
https://scholar.google.com/scholar?q=DistServe%3A+Disaggregating+Prefill+and+Decoding+for+Goodput-optimized+Large+Language+Model+Serving
4. Semantic Caching for Low-Cost LLM Serving: From Offline Learning to Online Adaptation — Xutong Liu, Baran Atalar, Xiangxiang Dai, Jinhang Zuo, Siwei Wang, John C. S. Lui, Wei Chen, Carlee Joe-Wong, 2026
https://scholar.google.com/scholar?q=Semantic+Caching+for+Low-Cost+LLM+Serving%3A+From+Offline+Learning+to+Online+Adaptation
5. POLAR: Online Learning for LoRA Adapter Caching and Routing in Edge LLM Serving — Shaoang Li, Jian Li, 2026
https://scholar.google.com/scholar?q=POLAR%3A+Online+Learning+for+LoRA+Adapter+Caching+and+Routing+in+Edge+LLM+Serving
6. Orca: A Distributed Serving System for Transformer-Based Generative Models — Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim, Byung-Gon Chun, 2022
https://scholar.google.com/scholar?q=Orca%3A+A+Distributed+Serving+System+for+Transformer-Based+Generative+Models
7. Efficient Memory Management for Large Language Model Serving with PagedAttention — Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, Ion Stoica, 2023
https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention
8. Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve — Amey Agrawal, Nitin Kedia, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav Gulavani, Alexey Tumanov, Ramachandran Ramjee, 2024
https://scholar.google.com/scholar?q=Taming+Throughput-Latency+Tradeoff+in+LLM+Inference+with+Sarathi-Serve
9. SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills — Amey Agrawal, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav S. Gulavani, Ramachandran Ramjee, 2023
https://scholar.google.com/scholar?q=SARATHI%3A+Efficient+LLM+Inference+by+Piggybacking+Decodes+with+Chunked+Prefills
10. Faster LLM Inference using DBMS-Inspired Preemption and Cache Replacement Policies — Kyoungmin Kim, Jiacheng Li, Kijae Hong, Anastasia Ailamaki, 2024
https://scholar.google.com/scholar?q=Faster+LLM+Inference+using+DBMS-Inspired+Preemption+and+Cache+Replacement+Policies
11. DualMap: Enabling Both Cache Affinity and Load Balancing for Distributed LLM Serving — Ying Yuan et al., 2026
https://scholar.google.com/scholar?q=DualMap%3A+Enabling+Both+Cache+Affinity+and+Load+Balancing+for+Distributed+LLM+Serving
12. Slice-Level Scheduling for High Throughput and Load Balanced LLM Serving — Ke Cheng et al., 2024
https://scholar.google.com/scholar?q=Slice-Level+Scheduling+for+High+Throughput+and+Load+Balanced+LLM+Serving
13. A Predictive and Synergistic Two-Layer Scheduling Framework for LLM Serving — Yue Zhang et al., 2025
https://scholar.google.com/scholar?q=A+Predictive+and+Synergistic+Two-Layer+Scheduling+Framework+for+LLM+Serving
14. KVLink: Accelerating Large Language Models via Efficient KV Cache Reuse — Jingbo Yang et al., 2025
https://scholar.google.com/scholar?q=KVLink%3A+Accelerating+Large+Language+Models+via+Efficient+KV+Cache+Reuse
15. HyperRAG: Enhancing Quality-Efficiency Tradeoffs in Retrieval-Augmented Generation with Reranker KV-Cache Reuse — Yuwei An et al., 2025
https://scholar.google.com/scholar?q=HyperRAG%3A+Enhancing+Quality-Efficiency+Tradeoffs+in+Retrieval-Augmented+Generation+with+Reranker+KV-Cache+Reuse
16. ProphetKV: User-Query-Driven Selective Recomputation for Efficient KV Cache Reuse in Retrieval-Augmented Generation — Shihao Wang et al., 2026
https://scholar.google.com/scholar?q=ProphetKV%3A+User-Query-Driven+Selective+Recomputation+for+Efficient+KV+Cache+Reuse+in+Retrieval-Augmented+Generation
17. dLoRA: Dynamically Orchestrating Requests and Adapters for LoRA LLM Serving — Bingyang Wu et al., 2024
https://scholar.google.com/scholar?q=dLoRA%3A+Dynamically+Orchestrating+Requests+and+Adapters+for+LoRA+LLM+Serving
18. SPAD: Specialized Prefill and Decode Hardware for Disaggregated LLM Inference — Hengrui Zhang et al., 2025
https://scholar.google.com/scholar?q=SPAD%3A+Specialized+Prefill+and+Decode+Hardware+for+Disaggregated+LLM+Inference
19. AI Post Transformers: Speculative Decoding in Real vLLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-speculative-decoding-in-real-vllm-servin-6f4e2b.mp3
20. AI Post Transformers: Lookahead Q-Cache for Consistent KV Eviction — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-25-lookahead-q-cache-for-consistent-kv-evic-d97b09.mp3
21. AI Post Transformers: LookaheadKV: Fast and Accurate KV Cache Eviction by Glimpsing into the Future without Generation — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-25-lookaheadkv-fast-and-accurate-kv-9cfc9f.mp3
22. AI Post Transformers: KV Cache TTL for Multi-Turn Agent Scheduling — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-09-kv-cache-ttl-for-multi-turn-agent-schedu-996bf1.mp3
23. AI Post Transformers: CacheFlow and 3D-Parallel KV Cache Restoration — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-01-cacheflow-and-3d-parallel-kv-cache-resto-8db883.mp3
24. AI Post Transformers: ContiguousKV for Faster LLM Prefill KV Reuse — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-20-contiguouskv-for-faster-llm-prefill-kv-r-59f545.mp3
25. AI Post Transformers: Breaking the Prefix Barrier with Shared KV Cache — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-24-breaking-the-prefix-barrier-with-shared-a5e5a6.mp3
26. AI Post Transformers: TokenDance for Multi-Agent KV Cache Sharing — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-22-tokendance-for-multi-agent-kv-cache-shar-aa9b99.mp3
27. AI Post Transformers: RetrievalAttention for Long-Context LLM Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-17-retrievalattention-for-long-context-llm-ddf566.mp3
28. AI Post Transformers: KVSwap for Disk-Aware Long-Context On-Device Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-16-kvswap-for-disk-aware-long-context-on-de-f3c15e.mp3
Interactive Visualization: Why LLM Serving Needs Mathematical Optimization

This episode explores EverMemOS, a memory system for long-lived AI agents that tries to organize past interactions into structured, higher-level semantic “scenes” instead of relying on flat retrieval alone. It explains why bigger context windows and standard RAG often fail when agents accumulate stale preferences, conflicting facts, and fragmented conversational traces, arguing that the real problem is not just forgetting but poorly organized remembering. The discussion walks through the paper’s core design, including MemCells, MemScenes, semantic consolidation, and reconstructive recollection, framing the system as a state-management layer around transformers rather than a new model architecture. A listener would find it interesting because it connects abstract memory research to practical agent failures and offers a concrete alternative for building assistants that can reason more reliably over long time horizons.

Interactive Visualization: EverMemOS for Long-Horizon Agent Memory
Sources:
1. EverMemOS: A Self-Organizing Memory Operating System for Structured Long-Horizon Reasoning — Chuanrui Hu, Xingze Gao, Zuyi Zhou, Dannong Xu, Yi Bai, Xintong Li, Hui Zhang, Tong Li, Chong Zhang, Lidong Bing, Yafeng Deng, 2026
http://arxiv.org/abs/2601.02163
2. MemoryBank: Enhancing Large Language Models with Long-Term Memory — Wanjun Zhong, Lianghong Guo, Qiqi Gao, He Ye, Yanlin Wang, 2023
https://scholar.google.com/scholar?q=MemoryBank%3A+Enhancing+Large+Language+Models+with+Long-Term+Memory
3. MemGPT: Towards LLMs as Operating Systems — Charles Packer, Sarah Wooders, Kevin Lin, Vivian Fang, Shishir G. Patil, Joseph E. Gonzalez, 2023
https://scholar.google.com/scholar?q=MemGPT%3A+Towards+LLMs+as+Operating+Systems
4. A Survey on the Memory Mechanism of Large Language Model based Agents — Zeyu Zhang, Xiaohe Bo, Chen Ma, Rui Li, Xu Chen, Quanyu Dai, Jieming Zhu, Zhenhua Dong, Ji-Rong Wen, 2024
https://scholar.google.com/scholar?q=A+Survey+on+the+Memory+Mechanism+of+Large+Language+Model+based+Agents
5. MemOS: A Memory OS for AI System — Zhiyu Li, Shichao Song, Chenyang Xi, Hanyu Wang, Chen Tang, Simin Niu, Ding Chen, et al., 2025
https://scholar.google.com/scholar?q=MemOS%3A+A+Memory+OS+for+AI+System
6. Memory OS of AI Agent — Jiazheng Kang, Mingming Ji, Zhe Zhao, Ting Bai, 2025
https://scholar.google.com/scholar?q=Memory+OS+of+AI+Agent
7. Zep: A Temporal Knowledge Graph Architecture for Agent Memory — Preston Rasmussen, Pavlo Paliychuk, Travis Beauvais, Jack Ryan, Daniel Chalef, 2025
https://scholar.google.com/scholar?q=Zep%3A+A+Temporal+Knowledge+Graph+Architecture+for+Agent+Memory
8. Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory — Prateek Chhikara, Dev Khant, Saket Aryan, Taranjeet Singh, Deshraj Yadav, 2025
https://scholar.google.com/scholar?q=Mem0%3A+Building+Production-Ready+AI+Agents+with+Scalable+Long-Term+Memory
9. LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory — Di Wu, Hongwei Wang, Wenhao Yu, Yuwei Zhang, Kai-Wei Chang, Dong Yu, 2024
https://scholar.google.com/scholar?q=LongMemEval%3A+Benchmarking+Chat+Assistants+on+Long-Term+Interactive+Memory
10. Evaluating Very Long-Term Conversational Memory of LLM Agents — Adyasha Maharana, Dong-Ho Lee, Sergey Tulyakov, Mohit Bansal, Francesco Barbieri, Yuwei Fang, 2024
https://scholar.google.com/scholar?q=Evaluating+Very+Long-Term+Conversational+Memory+of+LLM+Agents
11. BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack — Yuri/Yurii Kuratov et al., 2024
https://scholar.google.com/scholar?q=BABILong%3A+Testing+the+Limits+of+LLMs+with+Long+Context+Reasoning-in-a-Haystack
12. LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks — Yushi Bai et al., 2024/2025
https://scholar.google.com/scholar?q=LongBench+v2%3A+Towards+Deeper+Understanding+and+Reasoning+on+Realistic+Long-context+Multitasks
13. Memory-Aware and Uncertainty-Guided Retrieval for Multi-Hop Question Answering — Yuelyu Ji, Rui Meng, Zhuochun Li, Daqing He, 2025
https://scholar.google.com/scholar?q=Memory-Aware+and+Uncertainty-Guided+Retrieval+for+Multi-Hop+Question+Answering
14. BRIEF: Bridging Retrieval and Inference for Multi-hop Reasoning via Compression — Yuankai Li, Jia-Chen Gu, Di Wu, Kai-Wei Chang, Nanyun Peng, 2024/2025
https://scholar.google.com/scholar?q=BRIEF%3A+Bridging+Retrieval+and+Inference+for+Multi-hop+Reasoning+via+Compression
15. Explicit v.s. Implicit Memory: Exploring Multi-hop Complex Reasoning Over Personalized Information — Zeyu Zhang et al., 2025
https://scholar.google.com/scholar?q=Explicit+v.s.+Implicit+Memory%3A+Exploring+Multi-hop+Complex+Reasoning+Over+Personalized+Information
16. Handling Preference Drift in Capturing Dynamic User Preferences for Streaming Session-Based Recommendations — Oussama Alahoum, Boudjemaa Boudaa, Laouni Djafri, 2025/2026
https://scholar.google.com/scholar?q=Handling+Preference+Drift+in+Capturing+Dynamic+User+Preferences+for+Streaming+Session-Based+Recommendations
17. Episodic Memory in AI Agents Poses Risks That Should Be Studied and Mitigated — Chad DeChant, 2025
https://scholar.google.com/scholar?q=Episodic+Memory+in+AI+Agents+Poses+Risks+That+Should+Be+Studied+and+Mitigated
18. AI Post Transformers: Memory Sparse Attention for 100M-Token Scaling — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-07-memory-sparse-attention-for-100m-token-s-377cff.mp3
19. AI Post Transformers: Recursive Language Models for Arbitrarily Long Prompts — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-recursive-language-models-for-arbitraril-fbcd1c.mp3
20. AI Post Transformers: Neural Computers as Learned Latent Runtimes — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-11-neural-computers-as-learned-latent-runti-9fa282.mp3
21. AI Post Transformers: How Induction Heads Emerge in Transformers — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-03-how-induction-heads-emerge-in-transforme-a7bfcb.mp3
Interactive Visualization: EverMemOS for Long-Horizon Agent Memory

This episode explores a March 2026 paper arguing that LLM-based judges are an unreliable way to measure jailbreak success and adversarial robustness. It explains how modern safety evaluations rely on judge models to score harmful outputs, then walks through why those judges can break under attack shift, model shift, and data shift, sometimes degrading to near coin-flip reliability. The discussion connects this critique to benchmarks such as MT-Bench, HarmBench, and StrongREJECT, and examines how weaknesses in the judging pipeline can inflate or distort reported attack success rates. Listeners would find it interesting because it challenges whether many headline jailbreak results are exposing real model failures or simply failures in the grading system.

Interactive Visualization: When LLM Judges Become Coin Flips
Sources:
1. A Coin Flip for Safety: LLM Judges Fail to Reliably Measure Adversarial Robustness — Leo Schwinn, Moritz Ladenburger, Tim Beyer, Mehrnaz Mofakhami, Gauthier Gidel, Stephan Günnemann, 2026
http://arxiv.org/abs/2603.06594
2. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena — Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zi Lin, Zhuohan Li, Joseph E. Gonzalez, Ion Stoica and others, 2023
https://scholar.google.com/scholar?q=Judging+LLM-as-a-Judge+with+MT-Bench+and+Chatbot+Arena
3. Large Language Models are not Fair Evaluators — Peiyi Wang, Lei Li, Liang Chen, Zefan Cai, Dawei Zhu, Binghuai Lin, Yunbo Cao, Qi Liu, Tianyu Liu, Zhifang Sui, 2023
https://scholar.google.com/scholar?q=Large+Language+Models+are+not+Fair+Evaluators
4. An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Models are Task-specific Classifiers — Hui Huang, Yingqi Qu, Jing Liu, Muyun Yang, Tiejun Zhao, 2024
https://scholar.google.com/scholar?q=An+Empirical+Study+of+LLM-as-a-Judge+for+LLM+Evaluation%3A+Fine-tuned+Judge+Models+are+Task-specific+Classifiers
5. JudgeBench: A Benchmark for Evaluating LLM-based Judges — Sijun Tan, Siyuan Zhuang, Kyle Montgomery, William Y. Tang, Alejandro Cuadron, Chenguang Wang, Raluca Ada Popa, Ion Stoica, 2024
https://scholar.google.com/scholar?q=JudgeBench%3A+A+Benchmark+for+Evaluating+LLM-based+Judges
6. Universal and Transferable Adversarial Attacks on Aligned Language Models — Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J. Zico Kolter, Matt Fredrikson, 2023
https://scholar.google.com/scholar?q=Universal+and+Transferable+Adversarial+Attacks+on+Aligned+Language+Models
7. HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal — Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang, Norman Mu, Elham Sakhaee, Nathaniel Li, Steven Basart, Bo Li, David Forsyth, Dan Hendrycks, 2024
https://scholar.google.com/scholar?q=HarmBench%3A+A+Standardized+Evaluation+Framework+for+Automated+Red+Teaming+and+Robust+Refusal
8. A StrongREJECT for Empty Jailbreaks — Alexandra Souly, Qingyuan Lu, Dillon Bowen, Tu Trinh, Elvis Hsieh, Sana Pandey, Pieter Abbeel, Justin Svegliato, Scott Emmons, Olivia Watkins, Sam Toyer, 2024
https://scholar.google.com/scholar?q=A+StrongREJECT+for+Empty+Jailbreaks
9. JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models — Patrick Chao, Edoardo Debenedetti, Alexander Robey, Maksym Andriushchenko, Francesco Croce, Vikash Sehwag, Edgar Dobriban, Nicolas Flammarion, George J. Pappas, Florian Tramèr, Hamed Hassani, Eric Wong, 2024
https://scholar.google.com/scholar?q=JailbreakBench%3A+An+Open+Robustness+Benchmark+for+Jailbreaking+Large+Language+Models
10. Dataset Shift in Machine Learning — Joaquin Quiñonero-Candela, Masashi Sugiyama, Anton Schwaighofer, Neil D. Lawrence (editors), 2008
https://scholar.google.com/scholar?q=Dataset+Shift+in+Machine+Learning
11. Can You Trust Your Model's Uncertainty? Evaluating Predictive Uncertainty Under Dataset Shift — Yaniv Ovadia, Emily Fertig, Jie Ren, Zachary Nado, D. Sculley, Sebastian Nowozin, Joshua Dillon, Balaji Lakshminarayanan, Jasper Snoek, 2019
https://scholar.google.com/scholar?q=Can+You+Trust+Your+Model%27s+Uncertainty%3F+Evaluating+Predictive+Uncertainty+Under+Dataset+Shift
12. Measuring Robustness to Natural Distribution Shifts in Image Classification — Rohan Taori, Achal Dave, Vaishaal Shankar, Nicholas Carlini, Benjamin Recht, Ludwig Schmidt, 2020
https://scholar.google.com/scholar?q=Measuring+Robustness+to+Natural+Distribution+Shifts+in+Image+Classification
13. WILDS: A Benchmark of in-the-Wild Distribution Shifts — Pang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie, Marvin Zhang, Akshay Balsubramani, Weihua Hu, Michihiro Yasunaga, Percy Liang and others, 2021
https://scholar.google.com/scholar?q=WILDS%3A+A+Benchmark+of+in-the-Wild+Distribution+Shifts
14. LLMs Cannot Reliably Judge (Yet?): A Comprehensive Assessment on the Robustness of LLM-as-a-Judge — Shuaizhi Li, Chenxu Xu, Jiazhu Wang, Xianyu Gong, Cheng Chen, Jun Zhang, Junjie Wang, Kit Lam, and Shouling Ji, 2025
https://scholar.google.com/scholar?q=LLMs+Cannot+Reliably+Judge+%28Yet%3F%29%3A+A+Comprehensive+Assessment+on+the+Robustness+of+LLM-as-a-Judge
15. Confusion Is the Final Barrier: Rethinking Jailbreak Evaluation and Investigating the Real Misuse Threat of LLMs — Yixin Yan, Shichao Sun, Zhen Wang, Yifan Lin, Zeyu Duan, Zhenzhen Zheng, Mingyu Liu, Zhenfei Yin, and Jie Zhang, 2025
https://scholar.google.com/scholar?q=Confusion+Is+the+Final+Barrier%3A+Rethinking+Jailbreak+Evaluation+and+Investigating+the+Real+Misuse+Threat+of+LLMs
16. Know Thy Judge: On the Robustness Meta-Evaluation of LLM Safety Judges — Francisco Eiras et al., 2025
https://scholar.google.com/scholar?q=Know+Thy+Judge%3A+On+the+Robustness+Meta-Evaluation+of+LLM+Safety+Judges
17. Comparison Requires Valid Measurement: Rethinking Attack Success Rate Comparisons in AI Red Teaming — Alexandra Chouldechova, A. Feder Cooper, Solon Barocas, Abhinav Palia, Dan Vann, Hanna Wallach, 2025/2026
https://scholar.google.com/scholar?q=Comparison+Requires+Valid+Measurement%3A+Rethinking+Attack+Success+Rate+Comparisons+in+AI+Red+Teaming
18. How to Correctly Report LLM-as-a-Judge Evaluations — Chungpa Lee et al., 2025
https://scholar.google.com/scholar?q=How+to+Correctly+Report+LLM-as-a-Judge+Evaluations
19. Embracing Ambiguity: Bayesian Nonparametrics and Stakeholder Participation for Ambiguity-Aware Safety Evaluation — Yanan Long, 2025/2026
https://scholar.google.com/scholar?q=Embracing+Ambiguity%3A+Bayesian+Nonparametrics+and+Stakeholder+Participation+for+Ambiguity-Aware+Safety+Evaluation
20. AI Post Transformers: Multidimensional Safety Evaluation of Frontier AI Models — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/multidimensional-safety-evaluation-of-frontier-ai-models/
21. AI Post Transformers: LLM Benchmark Robustness to Linguistic Variation — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/llm-benchmark-robustness-to-linguistic-variation/
22. AI Post Transformers: AI Agent Traps and Prompt Injection — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-02-ai-agent-traps-and-prompt-injection-7ce4ba.mp3
Interactive Visualization: When LLM Judges Become Coin Flips

This episode explores a 2026 paper, Generative Modeling via Drifting, which argues that the hard transport process behind modern generative models can be moved into training so that inference becomes a single forward pass. It explains the core idea of a pushforward distribution, introduces the paper’s notion of a drifting field that nudges generated samples toward the data distribution during optimization, and frames equilibrium as the point where those updates no longer need to move samples. The discussion compares this approach with GANs, diffusion models, flow matching, and other fast one-step systems, highlighting the tradeoff between low-latency generation and the quality advantages of multi-step correction. A listener would find it interesting because it lays out a possible new generative modeling paradigm and tests whether one-shot generation can become more than just an accelerated approximation of diffusion.

Sources:
1. Generative Modeling via Drifting — Mingyang Deng, He Li, Tianhong Li, Yilun Du, Kaiming He, 2026
http://arxiv.org/abs/2602.04770
2. Generative Adversarial Nets — Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, Yoshua Bengio, 2014
https://neurips.cc/virtual/2014/poster/4618
3. Generative Moment Matching Networks — Yujia Li, Kevin Swersky, Rich Zemel, 2015
https://proceedings.mlr.press/v37/li15.html
4. Consistency Models — Yang Song, Prafulla Dhariwal, Mark Chen, Ilya Sutskever, 2023
https://icml.cc/virtual/2023/poster/24593
5. Adversarial Diffusion Distillation — Stability AI researchers, 2023
https://stability.ai/research/adversarial-diffusion-distillation
6. A Kernel Two-Sample Test — Arthur Gretton, Karsten M. Borgwardt, Malte J. Rasch, Bernhard Scholkopf, Alexander Smola, 2012
https://www.jmlr.org/beta/papers/v13/gretton12a.html
7. Wasserstein GAN — Martin Arjovsky, Soumith Chintala, Leon Bottou, 2017
https://icml.cc/virtual/2017/poster/799
8. Density Estimation using Real NVP — Laurent Dinh, Jascha Sohl-Dickstein, Samy Bengio, 2017
https://openreview.net/forum?id=HkpbnH9lx
9. Glow: Generative Flow with Invertible 1x1 Convolutions — Diederik P. Kingma, Prafulla Dhariwal, 2018
https://openai.com/index/glow/
10. Flow Matching for Generative Modeling — Yaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel, Matt Le, 2023
https://openreview.net/forum?id=PqvMRDCJT9t
11. Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow — Xingchao Liu, Chengyue Gong, Qiang Liu, 2022
https://openreview.net/forum?id=gWxpdtQpiYV
12. Large Scale GAN Training for High Fidelity Natural Image Synthesis — Andrew Brock, Jeff Donahue, Karen Simonyan, 2018
https://huggingface.co/papers/1809.11096
13. Diffusion Models Beat GANs on Image Synthesis — Prafulla Dhariwal, Alex Nichol, 2021
https://openreview.net/forum?id=AAWuCvzaVt
14. High-Resolution Image Synthesis with Latent Diffusion Models — Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, Bjorn Ommer, 2022
https://openaccess.thecvf.com/content/CVPR2022/html/Rombach_High-Resolution_Image_Synthesis_With_Latent_Diffusion_Models_CVPR_2022_paper.html
15. Scalable Diffusion Models with Transformers — William Peebles, Saining Xie, 2023
https://openaccess.thecvf.com/content/ICCV2023/html/Peebles_Scalable_Diffusion_Models_with_Transformers_ICCV_2023_paper.html
16. Denoising Diffusion Probabilistic Models — Jonathan Ho, Ajay Jain, Pieter Abbeel, 2020
https://scholar.google.com/scholar?q=Denoising+Diffusion+Probabilistic+Models
17. Progressive Distillation for Fast Sampling of Diffusion Models — Tim Salimans, Jonathan Ho, 2022
https://scholar.google.com/scholar?q=Progressive+Distillation+for+Fast+Sampling+of+Diffusion+Models
18. Unsupervised Image-to-Image Translation Networks — Ferenc Huszar and coauthors are not cited here; instead the more relevant cited moment-matching line is:, 2015
https://scholar.google.com/scholar?q=Unsupervised+Image-to-Image+Translation+Networks
19. Auto-Encoding Variational Bayes — Diederik P. Kingma, Max Welling, 2013
https://scholar.google.com/scholar?q=Auto-Encoding+Variational+Bayes
20. One-step diffusion with distribution matching distillation — approx. diffusion-distillation literature, recent
https://scholar.google.com/scholar?q=One-step+diffusion+with+distribution+matching+distillation
21. One-step diffusion distillation via deep equilibrium models — approx. diffusion-distillation / equilibrium-model authors, recent
https://scholar.google.com/scholar?q=One-step+diffusion+distillation+via+deep+equilibrium+models
22. Discrete Flow Matching — approx. flow-matching authors, recent
https://scholar.google.com/scholar?q=Discrete+Flow+Matching
23. Elucidating the design choice of probability paths in flow matching for forecasting — approx. forecasting / flow-matching authors, recent
https://scholar.google.com/scholar?q=Elucidating+the+design+choice+of+probability+paths+in+flow+matching+for+forecasting
24. Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation — approx. hybrid AR-diffusion authors, recent
https://scholar.google.com/scholar?q=Mixed+Autoregressive+and+Diffusion+Transformers+for+Continuous+Image+Generation
25. ACDiT: Interpolating autoregressive conditional modeling and diffusion transformer — approx. ACDiT authors, 2025
https://scholar.google.com/scholar?q=ACDiT%3A+Interpolating+autoregressive+conditional+modeling+and+diffusion+transformer
26. AI Post Transformers: Paris: Decentralized Open-Weight Diffusion Model — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/paris-decentralized-open-weight-diffusion-model/
Interactive Visualization: Generative Modeling via Drifting in One Step

This episode explores LAPS, a serving system for large language models that treats long prompt prefills and short multi-turn re-prefills as fundamentally different workloads instead of batching them together. It explains why user-perceived latency, especially time to first token, suffers when tiny follow-up requests get stuck behind large compute-heavy context loads, and how LAPS models the boundary between compute-bound and memory-bound prefills to separate them more intelligently. The discussion covers LAPS’s dual-queue design, its temporal and spatial disaggregation strategies, and engineering choices like short-request waiting windows, length-aware smart batching, and CUDA Graph execution. Listeners would find it interesting because it connects low-level scheduling and KV-cache behavior to the everyday experience of whether chat systems feel fast and responsive.

Interactive Visualization: LAPS for Length-Aware LLM Serving
Sources:
1. LAPS: A Length-Aware-Prefill LLM Serving System — Jianshu She, Zonghang Li, Hongchao Du, Shangyu Wu, Wenhao Zheng, Eric Xing, Zhengzhong Liu, Huaxiu Yao, Jason Xue, Qirong Ho, 2026
http://arxiv.org/abs/2601.11589
2. ORCA: A Distributed Serving System for Transformer-Based Generative Models — Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim, Byung-Gon Chun, 2022
https://scholar.google.com/scholar?q=ORCA%3A+A+Distributed+Serving+System+for+Transformer-Based+Generative+Models
3. Efficient Memory Management for Large Language Model Serving with PagedAttention — Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, Ion Stoica, 2023
https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention
4. SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills — Amey Agrawal, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav S. Gulavani, Ramachandran Ramjee, 2023
https://scholar.google.com/scholar?q=SARATHI%3A+Efficient+LLM+Inference+by+Piggybacking+Decodes+with+Chunked+Prefills
5. DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving — Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, Hao Zhang, 2024
https://scholar.google.com/scholar?q=DistServe%3A+Disaggregating+Prefill+and+Decoding+for+Goodput-optimized+Large+Language+Model+Serving
6. BucketServe: Bucket-Based Dynamic Batching for Smart and Efficient LLM Inference Serving — Wanyi Zheng, Minxian Xu, Shengye Song, Kejiang Ye, 2025
https://scholar.google.com/scholar?q=BucketServe%3A+Bucket-Based+Dynamic+Batching+for+Smart+and+Efficient+LLM+Inference+Serving
7. DéjàVu: KV-cache Streaming for Fast, Fault-tolerant Generative LLM Serving — Foteini Strati, Sara McAllister, Amar Phanishayee, Jakub Tarnawski, Ana Klimovic, 2024
https://scholar.google.com/scholar?q=D%C3%A9j%C3%A0Vu%3A+KV-cache+Streaming+for+Fast%2C+Fault-tolerant+Generative+LLM+Serving
8. Prefill-Decode Aggregation or Disaggregation? Unifying Both for Goodput-Optimized LLM Serving — approx. contemporary LLM systems authors, 2025
https://scholar.google.com/scholar?q=Prefill-Decode+Aggregation+or+Disaggregation%3F+Unifying+Both+for+Goodput-Optimized+LLM+Serving
9. FlowPrefill: Decoupling Preemption from Prefill Scheduling Granularity to Mitigate Head-of-Line Blocking in LLM Serving — approx. contemporary LLM serving authors, 2025
https://scholar.google.com/scholar?q=FlowPrefill%3A+Decoupling+Preemption+from+Prefill+Scheduling+Granularity+to+Mitigate+Head-of-Line+Blocking+in+LLM+Serving
10. BurstGPT: A Real-World Workload Dataset to Optimize LLM Serving Systems — approx. systems/data-center workload authors, 2025
https://scholar.google.com/scholar?q=BurstGPT%3A+A+Real-World+Workload+Dataset+to+Optimize+LLM+Serving+Systems
11. SageServe: Optimizing LLM Serving on Cloud Data Centers with Forecast Aware Auto-Scaling — approx. cloud systems authors, 2025
https://scholar.google.com/scholar?q=SageServe%3A+Optimizing+LLM+Serving+on+Cloud+Data+Centers+with+Forecast+Aware+Auto-Scaling
12. ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production — approx. production-serving measurement authors, 2025
https://scholar.google.com/scholar?q=ServeGen%3A+Workload+Characterization+and+Generation+of+Large+Language+Model+Serving+in+Production
13. Fairness in Serving Large Language Models — approx. theory/systems fairness authors, 2025
https://scholar.google.com/scholar?q=Fairness+in+Serving+Large+Language+Models
14. FairBatching: Fairness-Aware Batch Formation for LLM Inference — approx. LLM inference scheduling authors, 2025
https://scholar.google.com/scholar?q=FairBatching%3A+Fairness-Aware+Batch+Formation+for+LLM+Inference
15. Locality-Aware Fair Scheduling in LLM Serving — approx. LLM serving systems authors, 2025
https://scholar.google.com/scholar?q=Locality-Aware+Fair+Scheduling+in+LLM+Serving
16. AI Post Transformers: Splitwise: Phase-Split LLM Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-26-splitwise-phase-split-llm-inference-e8945b.mp3
17. AI Post Transformers: Prefill-as-a-Service for Cross-Datacenter KV Cache — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-19-prefill-as-a-service-for-cross-datacente-7560be.mp3
18. AI Post Transformers: Speculative Decoding in Real vLLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-speculative-decoding-in-real-vllm-servin-6f4e2b.mp3
19. AI Post Transformers: Breaking the Prefix Barrier with Shared KV Cache — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-24-breaking-the-prefix-barrier-with-shared-a5e5a6.mp3
20. AI Post Transformers: Computation-Bandwidth-Memory Trade-offs for AI Infrastructure — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-09-computation-bandwidth-memory-trade-offs-a83f2b.mp3
Interactive Visualization: LAPS for Length-Aware LLM Serving

This episode explores Paul Werbos’s 1990 paper on Backpropagation Through Time and explains how ordinary backpropagation extends to systems whose state evolves over time. It walks through the core idea of unrolling a recurrent or dynamic system into a time-indexed computation graph, then applying reverse-mode differentiation to compute exact gradients across both layers and time steps. The discussion also places BPTT in historical context, connecting it to earlier work on backpropagation, automatic differentiation, and alternative recurrent learning methods like real-time recurrent learning. Listeners would find it interesting because it shows how a foundational training method for sequence models, control systems, and differentiable simulations emerged from a simple but powerful reframing of memory and time in neural computation.

Interactive Visualization: Backpropagation Through Time Explained
Sources:
1. Backpropagation Through Time Explained
https://podcast.do-not-panic.com/uploaded-pdfs/2026-05-08T22-17-15-230Z-Backpropagation-through-time-what-it-does-and-how-to-do-it.pdf
2. Learning representations by back-propagating errors — David E. Rumelhart, Geoffrey E. Hinton, Ronald J. Williams, 1986
https://scholar.google.com/scholar?q=Learning+representations+by+back-propagating+errors
3. A Learning Algorithm for Continually Running Fully Recurrent Neural Networks — Ronald J. Williams, David Zipser, 1989
https://scholar.google.com/scholar?q=A+Learning+Algorithm+for+Continually+Running+Fully+Recurrent+Neural+Networks
4. Backpropagation Through Time: What It Does and How to Do It — Paul J. Werbos, 1990
https://scholar.google.com/scholar?q=Backpropagation+Through+Time%3A+What+It+Does+and+How+to+Do+It
5. Learning long-term dependencies with gradient descent is difficult — Yoshua Bengio, Patrice Simard, Paolo Frasconi, 1994
https://scholar.google.com/scholar?q=Learning+long-term+dependencies+with+gradient+descent+is+difficult
6. Long Short-Term Memory — Sepp Hochreiter, Jürgen Schmidhuber, 1997
https://scholar.google.com/scholar?q=Long+Short-Term+Memory
7. Taylor expansion of the accumulated rounding error — Seppo Linnainmaa, 1976
https://scholar.google.com/scholar?q=Taylor+expansion+of+the+accumulated+rounding+error
8. Fast Exact Multiplication by the Hessian — Barak A. Pearlmutter, 1994
https://scholar.google.com/scholar?q=Fast+Exact+Multiplication+by+the+Hessian
9. Automatic Differentiation in Machine Learning: a Survey — Atilim Gunes Baydin, Barak A. Pearlmutter, Alexey Andreyevich Radul, Jeffrey Mark Siskind, 2018
https://scholar.google.com/scholar?q=Automatic+Differentiation+in+Machine+Learning%3A+a+Survey
10. A review of automatic differentiation and its efficient implementation — Charles C. Margossian, 2019
https://scholar.google.com/scholar?q=A+review+of+automatic+differentiation+and+its+efficient+implementation
11. Generalization of Back-Propagation to Recurrent Neural Networks — Fernando J. Pineda, 1987
https://scholar.google.com/scholar?q=Generalization+of+Back-Propagation+to+Recurrent+Neural+Networks
12. Finding Structure in Time — Jeffrey L. Elman, 1990
https://scholar.google.com/scholar?q=Finding+Structure+in+Time
13. BP(lambda): Online Learning via Synthetic Gradients — approx. anonymous from snippet / modern deep learning authors, recent
https://scholar.google.com/scholar?q=BP%28lambda%29%3A+Online+Learning+via+Synthetic+Gradients
14. Streaming Propagation Through Time: A New Computational Paradigm for Recurrent Neural Networks — approx. modern recurrent-learning authors, recent
https://scholar.google.com/scholar?q=Streaming+Propagation+Through+Time%3A+A+New+Computational+Paradigm+for+Recurrent+Neural+Networks
15. Combining Truncated BPTT and Truncated RTRL for LSTM Training — Jakob Stefan Weber, recent
https://scholar.google.com/scholar?q=Combining+Truncated+BPTT+and+Truncated+RTRL+for+LSTM+Training
16. Second-order forward-mode optimization of recurrent neural networks for neuroscience — approx. modern neuroscience/optimization authors, recent
https://scholar.google.com/scholar?q=Second-order+forward-mode+optimization+of+recurrent+neural+networks+for+neuroscience
17. Sample-Based Hybrid Mode Control: Asymptotically Optimal Switching of Algorithmic and Non-Differentiable Control Modes — approx. modern control authors, recent
https://scholar.google.com/scholar?q=Sample-Based+Hybrid+Mode+Control%3A+Asymptotically+Optimal+Switching+of+Algorithmic+and+Non-Differentiable+Control+Modes
18. On the differentiability of the value function of switched linear systems under arbitrary and controlled switching — approx. control theory authors, recent
https://scholar.google.com/scholar?q=On+the+differentiability+of+the+value+function+of+switched+linear+systems+under+arbitrary+and+controlled+switching
19. AI Post Transformers: Long Short-Term Memory and Vanishing Gradients — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-19-long-short-term-memory-and-vanishing-gra-72448c.mp3
20. AI Post Transformers: When Spectral Gradient Updates Help Deep Learning — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-when-spectral-gradient-updates-help-deep-9c8441.mp3
Interactive Visualization: Backpropagation Through Time Explained

This episode explores Marvin Minsky’s 1961 paper on whether mostly random neural networks can learn useful behavior simply by reinforcing successful responses. It explains how the paper distinguishes rote memory, associative recall, pattern recognition, and true generalization, arguing that reward signals alone are not enough unless the system already has a meaningful notion of similarity between situations. The discussion places that idea in context with early machine learning work like Rosenblatt’s perceptron and Samuel’s checkers program, then connects it to later, more disciplined descendants such as echo state networks and random features. Listeners get a sharp historical view of a debate that still matters now: whether intelligence comes from discovering good representations or from selecting among structures that were already there.

Sources:
1. Learning in Random Nets and Generalization
https://stacks.stanford.edu/file/druid:yr384hg3073/yr384hg3073.pdf
2. Learning in Random Nets — Marvin Minsky, Oliver G. Selfridge, 1961
https://scholar.google.com/scholar?q=Learning+in+Random+Nets
3. The Perceptron: A Probabilistic Model for Information Storage and Organization in the Brain — Frank Rosenblatt, 1958
https://scholar.google.com/scholar?q=The+Perceptron%3A+A+Probabilistic+Model+for+Information+Storage+and+Organization+in+the+Brain
4. The Echo State Approach to Analysing and Training Recurrent Neural Networks — Herbert Jaeger, 2001
https://scholar.google.com/scholar?q=The+Echo+State+Approach+to+Analysing+and+Training+Recurrent+Neural+Networks
5. Random Features for Large-Scale Kernel Machines — Ali Rahimi, Benjamin Recht, 2007
https://scholar.google.com/scholar?q=Random+Features+for+Large-Scale+Kernel+Machines
6. The Use of Multiple Measurements in Taxonomic Problems — R. A. Fisher, 1936
https://scholar.google.com/scholar?q=The+Use+of+Multiple+Measurements+in+Taxonomic+Problems
7. Nearest Neighbor Pattern Classification — Thomas M. Cover, Peter E. Hart, 1967
https://scholar.google.com/scholar?q=Nearest+Neighbor+Pattern+Classification
8. Support-Vector Networks — Corinna Cortes, Vladimir Vapnik, 1995
https://scholar.google.com/scholar?q=Support-Vector+Networks
9. Gradient-Based Learning Applied to Document Recognition — Yann LeCun, Leon Bottou, Yoshua Bengio, Patrick Haffner, 1998
https://scholar.google.com/scholar?q=Gradient-Based+Learning+Applied+to+Document+Recognition
10. Some Studies in Machine Learning Using the Game of Checkers — Arthur L. Samuel, 1959
https://scholar.google.com/scholar?q=Some+Studies+in+Machine+Learning+Using+the+Game+of+Checkers
11. Generalization of Pattern Recognition in a Self-Organizing System — B. G. Farley and W. A. Clark, 1954
https://scholar.google.com/scholar?q=Generalization+of+Pattern+Recognition+in+a+Self-Organizing+System
12. A Heterarchy of Values Determined by the Topology of Nervous Nets — Warren S. McCulloch, 1945
https://scholar.google.com/scholar?q=A+Heterarchy+of+Values+Determined+by+the+Topology+of+Nervous+Nets
13. Asymptotics of Random Feature Regression Beyond the Linear Scaling Regime — Hong Hu, Yue M. Lu, Theodor Misiakiewicz, 2024
https://scholar.google.com/scholar?q=Asymptotics+of+Random+Feature+Regression+Beyond+the+Linear+Scaling+Regime
14. Power-Law Spectrum of the Random Feature Model — Elliot Paquette, Ke Liang Xiao, Yizhe Zhu, 2026
https://scholar.google.com/scholar?q=Power-Law+Spectrum+of+the+Random+Feature+Model
15. Local to Global: Learning Dynamics and Effect of Initialization for Transformers — Ashok Vardhan Makkuva, Marco Bondaschi, Chanakya Ekbote, Adway Girish, Alliot Nagle, Hyeji Kim, Michael Gastpar, 2024
https://scholar.google.com/scholar?q=Local+to+Global%3A+Learning+Dynamics+and+Effect+of+Initialization+for+Transformers
16. Augmenting Language Models with Long-Term Memory — Weizhi Wang, Li Dong, Hao Cheng, Xiaodong Liu, Xifeng Yan, Jianfeng Gao, Furu Wei, 2023
https://scholar.google.com/scholar?q=Augmenting+Language+Models+with+Long-Term+Memory
17. MemoryBank: Enhancing Large Language Models with Long-Term Memory — Wanjun Zhong, Lianghong Guo, Qiqi Gao, He Ye, Yanlin Wang, 2023
https://scholar.google.com/scholar?q=MemoryBank%3A+Enhancing+Large+Language+Models+with+Long-Term+Memory
18. HippoRAG: Neurobiologically Inspired Long-Term Memory for Large Language Models — Bernal Jiménez Gutiérrez, Yiheng Shu, Yu Gu, Michihiro Yasunaga, Yu Su, 2024
https://scholar.google.com/scholar?q=HippoRAG%3A+Neurobiologically+Inspired+Long-Term+Memory+for+Large+Language+Models
19. The Learnability of In-Context Learning — Noam Wies, Yoav Levine, Amnon Shashua, 2023
https://scholar.google.com/scholar?q=The+Learnability+of+In-Context+Learning
20. Understanding In-Context Learning in Transformers and LLMs by Learning to Learn Discrete Functions — Satwik Bhattamishra, Arkil Patel, Phil Blunsom, Varun Kanade, 2023
https://scholar.google.com/scholar?q=Understanding+In-Context+Learning+in+Transformers+and+LLMs+by+Learning+to+Learn+Discrete+Functions
21. Learning without Training: The Implicit Dynamics of In-Context Learning — Benoit Dherin, Michael Munn, Hanna Mazzawi, Michael Wunder, Javier Gonzalvo, 2025
https://scholar.google.com/scholar?q=Learning+without+Training%3A+The+Implicit+Dynamics+of+In-Context+Learning
22. AI Post Transformers: Long Short-Term Memory and Vanishing Gradients — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-19-long-short-term-memory-and-vanishing-gra-72448c.mp3
23. AI Post Transformers: Latent Space as a New Computational Paradigm — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-05-latent-space-as-a-new-computational-para-810f39.mp3
24. AI Post Transformers: When Spectral Gradient Updates Help Deep Learning — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-when-spectral-gradient-updates-help-deep-9c8441.mp3
25. AI Post Transformers: LeWorldModel: Stable Joint-Embedding World Models from Pixels — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-25-leworldmodel-stable-joint-embedding-worl-650f9f.mp3
Interactive Visualization: Learning in Random Nets and Generalization

This episode explores the 1959 paper on frog vision that argued the retina does far more than passively relay a camera-like image to the brain. It explains how experiments on single optic nerve fibers revealed specialized visual detectors tuned to ecologically relevant signals such as small moving dark objects, edges, dimming, and contrast changes, rather than raw brightness alone. The discussion connects these findings to modern machine learning ideas like preprocessing, receptive fields, sparse event-driven signals, and early feature extraction, while also emphasizing where biological retinal circuits differ sharply from engineered neural networks. A listener would find it interesting because it shows how a foundational neuroscience experiment anticipated core ideas in AI and neural coding by asking what information an animal actually needs to survive.

Interactive Visualization: What the Frog’s Eye Tells the Brain
Sources:
1. What the Frog’s Eye Tells the Brain
https://courses.csail.mit.edu/6.803/pdf/lettvin.pdf
2. The Response of Single Optic Nerve Fibers of the Vertebrate Eye to Illumination of the Retina — H. Keffer Hartline, 1938
https://scholar.google.com/scholar?q=The+Response+of+Single+Optic+Nerve+Fibers+of+the+Vertebrate+Eye+to+Illumination+of+the+Retina
3. Discharge Patterns and Functional Organization of Mammalian Retina — Stephen W. Kuffler, 1953
https://scholar.google.com/scholar?q=Discharge+Patterns+and+Functional+Organization+of+Mammalian+Retina
4. Anatomy and Physiology of Vision in the Frog (Rana pipiens) — Humberto R. Maturana, Jerome Y. Lettvin, Warren S. McCulloch, Walter H. Pitts, 1960
https://scholar.google.com/scholar?q=Anatomy+and+Physiology+of+Vision+in+the+Frog+%28Rana+pipiens%29
5. The Dynamic Receptive Fields of Retinal Ganglion Cells — Sophia Wienbar, Gregory W. Schwartz, 2018
https://scholar.google.com/scholar?q=The+Dynamic+Receptive+Fields+of+Retinal+Ganglion+Cells
6. What the Frog's Eye Tells the Frog's Brain — Jerome Y. Lettvin, Humberto R. Maturana, Warren S. McCulloch, Walter H. Pitts, 1959
https://scholar.google.com/scholar?q=What+the+Frog%27s+Eye+Tells+the+Frog%27s+Brain
7. Summation and Inhibition in the Frog's Retina — Horace B. Barlow, 1953
https://scholar.google.com/scholar?q=Summation+and+Inhibition+in+the+Frog%27s+Retina
8. The Mechanism of Directionally Selective Units in Rabbit's Retina — Horace B. Barlow, William R. Levick, 1965
https://scholar.google.com/scholar?q=The+Mechanism+of+Directionally+Selective+Units+in+Rabbit%27s+Retina
9. The Retina Dissects the Visual Scene into Distinct Features — Botond Roska, Markus Meister, 2014
https://scholar.google.com/scholar?q=The+Retina+Dissects+the+Visual+Scene+into+Distinct+Features
10. Possible Principles Underlying the Transformations of Sensory Messages — Horace B. Barlow, 1961
https://scholar.google.com/scholar?q=Possible+Principles+Underlying+the+Transformations+of+Sensory+Messages
11. The Neural Code of the Retina — Markus Meister, Michael J. Berry II, 1999
https://scholar.google.com/scholar?q=The+Neural+Code+of+the+Retina
12. Weak Pairwise Correlations Imply Strongly Correlated Network States in a Neural Population — Elad Schneidman, Michael J. Berry II, Ronen Segev, William Bialek, 2006
https://scholar.google.com/scholar?q=Weak+Pairwise+Correlations+Imply+Strongly+Correlated+Network+States+in+a+Neural+Population
13. Spatio-temporal Correlations and Visual Signalling in a Complete Neuronal Population — Jonathan W. Pillow, Jonathon Shlens, Liam Paninski, Alexander Sher, Alan M. Litke, E. J. Chichilnisky, Eero P. Simoncelli, 2008
https://scholar.google.com/scholar?q=Spatio-temporal+Correlations+and+Visual+Signalling+in+a+Complete+Neuronal+Population
14. Receptive Fields of Single Neurones in the Cat's Striate Cortex — D. H. Hubel and T. N. Wiesel, 1959
https://scholar.google.com/scholar?q=Receptive+Fields+of+Single+Neurones+in+the+Cat%27s+Striate+Cortex
15. Interpreting the retinal neural code for natural scenes: From computations to neurons — Maheswaranathan, McIntosh, Tanaka, Baccus et al., 2023
https://scholar.google.com/scholar?q=Interpreting+the+retinal+neural+code+for+natural+scenes%3A+From+computations+to+neurons
16. Spatial adaptation of primate retinal ganglion cells between artificial and natural stimuli — Vystrcilova, Sridhar, Burg, Gollisch, Ecker et al., 2025/2026
https://scholar.google.com/scholar?q=Spatial+adaptation+of+primate+retinal+ganglion+cells+between+artificial+and+natural+stimuli
17. Distributed feature representations of natural stimuli across parallel retinal pathways — Hsiang, Shen, Soto, Kerschensteiner et al., 2024
https://scholar.google.com/scholar?q=Distributed+feature+representations+of+natural+stimuli+across+parallel+retinal+pathways
18. Retinal motion statistics during natural locomotion — Muller, Matthis, Bonnen, Cormack, Huk, Hayhoe, 2023
https://scholar.google.com/scholar?q=Retinal+motion+statistics+during+natural+locomotion
19. Natural visual behavior and active sensing in the mouse — review by members of the Niell lab and colleagues, 2024
https://scholar.google.com/scholar?q=Natural+visual+behavior+and+active+sensing+in+the+mouse
20. A genetically defined tecto-thalamic pathway drives a system of superior-colliculus-dependent visual cortices — Brenner, Beltramo, Gerfen, Ruediger, Scanziani, 2023
https://scholar.google.com/scholar?q=A+genetically+defined+tecto-thalamic+pathway+drives+a+system+of+superior-colliculus-dependent+visual+cortices
Interactive Visualization: What the Frog’s Eye Tells the Brain

This episode explores a mechanistic interpretability study asking whether a language model can detect when a concept has been injected into its hidden activations and, in some cases, identify what that concept was. It explains the difference between detection and identification, walks through activation steering in the residual stream, and highlights the paper’s controlled experiments on Gemma3-27B across 500 concepts, including a strong result of moderate detection with zero false positives under several prompt styles. The discussion also focuses on the paper’s argument that this reporting behavior emerges mainly during post-training, especially preference optimization, rather than from pretraining alone. Listeners would find it interesting because it turns a provocative claim about model “introspection” into a concrete circuit-level question about what internal features and gates may be doing.

Sources:
1. Mechanisms of Introspective Awareness — Uzay Macar, Li Yang, Atticus Wang, Peter Wallich, Emmanuel Ameisen, Jack Lindsey, 2026
http://arxiv.org/abs/2603.21396
2. Emergent Introspective Awareness in Large Language Models — Jack Lindsey, 2025
https://scholar.google.com/scholar?q=Emergent+Introspective+Awareness+in+Large+Language+Models
3. Looking Inward: Language Models Can Learn About Themselves by Introspection — Felix J. Binder, James Chua, Tomek Korbak, Henry Sleight, John Hughes, Robert Long, Ethan Perez, Miles Turpin, Owain Evans, 2024
https://scholar.google.com/scholar?q=Looking+Inward%3A+Language+Models+Can+Learn+About+Themselves+by+Introspection
4. Activation Addition: Steering Language Models Without Optimization — Alexander Matt Turner, Lisa Thiergart, David Udell, Gavin Leech, Ulisse Mini, Monte MacDiarmid, Chris Olah, 2023
https://scholar.google.com/scholar?q=Activation+Addition%3A+Steering+Language+Models+Without+Optimization
5. Representation Engineering: A Top-Down Approach to AI Transparency — Andy Zou, Long Phan, Sarah Chen, James Campbell, Richard Ngo, Adam Jermyn, Stephen McAleer, Alexander Tamkin, 2023
https://scholar.google.com/scholar?q=Representation+Engineering%3A+A+Top-Down+Approach+to+AI+Transparency
6. Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet — Adly Templeton, Tom Conerly, Jonathan Marcus, Jack Lindsey, Trenton Bricken, Brian Chen, Adam Pearce, Craig Citro, and others, 2024
https://scholar.google.com/scholar?q=Scaling+Monosemanticity%3A+Extracting+Interpretable+Features+from+Claude+3+Sonnet
7. Circuit Tracing: Revealing Computational Graphs in Language Models — Emmanuel Ameisen, Jack Lindsey, Adam Pearce, Wes Gurnee, Nicholas L. Turner, Brian Chen, Craig Citro, and others, 2025
https://scholar.google.com/scholar?q=Circuit+Tracing%3A+Revealing+Computational+Graphs+in+Language+Models
8. Steering Vector Fields for Context-Aware Inference-Time Control in Large Language Models — authors unclear from snippet, 2025/2026
https://scholar.google.com/scholar?q=Steering+Vector+Fields+for+Context-Aware+Inference-Time+Control+in+Large+Language+Models
9. No Training Wheels: Steering Vectors for Bias Correction at Inference Time — authors unclear from snippet, 2025/2026
https://scholar.google.com/scholar?q=No+Training+Wheels%3A+Steering+Vectors+for+Bias+Correction+at+Inference+Time
10. A mechanistic understanding of alignment algorithms: A case study on DPO and toxicity — authors unclear from snippet, 2025/2026
https://scholar.google.com/scholar?q=A+mechanistic+understanding+of+alignment+algorithms%3A+A+case+study+on+DPO+and+toxicity
11. How Does DPO Reduce Toxicity? A Mechanistic Neuron-Level Analysis — authors unclear from snippet, 2025/2026
https://scholar.google.com/scholar?q=How+Does+DPO+Reduce+Toxicity%3F+A+Mechanistic+Neuron-Level+Analysis
12. Refusal in language models is mediated by a single direction — authors unclear from snippet, 2025/2026
https://scholar.google.com/scholar?q=Refusal+in+language+models+is+mediated+by+a+single+direction
13. Beyond I'm Sorry, I Can't: Dissecting Large-Language-Model Refusal — authors unclear from snippet, 2025/2026
https://scholar.google.com/scholar?q=Beyond+I%27m+Sorry%2C+I+Can%27t%3A+Dissecting+Large-Language-Model+Refusal
14. Surgical, cheap, and flexible: Mitigating false refusal in language models via single vector ablation — authors unclear from snippet, 2025/2026
https://scholar.google.com/scholar?q=Surgical%2C+cheap%2C+and+flexible%3A+Mitigating+false+refusal+in+language+models+via+single+vector+ablation
15. Residual stream analysis with multi-layer saes — authors unclear from snippet, 2025/2026
https://scholar.google.com/scholar?q=Residual+stream+analysis+with+multi-layer+saes
16. AI Post Transformers: Anthropic: Introspective Awareness in LLMs — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/anthropic-introspective-awareness-in-llms/
17. AI Post Transformers: Neural Chameleons and Evading Activation Monitors — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-14-neural-chameleons-and-evading-activation-bc470e.mp3
18. AI Post Transformers: Advancing Mechanistic Interpretability with Sparse Autoencoders — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/advancing-mechanistic-interpretability-with-sparse-autoencoders/
19. AI Post Transformers: How Induction Heads Emerge in Transformers — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-03-how-induction-heads-emerge-in-transforme-a7bfcb.mp3
20. AI Post Transformers: Self-Improving Pretraining With Post-Trained Models — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-02-self-improving-pretraining-with-post-tra-e37460.mp3
Interactive Visualization: How Models Detect Hidden Activation Steering

This episode explores an FPGA-based time-to-digital converter that combines careful delay-line layout with machine-learning-based calibration to achieve very fine timing measurements on real hardware. It explains how tapped-delay-line TDCs work, why real FPGA implementations suffer from nonuniform time bins and bubble errors, and why those imperfections matter for applications like LiDAR, medical imaging, particle physics, and high-speed communications. The discussion compares the new approach against earlier FPGA TDC work, arguing that the real contribution is not flashy AI but a practical learned decoder that maps a 940-bit raw hardware output into a more accurate time estimate after physical design has reduced as much noise as possible. Listeners would find it interesting because it gets specific about where machine learning genuinely helps in instrumentation: not replacing physics, but reducing calibration effort while preserving picosecond-level precision.

Sources:
1. Machine Learning Self-Calibrated FPGA Time-to-Digital Converter
https://podcast.do-not-panic.com/uploaded-pdfs/2026-05-08T03-07-35-153Z-1-s2.0-S2667305326000190-main.pdf
2. A 19.6 ps, FPGA-Based TDC With Multiple Channels for Open Source Applications — Matthew W. Fishburn, L. Harmen Menninga, Claudio Favi, Edoardo Charbon, 2013
https://scholar.google.com/scholar?q=A+19.6+ps%2C+FPGA-Based+TDC+With+Multiple+Channels+for+Open+Source+Applications
3. A low nonlinearity, missing-code free time-to-digital converter based on 28nm FPGAs with embedded bin-width calibrations — Haochang Chen, Yongliang Zhang, David Day-Uei Li, 2017
https://scholar.google.com/scholar?q=A+low+nonlinearity%2C+missing-code+free+time-to-digital+converter+based+on+28nm+FPGAs+with+embedded+bin-width+calibrations
4. A 19 ps Precision and 170 M Samples/s Time-to-Digital Converter Implemented in FPGA with Online Calibration — Mengdi Zhang, Ye Zhao, Zhengsheng Han, Fazhan Zhao, 2022
https://scholar.google.com/scholar?q=A+19+ps+Precision+and+170+M+Samples%2Fs+Time-to-Digital+Converter+Implemented+in+FPGA+with+Online+Calibration
5. Low-Resource Time-to-Digital Converters for Field Programmable Gate Arrays: A Review — Diego Real, David Calvo, 2024
https://scholar.google.com/scholar?q=Low-Resource+Time-to-Digital+Converters+for+Field+Programmable+Gate+Arrays%3A+A+Review
6. Calibration Methods for Time-to-Digital Converters — Wassim Khaddour, Wilfried Uhring, Foudil Dadouche, Norbert Dumas, Morgan Madec, 2023
https://scholar.google.com/scholar?q=Calibration+Methods+for+Time-to-Digital+Converters
7. Time Resolution Improvement Using Dual Delay Lines for Field-Programmable-Gate-Array-Based Time-to-Digital Converters with Real-Time Calibration — Yuan-Ho Chen, 2019
https://scholar.google.com/scholar?q=Time+Resolution+Improvement+Using+Dual+Delay+Lines+for+Field-Programmable-Gate-Array-Based+Time-to-Digital+Converters+with+Real-Time+Calibration
8. Novel machine learning-driven optimizing decoding solutions for FPGA-based time-to-digital converters — Fabio Garzetti, Nicola Lusardi, Enrico Ronconi, Andrea Costa, Angelo Geraci, 2024
https://scholar.google.com/scholar?q=Novel+machine+learning-driven+optimizing+decoding+solutions+for+FPGA-based+time-to-digital+converters
9. A novel FPGA-based time-to-digital converter featuring machine learning-aided self-calibration — Arash Amini Bardpareh, Eleonora Vacca, Davide Nicolini, Corrado De Sio, Sarah Azimi, Luca Sterpone, Elisa Fiorina, Emanuele Maria Data, Felix Mas Milian, 2026
https://scholar.google.com/scholar?q=A+novel+FPGA-based+time-to-digital+converter+featuring+machine+learning-aided+self-calibration
10. Multiple-tapped-delay-line hardware-linearisation technique based on wire load regulation — Dariusz Chaberski, Robert Frankowski, Marek Zielinski, Lukasz Zaworski, 2016
https://scholar.google.com/scholar?q=Multiple-tapped-delay-line+hardware-linearisation+technique+based+on+wire+load+regulation
11. 5.7 ps Resolution Time-to-Digital Converter Implementation Using Routing Path Delays — Roza Teklehaimanot Siecha, Getachew Alemu, Jeffrey Prinzie, Paul Leroux, 2023
https://scholar.google.com/scholar?q=5.7+ps+Resolution+Time-to-Digital+Converter+Implementation+Using+Routing+Path+Delays
12. Tapped delay line for compact time-to-digital converter on UltraScale FPGA and its coding method — Min Zhu, Xihan Qi, Tang Cui, Qiang Gao, 2023
https://scholar.google.com/scholar?q=Tapped+delay+line+for+compact+time-to-digital+converter+on+UltraScale+FPGA+and+its+coding+method
13. A High-Resolution (Machine Learning Self-Calibrated FPGA Time-to-Digital Converter

This episode explores the TensorFlow paper as a systems argument for unifying the full machine learning lifecycle, from mobile inference to large-scale distributed training, within a single stateful dataflow framework. It explains how TensorFlow represents computation as graphs with mutable state, why that mattered for device placement, parameter storage, checkpointing, and heterogeneous hardware, and how it aimed to improve on the limitations of DistBelief. The discussion also places the paper in the broader lineage of MapReduce, Dryad, Naiad, and parameter-server training, while debating whether TensorFlow truly generalized machine learning workflows or mainly fit the kinds of static, graph-friendly workloads large organizations like Google already needed. Listeners would find it interesting for its mix of technical history, distributed systems insight, and a clear-eyed look at the tradeoff between organizational scale, portability, and usability for everyday researchers.

Sources:
1. TensorFlow: Large-Scale Machine Learning on Heterogeneous Distributed Systems — Martín Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S. Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Ian Goodfellow, Andrew Harp, Geoffrey Irving, Michael Isard, Yangqing Jia, Rafal Jozefowicz, Lukasz Kaiser, Manjunath Kudlur, Josh Levenberg, Dan Mane, Rajat Monga, Sherry Moore, Derek Murray, Chris Olah, Mike Schuster, Jonathon Shlens, Benoit Steiner, Ilya Sutskever, Kunal Talwar, Paul Tucker, Vincent Vanhoucke, Vijay Vasudevan, Fernanda Viegas, Oriol Vinyals, Pete Warden, Martin Wattenberg, Martin Wicke, Yuan Yu, Xiaoqiang Zheng, 2016
http://arxiv.org/abs/1603.04467
2. MapReduce: Simplified Data Processing on Large Clusters — Jeffrey Dean and Sanjay Ghemawat, 2004
https://scholar.google.com/scholar?q=MapReduce%3A+Simplified+Data+Processing+on+Large+Clusters
3. Dryad: Distributed Data-Parallel Programs from Sequential Building Blocks — Michael Isard, Mihai Budiu, Yuan Yu, Andrew Birrell, and Dennis Fetterly, 2007
https://scholar.google.com/scholar?q=Dryad%3A+Distributed+Data-Parallel+Programs+from+Sequential+Building+Blocks
4. Large Scale Distributed Deep Networks — Jeffrey Dean, Greg Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Quoc V. Le, Mark Z. Mao, Marc'Aurelio Ranzato, Andrew Senior, Paul Tucker, Ke Yang, and Andrew Y. Ng, 2012
https://scholar.google.com/scholar?q=Large+Scale+Distributed+Deep+Networks
5. TensorFlow: Large-Scale Machine Learning on Heterogeneous Distributed Systems — Martín Abadi, Ashish Agarwal, Paul Barham, Jeffrey Dean, Rajat Monga, and many others, 2016
https://scholar.google.com/scholar?q=TensorFlow%3A+Large-Scale+Machine+Learning+on+Heterogeneous+Distributed+Systems
6. Naiad: A Timely Dataflow System — Frank McSherry, Derek G. Murray, Rebecca Isaacs, and Michael Isard, 2013
https://scholar.google.com/scholar?q=Naiad%3A+A+Timely+Dataflow+System
7. Project Adam: Building an Efficient and Scalable Deep Learning Training System — Trishul Chilimbi, Yutaka Suzue, Johnson Apacible, and Karthik Kalyanaraman, 2014
https://scholar.google.com/scholar?q=Project+Adam%3A+Building+an+Efficient+and+Scalable+Deep+Learning+Training+System
8. Parameter Server for Distributed Machine Learning — Mu Li, David G. Andersen, Alexander J. Smola, and Kai Yu, 2014
https://scholar.google.com/scholar?q=Parameter+Server+for+Distributed+Machine+Learning
9. Caffe: Convolutional Architecture for Fast Feature Embedding — Yangqing Jia, Evan Shelhamer, Jeff Donahue, Sergey Karayev, Jonathan Long, Ross Girshick, Sergio Guadarrama, and Trevor Darrell, 2014
https://scholar.google.com/scholar?q=Caffe%3A+Convolutional+Architecture+for+Fast+Feature+Embedding
10. Theano: A CPU and GPU Math Compiler in Python — James Bergstra, Olivier Breuleux, Frédéric Bastien, Pascal Lamblin, Razvan Pascanu, Guillaume Desjardins, Joseph Turian, David Warde-Farley, and Yoshua Bengio, 2010
https://scholar.google.com/scholar?q=Theano%3A+A+CPU+and+GPU+Math+Compiler+in+Python
11. MXNet: A Flexible and Efficient Machine Learning Library for Heterogeneous Distributed Systems — Tianqi Chen, Mu Li, Yutian Li, Min Lin, Naiyan Wang, Minjie Wang, Tianjun Xiao, Bing Xu, Chiyuan Zhang, and Zheng Zhang, 2015
https://scholar.google.com/scholar?q=MXNet%3A+A+Flexible+and+Efficient+Machine+Learning+Library+for+Heterogeneous+Distributed+Systems
12. SIMPLE: Efficient Temporal Graph Neural Network Training at Scale with Dynamic Data Placement — Shihong Gao, Yiming Li, Xin Zhang, Yanyan Shen, Yingxia Shao, Lei Chen, 2024
https://scholar.google.com/scholar?q=SIMPLE%3A+Efficient+Temporal+Graph+Neural+Network+Training+at+Scale+with+Dynamic+Data+Placement
13. Strategy-Switch: From All-Reduce to Parameter Server for Faster Efficient Training — Nikodimos Provatas, Iasonas Chalas, Ioannis Konstantinou, Nectarios Koziris, 2025
https://scholar.google.com/scholar?q=Strategy-Switch%3A+From+All-Reduce+to+Parameter+Server+for+Faster+Efficient+Training
14. Dissecting the Runtime Performance of the Training, Fine-tuning, and Inference of Large Language Models — Longteng Zhang, Xiang Liu, Zeyu Li, Xinglin Pan, Peijie Dong, Ruibo Fan, Rui Guo, Xin Wang, Qiong Luo, Shaohuai Shi, Xiaowen Chu, 2023
https://scholar.google.com/scholar?q=Dissecting+the+Runtime+Performance+of+the+Training%2C+Fine-tuning%2C+and+Inference+of+Large+Language+Models
15. AI Post Transformers: ONNX Ecosystem, Optimization, and Deployment — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/onnx-ecosystem-optimization-and-deployment/
Interactive Visualization: TensorFlow for Distributed Machine Learning Systems

This episode explores SGLang, a system for making complex language model workflows run faster by treating them as full programs rather than single prompt-response calls. It explains how modern LLM applications involve branching, tool use, retries, and structured outputs, then examines SGLang’s co-design of a Python-embedded language with a specialized runtime that can optimize those patterns directly. The discussion highlights ideas like KV-cache reuse through RadixAttention, grammar-constrained decoding for reliable JSON output, and why these systems techniques matter more than just nicer prompt scripting. Listeners would find it interesting because it connects practical agent-style LLM engineering to deeper questions about compilers, serving infrastructure, and whether headline speedups really hold across real workloads.

Sources:
1. SGLang: Efficient Execution of Structured Language Model Programs — Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun, Jeff Huang, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonzalez, Clark Barrett, Ying Sheng, 2023
http://arxiv.org/abs/2312.07104
2. Grammar-Constrained Decoding for Structured NLP Tasks without Finetuning — Saibo Geng, Martin Josifoski, Maxime Peyrard, Robert West, 2023
https://scholar.google.com/scholar?q=Grammar-Constrained+Decoding+for+Structured+NLP+Tasks+without+Finetuning
3. SGLang: Efficient Execution of Structured Language Model Programs — Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun, Jeff Huang, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonzalez, Clark Barrett, Ying Sheng, 2024
https://scholar.google.com/scholar?q=SGLang%3A+Efficient+Execution+of+Structured+Language+Model+Programs
4. XGrammar: Flexible and Efficient Structured Generation Engine for Large Language Models — Yixin Dong, Charlie F. Ruan, Yaxing Cai, Ruihang Lai, Ziyi Xu, Yilong Zhao, Tianqi Chen, 2024
https://scholar.google.com/scholar?q=XGrammar%3A+Flexible+and+Efficient+Structured+Generation+Engine+for+Large+Language+Models
5. Generating Structured Outputs from Language Models: Benchmark and Studies — Saibo Geng, Hudson Cooper, Michał Moskal, Samuel Jenkins, Julian Berman, Nathan Ranchin, Robert West, Eric Horvitz, Harsha Nori, 2025
https://scholar.google.com/scholar?q=Generating+Structured+Outputs+from+Language+Models%3A+Benchmark+and+Studies
6. Language Model Cascades — David Dohan, Winnie Xu, Aitor Lewkowycz, Jacob Austin, David Bieber, Raphael Gontijo Lopes, Yuhuai Wu, Henryk Michalewski, Rif A. Saurous, Jascha Sohl-Dickstein, Kevin Murphy, Charles Sutton, 2022
https://scholar.google.com/scholar?q=Language+Model+Cascades
7. Prompting Is Programming: A Query Language for Large Language Models — Luca Beurer-Kellner, Marc Fischer, Martin Vechev, 2023
https://scholar.google.com/scholar?q=Prompting+Is+Programming%3A+A+Query+Language+for+Large+Language+Models
8. DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines — Omar Khattab, Arnav Singhvi, Paridhi Maheshwari, Zhiyuan Zhang, Keshav Santhanam, Sri Vardhamanan, Saiful Haq, Ashutosh Sharma, Thomas T. Joshi, Hanna Moazam, Heather Miller, Matei Zaharia, Christopher Potts, 2023
https://scholar.google.com/scholar?q=DSPy%3A+Compiling+Declarative+Language+Model+Calls+into+Self-Improving+Pipelines
9. vLLM: Easy, Fast, and Cheap LLM Serving with PagedAttention — Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, et al., 2023
https://scholar.google.com/scholar?q=vLLM%3A+Easy%2C+Fast%2C+and+Cheap+LLM+Serving+with+PagedAttention
10. Guidance: A Guidance Language for Controlling Large Language Models — Microsoft Research and collaborators, 2023
https://scholar.google.com/scholar?q=Guidance%3A+A+Guidance+Language+for+Controlling+Large+Language+Models
11. LMQL: A Programming Language for Large Language Models — Luca Beurer-Kellner, Marc Fischer, Martin Vechev, 2023
https://scholar.google.com/scholar?q=LMQL%3A+A+Programming+Language+for+Large+Language+Models
12. Outlines — Thibault Glaunec and contributors, 2023
https://scholar.google.com/scholar?q=Outlines
13. KVLink: Accelerating Large Language Models via Efficient KV Cache Reuse — Jingbo Yang, Bairu Hou, Wei Wei, Yujia Bao, Shiyu Chang, 2025
https://scholar.google.com/scholar?q=KVLink%3A+Accelerating+Large+Language+Models+via+Efficient+KV+Cache+Reuse
14. Key, Value, Compress: A Systematic Exploration of KV Cache Compression Techniques — Neusha Javidnia, Bita Rouhani, Farinaz Koushanfar, 2025
https://scholar.google.com/scholar?q=Key%2C+Value%2C+Compress%3A+A+Systematic+Exploration+of+KV+Cache+Compression+Techniques
15. KV-CAR: KV Cache Compression using Autoencoders and KV Reuse in Large Language Models — Sourjya Roy, Shrihari Sridharan, Surya Selvam, Anand Raghunathan, 2025
https://scholar.google.com/scholar?q=KV-CAR%3A+KV+Cache+Compression+using+Autoencoders+and+KV+Reuse+in+Large+Language+Models
16. Grammar-Aligned Decoding — Kanghee Park, Jiayu Wang, Taylor Berg-Kirkpatrick, Nadia Polikarpova, Loris D'Antoni, 2024
https://scholar.google.com/scholar?q=Grammar-Aligned+Decoding
17. Grammar-Constrained Decoding Makes Large Language Models Better Logical Parsers — Federico Raspanti, Tanir Ozcelebi, Mike J. Holenderski, 2025
https://scholar.google.com/scholar?q=Grammar-Constrained+Decoding+Makes+Large+Language+Models+Better+Logical+Parsers
18. Marconi: Prefix Caching for the Era of Hybrid LLMs — Rui Pan, Zhuang Wang, Zhen Jia, Can Karakus, Luca Zancato, Tri Dao, Yida Wang, Ravi Netravali, 2025
https://scholar.google.com/scholar?q=Marconi%3A+Prefix+Caching+for+the+Era+of+Hybrid+LLMs
19. Towards Efficient Agents: A Co-Design of Inference Architecture and System — Weizhe Lin, Hui-Ling Zhen, Shuai Yang, Xian Wang, Renxi Liu, Hanting Chen, Wangze Zhang, Chuansai Zhou, Yiming Li, Chen Chen, Xing Li, Zhiyuan Yang, Xiaosong Li, Xianzhi Yu, Zhenhua Dong, Mingxuan Yuan, Yunhe Wang, 2025
https://scholar.google.com/scholar?q=Towards+Efficient+Agents%3A+A+Co-Design+of+Inference+Architecture+and+System
20. Optimizing Agentic Language Model Inference via Speculative Tool Calls — Daniel Nichols, Prajwal Singhania, Charles Jekel, Abhinav Bhatele, Harshitha Menon, 2025
https://scholar.google.com/scholar?q=Optimizing+Agentic+Language+Model+Inference+via+Speculative+Tool+Calls
21. AI Post Transformers: SGLang: Efficient Language Model Program Execution — Hal Turing & Dr. Ada Shannon, Sun,
https://podcast.do-not-panic.com/episodes/sglang-efficient-language-model-program-execution/
22. AI Post Transformers: Breaking the Prefix Barrier with Shared KV Cache — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-24-breaking-the-prefix-barrier-with-shared-a5e5a6.mp3
23. AI Post Transformers: Continuous Batching for LLM Inference: Throughput and Latency Gains — Hal Turing & Dr. Ada Shannon, Mon,
https://podcast.do-not-panic.com/episodes/continuous-batching-for-llm-inference-throughput-and-latency-gains/
24. AI Post Transformers: Speculative Decoding in Real vLLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-speculative-decoding-in-real-vllm-servin-6f4e2b.mp3
25. AI Post Transformers: KV Cache TTL for Multi-Turn Agent Scheduling — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-09-kv-cache-ttl-for-multi-turn-agent-schedu-996bf1.mp3
26. AI Post Transformers: TokenDance for Multi-Agent KV Cache Sharing — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-22-tokendance-for-multi-agent-kv-cache-shar-aa9b99.mp3
Interactive Visualization: SGLang for Faster Structured LLM Programs

This episode explores CL-BENCH, a benchmark designed to test whether language models can actually learn task-specific knowledge from long, messy context and then reason with it, rather than merely retrieving facts or mimicking examples. It explains the distinction between long-context understanding, in-context learning, and the stronger notion of context learning, using examples like legal codes, product manuals, and experimental notebooks to show what real-world adaptation demands. The discussion highlights how the benchmark’s 500 contexts, 1,899 tasks, and dense binary verification rubrics are built to stress models on rule-following, procedural reasoning, and inferring governing relationships from data. Listeners would find it interesting because it gets at a central question in modern AI: whether bigger context windows actually make systems more capable, or just better at holding more text without truly learning from it.

Interactive Visualization: Can Models Learn from Long Context?
Sources:
1. CL-bench: A Benchmark for Context Learning — Shihan Dou, Ming Zhang, Zhangyue Yin, Chenhao Huang, Yujiong Shen, Junzhe Wang, Jiayi Chen, Yuchen Ni, Junjie Ye, Cheng Zhang, Huaibing Xie, Jianglu Hu, Shaolei Wang, Weichao Wang, Yanling Xiao, Yiting Liu, Zenan Xu, Zhen Guo, Pluto Zhou, Tao Gui, Zuxuan Wu, Xipeng Qiu, Qi Zhang, Xuanjing Huang, Yu-Gang Jiang, Di Wang, Shunyu Yao, 2026
http://arxiv.org/abs/2602.03587
2. Language Models are Few-Shot Learners — Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan and others, 2020
https://scholar.google.com/scholar?q=Language+Models+are+Few-Shot+Learners
3. MetaICL: Learning to Learn In Context — Sewon Min, Mike Lewis, Luke Zettlemoyer, Hannaneh Hajishirzi, 2021
https://scholar.google.com/scholar?q=MetaICL%3A+Learning+to+Learn+In+Context
4. Transformers learn in-context by gradient descent — Johannes von Oswald, Eyvind Niklasson, Ettore Randazzo, Joao Sacramento, Alexander Mordvintsev, Andrey Zhmoginov, Max Vladymyrov, 2022
https://scholar.google.com/scholar?q=Transformers+learn+in-context+by+gradient+descent
5. Lost in the Middle: How Language Models Use Long Contexts — Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, Percy Liang, 2023
https://scholar.google.com/scholar?q=Lost+in+the+Middle%3A+How+Language+Models+Use+Long+Contexts
6. LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding — Yushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu, Jiankai Tang, Zhidian Huang, Zhengxiao Du, Xiao Liu, Aohan Zeng, Lei Hou, Yuxiao Dong, Jie Tang, Juanzi Li, 2023
https://scholar.google.com/scholar?q=LongBench%3A+A+Bilingual%2C+Multitask+Benchmark+for+Long+Context+Understanding
7. BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack — Yurii Kuratov, Aydar Bulatov, Petr Anokhin, Ivan Rodkin, Dmitry Sorokin, Artyom Sorokin, Mikhail Burtsev, 2024
https://scholar.google.com/scholar?q=BABILong%3A+Testing+the+Limits+of+LLMs+with+Long+Context+Reasoning-in-a-Haystack
8. LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks — Yushi Bai, Shangqing Tu, Jiajie Zhang, Hao Peng, Xiaozhi Wang, Xin Lv, Shulin Cao, Jiazheng Xu, Lei Hou, Yuxiao Dong, Jie Tang, Juanzi Li, 2024
https://scholar.google.com/scholar?q=LongBench+v2%3A+Towards+Deeper+Understanding+and+Reasoning+on+Realistic+Long-context+Multitasks
9. NoLiMa: Long-Context Evaluation Beyond Literal Matching — Ali Modarressi, Hanieh Deilamsalehy, Franck Dernoncourt, Trung Bui, Ryan A. Rossi, David Seunghyun Yoon, Hinrich Schutze, 2025
https://scholar.google.com/scholar?q=NoLiMa%3A+Long-Context+Evaluation+Beyond+Literal+Matching
10. LongReason: A Synthetic Long-Context Reasoning Benchmark via Context Expansion — Zhan Ling et al., 2025
https://scholar.google.com/scholar?q=LongReason%3A+A+Synthetic+Long-Context+Reasoning+Benchmark+via+Context+Expansion
11. DocPuzzle: A Process-Aware Benchmark for Evaluating Realistic Long-Context Reasoning Capabilities — Tianyi Zhuang et al., 2025
https://scholar.google.com/scholar?q=DocPuzzle%3A+A+Process-Aware+Benchmark+for+Evaluating+Realistic+Long-Context+Reasoning+Capabilities
12. In-Context Learning Creates Task Vectors — Roee Hendel, Mor Geva, Amir Globerson, 2023
https://scholar.google.com/scholar?q=In-Context+Learning+Creates+Task+Vectors
13. In-context Vectors: Making In Context Learning More Effective and Controllable Through Latent Space Steering — Sheng Liu, Haotian Ye, Lei Xing, James Zou, 2024
https://scholar.google.com/scholar?q=In-context+Vectors%3A+Making+In+Context+Learning+More+Effective+and+Controllable+Through+Latent+Space+Steering
14. Task Vectors in In-Context Learning: Emergence, Formation, and Benefit — Liu Yang, Ziqian Lin, Kangwook Lee, Dimitris Papailiopoulos, Robert Nowak, 2025
https://scholar.google.com/scholar?q=Task+Vectors+in+In-Context+Learning%3A+Emergence%2C+Formation%2C+and+Benefit
15. Learn to Memorize: Scalable Continual Learning in Semiparametric Models with Mixture-of-Neighbors Induction Memory — Guangyue Peng, Tao Ge, Wen Luo, Wei Li, Houfeng Wang, 2025
https://scholar.google.com/scholar?q=Learn+to+Memorize%3A+Scalable+Continual+Learning+in+Semiparametric+Models+with+Mixture-of-Neighbors+Induction+Memory
16. AI Post Transformers: How Induction Heads Emerge in Transformers — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-03-how-induction-heads-emerge-in-transforme-a7bfcb.mp3
17. AI Post Transformers: Real Context Size and Context Rot — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-07-real-context-size-and-context-rot-56cbb4.mp3
18. AI Post Transformers: DeepSeek-V4 and Practical Million-Token Context — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-25-deepseek-v4-and-practical-million-token-6f4de1.mp3
19. AI Post Transformers: In-Place Test-Time Training for Transformers — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-09-in-place-test-time-training-for-transfor-d0b976.mp3
20. AI Post Transformers: Training LLMs for Divide-and-Conquer Reasoning — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-04-training-llms-for-divide-and-conquer-rea-ea6e22.mp3
21. AI Post Transformers: Inverse IFEval: Unlearning LLM Cognitive Inertia — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/inverse-ifeval-unlearning-llm-cognitive-inertia/
Interactive Visualization: Can Models Learn from Long Context?

This episode explores the 1987 paper on synchronous data flow and how it turns stream-processing programs into analyzable graphs with fixed token production and consumption rates. It explains how those fixed rates let a compiler precompute a repeating execution schedule, prove steady-state consistency through balance equations, and allocate bounded buffers ahead of time instead of relying on expensive runtime scheduling. The discussion highlights why that tradeoff works so well for digital signal processing workloads like filtering, resampling, and codecs, while also showing why the model is too restrictive for messier software with irregular control flow. Listeners would find it interesting because it shows how a carefully limited programming model can unlock strong guarantees about performance, memory use, and parallel execution.

Sources:
1. Synchronous Data Flow for Signal Processing
https://ptolemy.berkeley.edu/publications/papers/87/synchdataflow/synchdataflow.pdf
2. Synchronous Data Flow — Edward A. Lee and David G. Messerschmitt, 1987
https://scholar.google.com/scholar?q=Synchronous+Data+Flow
3. Dataflow Process Networks — Edward A. Lee and Thomas M. Parks, 1995
https://scholar.google.com/scholar?q=Dataflow+Process+Networks
4. Cycle-Static Dataflow — Greet Bilsen, Marc Engels, Rudy Lauwereins, and Jean Peperstraete, 1996
https://scholar.google.com/scholar?q=Cycle-Static+Dataflow
5. Static Scheduling of Synchronous Data Flow Programs for Digital Signal Processing — Edward A. Lee and David G. Messerschmitt, 1987
https://scholar.google.com/scholar?q=Static+Scheduling+of+Synchronous+Data+Flow+Programs+for+Digital+Signal+Processing
6. Bounded Scheduling of Process Networks — Thomas M. Parks, 1995
https://scholar.google.com/scholar?q=Bounded+Scheduling+of+Process+Networks
7. Synthesis of Embedded Software from Synchronous Dataflow Specifications — Shuvra S. Bhattacharyya, Praveen K. Murthy, and Edward A. Lee, 1999
https://scholar.google.com/scholar?q=Synthesis+of+Embedded+Software+from+Synchronous+Dataflow+Specifications
8. StreamIt: A Language for Streaming Applications — William Thies, Michal Karczmarek, and Saman Amarasinghe, 2002
https://scholar.google.com/scholar?q=StreamIt%3A+A+Language+for+Streaming+Applications
9. Memory Management for Dataflow Programming of Multirate Signal Processing Algorithms — Shuvra S. Bhattacharyya and Edward A. Lee, 1994
https://scholar.google.com/scholar?q=Memory+Management+for+Dataflow+Programming+of+Multirate+Signal+Processing+Algorithms
10. Joint Minimization of Code and Data for Synchronous Dataflow Programs — Praveen K. Murthy, Shuvra S. Bhattacharyya, and Edward A. Lee, 1994
https://scholar.google.com/scholar?q=Joint+Minimization+of+Code+and+Data+for+Synchronous+Dataflow+Programs
11. Buffer Merging: A Powerful Technique for Reducing Memory Requirements of Synchronous Dataflow Specifications — Praveen K. Murthy and Shuvra S. Bhattacharyya, 2000
https://scholar.google.com/scholar?q=Buffer+Merging%3A+A+Powerful+Technique+for+Reducing+Memory+Requirements+of+Synchronous+Dataflow+Specifications
12. Pipeline Interleaved Programmable DSP's: Synchronous Data Flow Programming — Edward A. Lee and David G. Messerschmitt, 1987
https://scholar.google.com/scholar?q=Pipeline+Interleaved+Programmable+DSP%27s%3A+Synchronous+Data+Flow+Programming
13. Multirate Digital Filters, Filter Banks, Polyphase Networks, and Applications: A Tutorial — P. P. Vaidyanathan, 1990
https://scholar.google.com/scholar?q=Multirate+Digital+Filters%2C+Filter+Banks%2C+Polyphase+Networks%2C+and+Applications%3A+A+Tutorial
14. The Semantics of a Simple Language for Parallel Programming — Gilles Kahn, 1974
https://scholar.google.com/scholar?q=The+Semantics+of+a+Simple+Language+for+Parallel+Programming
15. First Version of a Data Flow Procedure Language — Jack B. Dennis, 1974
https://scholar.google.com/scholar?q=First+Version+of+a+Data+Flow+Procedure+Language
16. On the Boundedness of Process Networks — Gilles Kahn and David B. MacQueen, 1977
https://scholar.google.com/scholar?q=On+the+Boundedness+of+Process+Networks
17. Algorithm Design for Signal Processing — Charles S. Burrus, 1982
https://scholar.google.com/scholar?q=Algorithm+Design+for+Signal+Processing
18. DynVec: An End-to-End Framework for Efficient Vector-Dataflow Execution — approximate; recent systems/compiler authors, recent
https://scholar.google.com/scholar?q=DynVec%3A+An+End-to-End+Framework+for+Efficient+Vector-Dataflow+Execution
19. Compiler discovered dynamic scheduling of irregular code in high-level synthesis — approximate; recent HLS/compiler authors, recent
https://scholar.google.com/scholar?q=Compiler+discovered+dynamic+scheduling+of+irregular+code+in+high-level+synthesis
20. Dataflow Models of computation for programming heterogeneous multicores — approximate; recent embedded/parallel-systems authors, recent
https://scholar.google.com/scholar?q=Dataflow+Models+of+computation+for+programming+heterogeneous+multicores
21. Heuristic & Expert-Guided Buffer Sizing for Neural Network Inference Applications on FPGAs — approximate; recent FPGA/dataflow authors, recent
https://scholar.google.com/scholar?q=Heuristic+%26+Expert-Guided+Buffer+Sizing+for+Neural+Network+Inference+Applications+on+FPGAs
22. Sgcn: Exploiting compressed-sparse features in deep graph convolutional network accelerators — approximate; recent accelerator authors, recent
https://scholar.google.com/scholar?q=Sgcn%3A+Exploiting+compressed-sparse+features+in+deep+graph+convolutional+network+accelerators
23. Safe shared state in dataflow systems — approximate; recent programming-systems authors, recent
https://scholar.google.com/scholar?q=Safe+shared+state+in+dataflow+systems
24. An Intermediate Representation for Stateful Dataflows — approximate; recent systems authors, recent
https://scholar.google.com/scholar?q=An+Intermediate+Representation+for+Stateful+Dataflows
25. AI Post Transformers: FlatAttention for Tile-Based Accelerator Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-flatattention-for-tile-based-accelerator-56e6ca.mp3
26. AI Post Transformers: Caffeine: A Unified FPGA for CNNs — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-06-caffeine-a-unified-fpga-for-cnns-e8acbe.mp3
27. AI Post Transformers: Caffe and the Rise of CNN Frameworks — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-06-caffe-and-the-rise-of-cnn-frameworks-cf15f3.mp3
Interactive Visualization: Synchronous Data Flow for Signal Processing

This episode explores FP-DNN, a 2017 framework that aims to compile TensorFlow-era neural networks onto FPGAs automatically, reducing the need for hand-designed accelerators for each model. It explains how the system maps convolutional layers, fully connected layers, and parts of LSTM computation into a shared matrix-multiplication core, while combining hand-tuned RTL for performance-critical components with HLS-generated logic for orchestration and layer-specific handling. The discussion highlights why this hybrid design matters for performance-per-watt, latency, and communication efficiency, especially as deeper CNNs and recurrent models were pushing hardware limits. Listeners would find it interesting for its clear look at an early attempt to turn FPGA deployment from an expert-only craft into a more reusable compiler-driven workflow, while also showing where the paper’s claims about broad model coverage may be too optimistic.

Sources:
1. Automating DNN Compilation for FPGA Accelerators
https://ceca.pku.edu.cn/media/lw/e3d0e0cd92452e0504b148220d442b9a.pdf
2. A Survey of FPGA-based Neural Network Inference Accelerators — Kaiyuan Guo, Shulin Zeng, Jincheng Yu, Yu Wang, Huazhong Yang, 2019
https://scholar.google.com/scholar?q=A+Survey+of+FPGA-based+Neural+Network+Inference+Accelerators
3. DeepBurning: Automatic Generation of FPGA-based Learning Accelerators for the Neural Network Family — Ying Wang, Jie Xu, Yudeng Sun, Baohua Cao, Chunyuan Xu, Yibo Kong, Chundao Han, Xuan Wang, 2016
https://scholar.google.com/scholar?q=DeepBurning%3A+Automatic+Generation+of+FPGA-based+Learning+Accelerators+for+the+Neural+Network+Family
4. FP-DNN: An Automated Framework for Mapping Deep Neural Networks onto FPGAs with RTL-HLS Hybrid Templates — Yijin Guan, Hao Liang, Ningyi Xu, Wenqiang Wang, Shaoshuai Shi, Xi Chen, Guangyu Sun, Wei Zhang, Jason Cong, 2017
https://scholar.google.com/scholar?q=FP-DNN%3A+An+Automated+Framework+for+Mapping+Deep+Neural+Networks+onto+FPGAs+with+RTL-HLS+Hybrid+Templates
5. DNNBuilder: An Automated Tool for Building High-Performance DNN Hardware Accelerators for FPGAs — Xiaofan Zhang, Junsong Wang, Chao Zhu, Yonghua Lin, Jinjun Xiong, Wen-Mei Hwu, Deming Chen, 2018
https://scholar.google.com/scholar?q=DNNBuilder%3A+An+Automated+Tool+for+Building+High-Performance+DNN+Hardware+Accelerators+for+FPGAs
6. From High-Level Deep Neural Models to FPGAs — Hardik Sharma, Jongse Park, Emmanuel Amaro, Bradley Thwaites, Priyanka Kotha, Anmol Gupta, Joon Kyung Kim, Asit Mishra, and Hsien-Hsin S. Lee, 2016
https://scholar.google.com/scholar?q=From+High-Level+Deep+Neural+Models+to+FPGAs
7. Caffeine: Towards Uniformed Representation and Acceleration for Deep Convolutional Neural Networks — Chen Zhang, Zhenman Fang, Peipei Zhou, Peichen Pan, and Jason Cong, 2016
https://scholar.google.com/scholar?q=Caffeine%3A+Towards+Uniformed+Representation+and+Acceleration+for+Deep+Convolutional+Neural+Networks
8. Throughput-Optimized OpenCL-Based FPGA Accelerator for Large-Scale Convolutional Neural Networks — Naveen Suda, Vikas Chandra, Ganesh Dasika, Abinash Mohanty, Yufei Ma, Sarita Vrudhula, Jae-sun Seo, and Yu Cao, 2016
https://scholar.google.com/scholar?q=Throughput-Optimized+OpenCL-Based+FPGA+Accelerator+for+Large-Scale+Convolutional+Neural+Networks
9. Going Deeper with Embedded FPGA Platform for Convolutional Neural Network — Jiantao Qiu, Jie Wang, Song Yao, Kai Guo, Boxun Li, Erjin Zhou, Jincheng Yu, Tianqi Tang, Ningyi Xu, and Song Wang, 2016
https://scholar.google.com/scholar?q=Going+Deeper+with+Embedded+FPGA+Platform+for+Convolutional+Neural+Network
10. Optimizing FPGA-Based Accelerator Design for Deep Convolutional Neural Networks — Chen Zhang, Peng Li, Guangyu Sun, Yijin Guan, Bingjun Xiao, and Jason Cong, 2015
https://scholar.google.com/scholar?q=Optimizing+FPGA-Based+Accelerator+Design+for+Deep+Convolutional+Neural+Networks
11. Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding — Song Han, Huizi Mao, and William J. Dally, 2015
https://scholar.google.com/scholar?q=Deep+Compression%3A+Compressing+Deep+Neural+Networks+with+Pruning%2C+Trained+Quantization+and+Huffman+Coding
12. BladeDISC: Optimizing Dynamic Shape Machine Learning Workloads via Compiler Approach — Zhen Zheng et al., 2023
https://scholar.google.com/scholar?q=BladeDISC%3A+Optimizing+Dynamic+Shape+Machine+Learning+Workloads+via+Compiler+Approach
13. TSCompiler: efficient compilation framework for dynamic-shape models — Xiang Luo, Chen Zhang, Chenbo Geng, Yanzhi Yi, Jiahui Hu, Renwei Zhang, Zhen Zhang, Gianpietro Consolaro, Fan Yang, Tun Lu, Ning Gu, Li Shang, 2024
https://scholar.google.com/scholar?q=TSCompiler%3A+efficient+compilation+framework+for+dynamic-shape+models
14. TATAA: Programmable Mixed-Precision Transformer Acceleration with a Transformable Arithmetic Architecture — Jiajun Wu, Mo Song, Jingmin Zhao, Yizhao Gao, Jia Li, Hayden Kwok-Hay So, 2024
https://scholar.google.com/scholar?q=TATAA%3A+Programmable+Mixed-Precision+Transformer+Acceleration+with+a+Transformable+Arithmetic+Architecture
15. FPGA Acceleration With Hessian-Based Comprehensive Intra-Layer Mixed-Precision Quantization for Transformer Models — Woohong Byun, Jongseok Woo, Saibal Mukhopadhyay, 2025
https://scholar.google.com/scholar?q=FPGA+Acceleration+With+Hessian-Based+Comprehensive+Intra-Layer+Mixed-Precision+Quantization+for+Transformer+Models
16. Understand and Accelerate Memory Processing Pipeline for Disaggregated LLM Inference — Zifan He, Rui Ma, Yizhou Sun, Jason Cong, 2026
https://scholar.google.com/scholar?q=Understand+and+Accelerate+Memory+Processing+Pipeline+for+Disaggregated+LLM+Inference
17. AI Post Transformers: FPGA Neural Network Accelerators for Space — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-26-fpga-neural-network-accelerators-for-spa-3087ae.mp3
18. AI Post Transformers: Continuous Batching for LLM Inference: Throughput and Latency Gains — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/continuous-batching-for-llm-inference-throughput-and-latency-gains/
19. AI Post Transformers: Advancements in Efficient KV Cache Quantization and Management — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/advancements-in-efficient-kv-cache-quantization-and-management/
Interactive Visualization: Automating DNN Compilation for FPGA Accelerators

This episode explores why Caffe mattered as a systems breakthrough for the early CNN era, even though it did not introduce a new learning algorithm. It explains how the framework helped researchers and engineers move from handcrafted vision features to learned feature embeddings, and why separating model definition from implementation made experimentation and deployment far more practical. The discussion highlights Caffe’s use of declarative Protocol Buffers configurations, directed acyclic graph model structure, and the blob abstraction that hid CPU versus GPU details while supporting modular extensions. Listeners would find it interesting for its clear account of how deep learning became usable at scale in 2014, and for its nuanced take on Caffe’s evidence: strong engineering promises, impressive throughput figures, and a major role in shaping the emerging model-development ecosystem.

Interactive Visualization: Caffe and the Rise of CNN Frameworks
Sources:
1. Caffe: Convolutional Architecture for Fast Feature Embedding — Yangqing Jia, Evan Shelhamer, Jeff Donahue, Sergey Karayev, Jonathan Long, Ross Girshick, Sergio Guadarrama, Trevor Darrell, 2014
http://arxiv.org/abs/1408.5093
2. ImageNet Classification with Deep Convolutional Neural Networks — Alex Krizhevsky, Ilya Sutskever, Geoffrey E. Hinton, 2012
https://scholar.google.com/scholar?q=ImageNet+Classification+with+Deep+Convolutional+Neural+Networks
3. OverFeat: Integrated Recognition, Localization and Detection using Convolutional Networks — Pierre Sermanet, David Eigen, Xiang Zhang, Michael Mathieu, Rob Fergus, Yann LeCun, 2013
https://scholar.google.com/scholar?q=OverFeat%3A+Integrated+Recognition%2C+Localization+and+Detection+using+Convolutional+Networks
4. Decaf: A Deep Convolutional Activation Feature for Generic Visual Recognition — Jeff Donahue, Yangqing Jia, Oriol Vinyals, Judy Hoffman, Ning Zhang, Eric Tzeng, Trevor Darrell, 2014
https://scholar.google.com/scholar?q=Decaf%3A+A+Deep+Convolutional+Activation+Feature+for+Generic+Visual+Recognition
5. cuda-convnet — Alex Krizhevsky, 2012
https://scholar.google.com/scholar?q=cuda-convnet
6. Theano: new features and speed improvements — Frédéric Bastien, Pascal Lamblin, Razvan Pascanu, James Bergstra, Ian Goodfellow, Arnaud Bergeron, Nicolas Bouchard, David Warde-Farley, Yoshua Bengio and others, 2012
https://scholar.google.com/scholar?q=Theano%3A+new+features+and+speed+improvements
7. Pylearn2: a machine learning research library — Ian Goodfellow, David Warde-Farley, Mehdi Mirza, Aaron Courville, Yoshua Bengio, 2013
https://scholar.google.com/scholar?q=Pylearn2%3A+a+machine+learning+research+library
8. Torch7 — Ronan Collobert, Koray Kavukcuoglu, Clément Farabet and others, 2011
https://scholar.google.com/scholar?q=Torch7
9. Efficient inference of Vision Transformer with structural pruning and operator fusion on GPU — unknown from snippet, recent
https://scholar.google.com/scholar?q=Efficient+inference+of+Vision+Transformer+with+structural+pruning+and+operator+fusion+on+GPU
10. I-ViT: Integer-only quantization for efficient vision transformer inference — unknown from snippet, recent
https://scholar.google.com/scholar?q=I-ViT%3A+Integer-only+quantization+for+efficient+vision+transformer+inference
11. DeViT: Decomposing vision transformers for collaborative inference in edge devices — unknown from snippet, recent
https://scholar.google.com/scholar?q=DeViT%3A+Decomposing+vision+transformers+for+collaborative+inference+in+edge+devices
12. Raman: A reconfigurable and sparse TinyML accelerator for inference on edge — unknown from snippet, recent
https://scholar.google.com/scholar?q=Raman%3A+A+reconfigurable+and+sparse+TinyML+accelerator+for+inference+on+edge
13. Hardware accelerator design for sparse DNN inference and training: A tutorial — unknown from snippet, recent
https://scholar.google.com/scholar?q=Hardware+accelerator+design+for+sparse+DNN+inference+and+training%3A+A+tutorial
14. Inference serving with end-to-end latency SLOs over dynamic edge networks — unknown from snippet, recent
https://scholar.google.com/scholar?q=Inference+serving+with+end-to-end+latency+SLOs+over+dynamic+edge+networks
15. Training data attribution via approximate unrolling — unknown from snippet, recent
https://scholar.google.com/scholar?q=Training+data+attribution+via+approximate+unrolling
16. Exploring Training Data Attribution under Limited Access Constraints — unknown from snippet, recent
https://scholar.google.com/scholar?q=Exploring+Training+Data+Attribution+under+Limited+Access+Constraints
17. DATE-LM: Benchmarking Data Attribution Evaluation for Large Language Models — unknown from snippet, recent
https://scholar.google.com/scholar?q=DATE-LM%3A+Benchmarking+Data+Attribution+Evaluation+for+Large+Language+Models
18. AI Post Transformers: Speculative Decoding in Real vLLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-speculative-decoding-in-real-vllm-servin-6f4e2b.mp3
19. AI Post Transformers: GPT-NeoX: Large-Scale Autoregressive Language Modeling in PyTorch — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/gpt-neox-large-scale-autoregressive-language-modeling-in-pytorch/
20. AI Post Transformers: NVMe Offload on Colossal AI: Breaking the GPU Memory Wall — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/nvme-offload-on-colossal-ai-breaking-the-gpu-memory-wall/
21. AI Post Transformers: Mamba-3 for Efficient Sequence Modeling — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-16-mamba-3-for-efficient-sequence-modeling-97a22a.mp3
Interactive Visualization: Caffe and the Rise of CNN Frameworks

This episode explores the 2016 Caffeine FPGA accelerator and its central claim that a single FPGA design can handle an entire CNN efficiently, rather than excelling at convolutions while bottlenecking on fully connected layers. It explains why that mattered in the AlexNet-to-VGG era, when convolutional layers were compute-bound but dense layers often became communication-bound because moving weights and activations through memory was the real constraint. The discussion focuses on Caffeine’s main technical idea: a unified matrix-multiplication-oriented representation that supports both convolution and fully connected layers without the heavy data expansion of standard `im2col` approaches, plus memory-access scheduling choices such as weight-major mapping to improve reuse and burst efficiency. Listeners would find it interesting because the episode makes a precise systems argument about how hardware performance depends not just on arithmetic throughput, but on matching dataflow, buffering, and bandwidth to the structure of the network.

Interactive Visualization: Caffeine: A Unified FPGA for CNNs
Sources:
1. Caffeine: A Unified FPGA for CNNs
https://ceca.pku.edu.cn/media/lw/83b308c75c56a94fbf706b92dbe57917.pdf
2. Gradient-Based Learning Applied to Document Recognition — Yann LeCun, Léon Bottou, Yoshua Bengio, Patrick Haffner, 1998
https://scholar.google.com/scholar?q=Gradient-Based+Learning+Applied+to+Document+Recognition
3. ImageNet Classification with Deep Convolutional Neural Networks — Alex Krizhevsky, Ilya Sutskever, Geoffrey E. Hinton, 2012
https://scholar.google.com/scholar?q=ImageNet+Classification+with+Deep+Convolutional+Neural+Networks
4. Very Deep Convolutional Networks for Large-Scale Image Recognition — Karen Simonyan, Andrew Zisserman, 2014
https://scholar.google.com/scholar?q=Very+Deep+Convolutional+Networks+for+Large-Scale+Image+Recognition
5. Deep Residual Learning for Image Recognition — Kaiming He, Xiangyu Zhang, Shaoqing Ren, Jian Sun, 2016
https://scholar.google.com/scholar?q=Deep+Residual+Learning+for+Image+Recognition
6. DianNao: A Small-Footprint High-Throughput Accelerator for Ubiquitous Machine-Learning — Tianshi Chen, Zidong Du, Ninghui Sun, Jia Wang, Chengyong Wu, Yunji Chen, Olivier Temam, 2014
https://scholar.google.com/scholar?q=DianNao%3A+A+Small-Footprint+High-Throughput+Accelerator+for+Ubiquitous+Machine-Learning
7. Eyeriss: An Energy-Efficient Reconfigurable Accelerator for Deep Convolutional Neural Networks — Yu-Hsin Chen, Tushar Krishna, Joel S. Emer, Vivienne Sze, 2016
https://scholar.google.com/scholar?q=Eyeriss%3A+An+Energy-Efficient+Reconfigurable+Accelerator+for+Deep+Convolutional+Neural+Networks
8. Caffeine: Towards Uniformed Representation and Acceleration for Deep Convolutional Neural Networks — Chen Zhang, Zhenman Fang, Peipei Zhou, Peichen Pan, Jason Cong, 2016
https://scholar.google.com/scholar?q=Caffeine%3A+Towards+Uniformed+Representation+and+Acceleration+for+Deep+Convolutional+Neural+Networks
9. In-Datacenter Performance Analysis of a Tensor Processing Unit — Norman P. Jouppi and many colleagues at Google, 2017
https://scholar.google.com/scholar?q=In-Datacenter+Performance+Analysis+of+a+Tensor+Processing+Unit
10. Optimizing FPGA-Based Accelerator Design for Deep Convolutional Neural Networks — Chen Zhang, Peng Li, Guangyu Sun, Yijin Guan, Bingjun Xiao, Jason Cong, 2015
https://scholar.google.com/scholar?q=Optimizing+FPGA-Based+Accelerator+Design+for+Deep+Convolutional+Neural+Networks
11. Going Deeper with Embedded FPGA Platform for Convolutional Neural Network — Jiantao Qiu, Jingsheng Wang, Song Yao, Kaiyuan Guo, Boxun Li, Erjin Zhou, Jincheng Yu, Tianqi Tang, Ningyi Xu, Sen Song, Yu Wang, Huazhong Yang, 2016
https://scholar.google.com/scholar?q=Going+Deeper+with+Embedded+FPGA+Platform+for+Convolutional+Neural+Network
12. fpgaConvNet: A Framework for Mapping Convolutional Neural Networks on FPGAs — Stylianos I. Venieris, Christos-Savvas Bouganis, 2016
https://scholar.google.com/scholar?q=fpgaConvNet%3A+A+Framework+for+Mapping+Convolutional+Neural+Networks+on+FPGAs
13. A high-performance FPGA-based depthwise separable convolution accelerator — approximate; recent FPGA accelerator authors, recent
https://scholar.google.com/scholar?q=A+high-performance+FPGA-based+depthwise+separable+convolution+accelerator
14. Fpga-based acceleration for convolutional neural networks: A comprehensive review — approximate; review authors, recent
https://scholar.google.com/scholar?q=Fpga-based+acceleration+for+convolutional+neural+networks%3A+A+comprehensive+review
15. Mobile-X: Dedicated FPGA implementation of the MobileNet accelerator optimizing depthwise separable convolution — approximate; Mobile-X authors, recent
https://scholar.google.com/scholar?q=Mobile-X%3A+Dedicated+FPGA+implementation+of+the+MobileNet+accelerator+optimizing+depthwise+separable+convolution
16. Design of a convolutional neural network accelerator based on on-chip data reordering — approximate; accelerator authors, recent
https://scholar.google.com/scholar?q=Design+of+a+convolutional+neural+network+accelerator+based+on+on-chip+data+reordering
17. Energy-efficient and high-throughput CNN inference engine based on memory-sharing and data-reusing for edge applications — approximate; edge-CNN accelerator authors, recent
https://scholar.google.com/scholar?q=Energy-efficient+and+high-throughput+CNN+inference+engine+based+on+memory-sharing+and+data-reusing+for+edge+applications
18. An efficient sparse CNN inference accelerator with balanced intra-and inter-PE workload — approximate; sparse accelerator authors, recent
https://scholar.google.com/scholar?q=An+efficient+sparse+CNN+inference+accelerator+with+balanced+intra-and+inter-PE+workload
19. Hardware accelerator design for sparse DNN inference and training: A tutorial — approximate; tutorial authors, recent
https://scholar.google.com/scholar?q=Hardware+accelerator+design+for+sparse+DNN+inference+and+training%3A+A+tutorial
20. AI Post Transformers: FlatAttention for Tile-Based Accelerator Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-flatattention-for-tile-based-accelerator-56e6ca.mp3
21. AI Post Transformers: RFNoC SISO Processor via High-Level Synthesis — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-06-rfnoc-siso-processor-via-high-level-synt-c892f3.mp3
22. AI Post Transformers: Computation-Bandwidth-Memory Trade-offs for AI Infrastructure — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-09-computation-bandwidth-memory-trade-offs-a83f2b.mp3
23. AI Post Transformers: Los Alamos: overcoming the memory wall fighting sparse memory access — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/los-alamos-overcoming-the-memory-wall-fighting-sparse-memory-access/
Interactive Visualization: Caffeine: A Unified FPGA for CNNs

This episode explores how the CMS experiment uses machine learning inside its Level-1 endcap muon trigger, where hardware must estimate muon momentum within roughly 500 nanoseconds while filtering an enormous stream of proton-collision data. It explains why boosted decision trees were chosen over neural networks: not because they are trendier, but because they fit strict FPGA constraints around deterministic latency, fixed-point arithmetic, and bounded memory. A central finding is that the online system does not run the trees directly; instead, the model is trained offline and compiled into a massive precomputed lookup table, turning inference into a single fast memory access. The discussion is especially interesting because it shows machine learning as a systems-and-hardware co-design problem, grounded in detector physics, feature engineering, and the practical realities of deploying learned functions in one of the harshest real-time environments in science.

Sources:
1. Boosted Decision Trees for CMS Muon Triggers
https://indico.cern.ch/event/567550/papers/2629686/files/6172-acat_bdt_l1t.pdf
2. Applications and Techniques for Fast Machine Learning in Science — Allison McCarn Deiana, Nhan Tran, Joshua Agar, Michaela Blott, Giuseppe Di Guglielmo, Javier Duarte, Philip Harris, Mia Liu, Mark Neubauer, Jennifer Ngadiuba, Maurizio Pierini and many others, 2022
https://scholar.google.com/scholar?q=Applications+and+Techniques+for+Fast+Machine+Learning+in+Science
3. Fast inference of Boosted Decision Trees in FPGAs for particle physics — Sioni Summers, Giuseppe Di Guglielmo, Javier Duarte, Philip Harris, Duc Hoang, Sergo Jindariani, Edward Kreinar, Vladimir Loncar, Jennifer Ngadiuba, Maurizio Pierini, Dylan Rankin, Nhan Tran and Zhenbin Wu, 2020
https://scholar.google.com/scholar?q=Fast+inference+of+Boosted+Decision+Trees+in+FPGAs+for+particle+physics
4. Automatic heterogeneous quantization of deep neural networks for low-latency inference on the edge for particle detectors — Claudionor N. Coelho Jr, Aki Kuusela, Shan Li, Hao Zhuang, Jennifer Ngadiuba, Thea Klaeboe Aarrestad, Vladimir Loncar, Maurizio Pierini, Adrian Alan Pol and Sioni Summers, 2021
https://scholar.google.com/scholar?q=Automatic+heterogeneous+quantization+of+deep+neural+networks+for+low-latency+inference+on+the+edge+for+particle+detectors
5. Serving DNNs in Real Time at Datacenter Scale with Project Brainwave — Eric Chung, Jeremy Fowers, Kalin Ovtcharov, Michael Papamichael, Adrian Caulfield, Todd Massengill, Ming Liu, Mahdi Ghandi, Daniel Lo and others, 2018
https://scholar.google.com/scholar?q=Serving+DNNs+in+Real+Time+at+Datacenter+Scale+with+Project+Brainwave
6. The CMS Trigger System — CMS Collaboration, not specified in excerpt
https://scholar.google.com/scholar?q=The+CMS+Trigger+System
7. The CMS Endcap Muon Track Finder — CMS Collaboration or EMTF-related authors, not specified in excerpt
https://scholar.google.com/scholar?q=The+CMS+Endcap+Muon+Track+Finder
8. TMVA: Toolkit for Multivariate Data Analysis — Andreas Hoecker and collaborators, not specified in excerpt
https://scholar.google.com/scholar?q=TMVA%3A+Toolkit+for+Multivariate+Data+Analysis
9. Fast Machine Learning for Science: how accelerated hardware and software are enabling real-time data analysis at the edge — Javier Duarte and collaborators, 2022
https://scholar.google.com/scholar?q=Fast+Machine+Learning+for+Science%3A+how+accelerated+hardware+and+software+are+enabling+real-time+data+analysis+at+the+edge
10. hls4ml: An Open-Source Codesign Workflow to Empower Scientific Low-Power Machine Learning Devices — Giuseppe Di Guglielmo, Javier Duarte and collaborators, 2021
https://scholar.google.com/scholar?q=hls4ml%3A+An+Open-Source+Codesign+Workflow+to+Empower+Scientific+Low-Power+Machine+Learning+Devices
11. End-to-end codesign of Hessian-aware quantized neural networks for FPGAs and ASICs — Javier Campos, Zhen Dong, Javier Duarte, Nhan Tran, et al., 2023
https://scholar.google.com/scholar?q=End-to-end+codesign+of+Hessian-aware+quantized+neural+networks+for+FPGAs+and+ASICs
12. FPGA-QNN: Quantized Neural Network Hardware Acceleration on FPGAs — Mustafa Tasci, Ayhan Istanbullu, Vedat Tumen, Selahattin Kosunalp, 2025
https://scholar.google.com/scholar?q=FPGA-QNN%3A+Quantized+Neural+Network+Hardware+Acceleration+on+FPGAs
13. An FPGA-Based Time-to-Digital Converter with Online Dual-Chain Calibration — Zhengsen Jia, Yuzhuo Wang, Jie Ding, Qian Xu, et al., 2025
https://scholar.google.com/scholar?q=An+FPGA-Based+Time-to-Digital+Converter+with+Online+Dual-Chain+Calibration
14. A Novel FPGA-based Time-to-Digital Converter featuring Machine Learning-Aided Self-Calibration — Arash Amini Bardpareh, Eleonora Vacca, Davide Nicolini, Luca Sterpone, et al., 2026
https://scholar.google.com/scholar?q=A+Novel+FPGA-based+Time-to-Digital+Converter+featuring+Machine+Learning-Aided+Self-Calibration
15. AI Post Transformers: FlatAttention for Tile-Based Accelerator Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-flatattention-for-tile-based-accelerator-56e6ca.mp3
16. AI Post Transformers: SolidAttention: Co-Designing Sparse Attention and SSD I/O — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-18-solidattention-co-designing-sparse-atten-5a8622.mp3
Interactive Visualization: Boosted Decision Trees for CMS Muon Triggers

This episode explores how a 2018 paper brings neural network inference into the Level-1 trigger at the Large Hadron Collider, where event decisions must be made under sub-microsecond latency constraints. It explains why FPGAs are a natural fit for this setting, emphasizing batch-one, deterministic inference and the hardware realities that make model size, timing, memory use, and routing just as important as accuracy. The discussion centers on a compact dense network for jet substructure classification, using 16 engineered features to distinguish quark, gluon, W, Z, and top jets while preserving rare physics signals. It also highlights the paper’s broader argument: tools like High-Level Synthesis and hls4ml can let physicists deploy hardware-aware ML workflows directly, making real-time AI a practical part of scientific instrumentation rather than just a benchmark exercise.

Interactive Visualization: Fast FPGA Inference for LHC Triggers
Sources:
1. Fast inference of deep neural networks in FPGAs for particle physics — Javier Duarte, Song Han, Philip Harris, Sergo Jindariani, Edward Kreinar, Benjamin Kreis, Jennifer Ngadiuba, Maurizio Pierini, Ryan Rivera, Nhan Tran, Zhenbin Wu, 2018
http://arxiv.org/abs/1804.06913
2. A Survey on Performance Optimization of High-Level Synthesis Tools — Lan Huang, Da-Lin Li, Kang-Ping Wang, Teng Gao, Adriano Tavares, 2020
https://scholar.google.com/scholar?q=A+Survey+on+Performance+Optimization+of+High-Level+Synthesis+Tools
3. FINN-R: An End-to-End Deep-Learning Framework for Fast Exploration of Quantized Neural Networks — Michaela Blott, Thomas B. Preusser, Nicholas J. Fraser, Giulio Gambardella, Kenneth O'Brien, Yaman Umuroglu, Miriam Leeser, Kees Vissers, 2018
https://scholar.google.com/scholar?q=FINN-R%3A+An+End-to-End+Deep-Learning+Framework+for+Fast+Exploration+of+Quantized+Neural+Networks
4. Fast inference of deep neural networks in FPGAs for particle physics — Javier Duarte, Song Han, Philip Harris, Sergo Jindariani, Edward Kreinar, Jennifer Ngadiuba, Maurizio Pierini, Nhan Tran, Zhenbin Wu, et al., 2018
https://scholar.google.com/scholar?q=Fast+inference+of+deep+neural+networks+in+FPGAs+for+particle+physics
5. Fast convolutional neural networks on FPGAs with hls4ml — Thea Aarrestad, Vladimir Loncar, Nicolo Ghielmetti, Maurizio Pierini, Sioni Summers, Jennifer Ngadiuba, Javier Duarte, Philip Harris, et al., 2021
https://scholar.google.com/scholar?q=Fast+convolutional+neural+networks+on+FPGAs+with+hls4ml
6. FINN: A Framework for Fast, Scalable Binarized Neural Network Inference — Yaman Umuroglu, Nicholas J. Fraser, Giulio Gambardella, Michaela Blott, Philip H. W. Leong, Magnus Jahre, Kees Vissers, 2017
https://scholar.google.com/scholar?q=FINN%3A+A+Framework+for+Fast%2C+Scalable+Binarized+Neural+Network+Inference
7. Serving DNNs in Real Time at Datacenter Scale with Project Brainwave — Eric Chung, Jeremy Fowers, Kalin Ovtcharov, Michael Papamichael, Adrian Caulfield, Todd Massengill, Ming Liu, Doug Burger, et al., 2018
https://scholar.google.com/scholar?q=Serving+DNNs+in+Real+Time+at+Datacenter+Scale+with+Project+Brainwave
8. Automatic heterogeneous quantization of deep neural networks for low-latency inference on the edge for particle detectors — Claudionor N. Coelho Jr., Aki Kuusela, Shan Li, Hao Zhuang, Jennifer Ngadiuba, Thea K. Aarrestad, Vladimir Loncar, Maurizio Pierini, et al., 2021
https://scholar.google.com/scholar?q=Automatic+heterogeneous+quantization+of+deep+neural+networks+for+low-latency+inference+on+the+edge+for+particle+detectors
9. Learning both Weights and Connections for Efficient Neural Network — Song Han, Jeff Pool, John Tran, William J. Dally, 2015
https://scholar.google.com/scholar?q=Learning+both+Weights+and+Connections+for+Efficient+Neural+Network
10. Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding — Song Han, Huizi Mao, William J. Dally, 2016
https://scholar.google.com/scholar?q=Deep+Compression%3A+Compressing+Deep+Neural+Networks+with+Pruning%2C+Trained+Quantization+and+Huffman+Coding
11. Jet Substructure at the Large Hadron Collider: A Review of Recent Advances in Theory and Machine Learning — Andrew J. Larkoski, Ian Moult, Benjamin Nachman, 2020
https://scholar.google.com/scholar?q=Jet+Substructure+at+the+Large+Hadron+Collider%3A+A+Review+of+Recent+Advances+in+Theory+and+Machine+Learning
12. Deep-learning Top Taggers or The End of QCD? — Gregor Kasieczka, Tilman Plehn, Michael Russell, Torben Schell, 2017
https://scholar.google.com/scholar?q=Deep-learning+Top+Taggers+or+The+End+of+QCD%3F
13. From High-Level Deep Neural Models to FPGAs — Hardik Sharma, Jongse Park, Divya Mahajan, Emmanuel Amaro, Joon Kyung Kim, Chenkai Shao, Asit Mishra, Hadi Esmaeilzadeh, 2016
https://scholar.google.com/scholar?q=From+High-Level+Deep+Neural+Models+to+FPGAs
14. Distance-Weighted Graph Neural Networks on FPGAs for Real-Time Particle Reconstruction in High Energy Physics — Yutaro Iiyama, Gianluca Cerminara, Abhijay Gupta, Jan Kieseler, Vladimir Loncar, Maurizio Pierini, Shah Rukh Qasim and collaborators, 2021
https://scholar.google.com/scholar?q=Distance-Weighted+Graph+Neural+Networks+on+FPGAs+for+Real-Time+Particle+Reconstruction+in+High+Energy+Physics
15. Low latency transformer inference on FPGAs for physics applications with hls4ml — not confirmed from snippet; likely hls4ml/particle-physics collaboration, recent
https://scholar.google.com/scholar?q=Low+latency+transformer+inference+on+FPGAs+for+physics+applications+with+hls4ml
16. Optimizing transformer models for low-latency inference: techniques, architectures, and code implementations — not confirmed from snippet, recent
https://scholar.google.com/scholar?q=Optimizing+transformer+models+for+low-latency+inference%3A+techniques%2C+architectures%2C+and+code+implementations
17. Low-bit mixed-precision quantization and acceleration of CNN for FPGA deployment — not confirmed from snippet, recent
https://scholar.google.com/scholar?q=Low-bit+mixed-precision+quantization+and+acceleration+of+CNN+for+FPGA+deployment
18. MPQA: Mixed-Precision Quantization Accelerator for CNN Inference — not confirmed from snippet, recent
https://scholar.google.com/scholar?q=MPQA%3A+Mixed-Precision+Quantization+Accelerator+for+CNN+Inference
19. Fine-grained structured sparse computing for FPGA-based AI inference — not confirmed from snippet, recent
https://scholar.google.com/scholar?q=Fine-grained+structured+sparse+computing+for+FPGA-based+AI+inference
20. Efficient CNN inference acceleration on FPGAs: a pattern pruning-driven approach — not confirmed from snippet, recent
https://scholar.google.com/scholar?q=Efficient+CNN+inference+acceleration+on+FPGAs%3A+a+pattern+pruning-driven+approach
21. Online Learning Extreme Learning Machine with Low-Complexity Predictive Plasticity Rule and FPGA Implementation — not confirmed from snippet, recent
https://scholar.google.com/scholar?q=Online+Learning+Extreme+Learning+Machine+with+Low-Complexity+Predictive+Plasticity+Rule+and+FPGA+Implementation
22. An FPGA architecture for online learning using the Tsetlin machine — not confirmed from snippet, recent
https://scholar.google.com/scholar?q=An+FPGA+architecture+for+online+learning+using+the+Tsetlin+machine
23. AI Post Transformers: FPGA Neural Network Accelerators for Space — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-26-fpga-neural-network-accelerators-for-spa-3087ae.mp3
Interactive Visualization: Fast FPGA Inference for LHC Triggers

This episode explores how boosted decision trees can be compiled directly into FPGA firmware for ultra-low-latency particle-physics triggers at the Large Hadron Collider. It explains why this setting favors shallow, quantized tree ensembles over larger neural networks: trigger decisions must happen within a tiny hardware budget, with strict limits on latency, power, and on-chip resources. The discussion focuses on a concrete benchmark where a 100-tree, depth-4 gradient-boosted model for five-class jet tagging is mapped to a Xilinx VU9P FPGA and compared against a similarly deployed multilayer perceptron. Listeners would find it interesting because it shows how model choice changes when every nanosecond matters, and how familiar ML methods can become hardwired decision circuits rather than conventional software inference.

Interactive Visualization: Fast FPGA BDT Inference for LHC Triggers
Sources:
1. Fast inference of Boosted Decision Trees in FPGAs for particle physics — Sioni Summers, Giuseppe Di Guglielmo, Javier Duarte, Philip Harris, Duc Hoang, Sergo Jindariani, Edward Kreinar, Vladimir Loncar, Jennifer Ngadiuba, Maurizio Pierini, Dylan Rankin, Nhan Tran, Zhenbin Wu, 2020
http://arxiv.org/abs/2002.02534
2. Greedy Function Approximation: A Gradient Boosting Machine — Jerome H. Friedman, 2001
https://scholar.google.com/scholar?q=Greedy+Function+Approximation%3A+A+Gradient+Boosting+Machine
3. XGBoost: A Scalable Tree Boosting System — Tianqi Chen, Carlos Guestrin, 2016
https://scholar.google.com/scholar?q=XGBoost%3A+A+Scalable+Tree+Boosting+System
4. LightGBM: A Highly Efficient Gradient Boosting Decision Tree — Guolin Ke, Qi Meng, Thomas Finley, Taifeng Wang, Wei Chen, Weidong Ma, Qiwei Ye, Tie-Yan Liu, 2017
https://scholar.google.com/scholar?q=LightGBM%3A+A+Highly+Efficient+Gradient+Boosting+Decision+Tree
5. CatBoost: unbiased boosting with categorical features — Liudmila Prokhorenkova, Gleb Gusev, Aleksandr Vorobev, Anna Veronika Dorogush, Andrey Gulin, 2018
https://scholar.google.com/scholar?q=CatBoost%3A+unbiased+boosting+with+categorical+features
6. Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference — Benoit Jacob, Skirmantas Kligys, Bo Chen, Menglong Zhu, Matthew Tang, Andrew Howard, Hartwig Adam, Dmitry Kalenichenko, 2018
https://scholar.google.com/scholar?q=Quantization+and+Training+of+Neural+Networks+for+Efficient+Integer-Arithmetic-Only+Inference
7. Quantizing deep convolutional networks for efficient inference: A whitepaper — Raghuraman Krishnamoorthi, 2018
https://scholar.google.com/scholar?q=Quantizing+deep+convolutional+networks+for+efficient+inference%3A+A+whitepaper
8. Post-training 4-bit quantization of convolution networks for rapid-deployment — Ron Banner, Yury Nahshan, Elad Hoffer, Daniel Soudry, 2019
https://scholar.google.com/scholar?q=Post-training+4-bit+quantization+of+convolution+networks+for+rapid-deployment
9. GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers — Elias Frantar, Saleh Ashkboos, Torsten Hoefler, Dan Alistarh, 2022
https://scholar.google.com/scholar?q=GPTQ%3A+Accurate+Post-Training+Quantization+for+Generative+Pre-trained+Transformers
10. AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration — Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Wei-Ming Chen, Wei-Chen Wang, Guangxuan Xiao, Xingyu Dang, Chuang Gan, Song Han, 2023
https://scholar.google.com/scholar?q=AWQ%3A+Activation-aware+Weight+Quantization+for+LLM+Compression+and+Acceleration
11. Fast inference of deep neural networks in FPGAs for particle physics — Javier Duarte, Song Han, Philip Harris, Sergo Jindariani, Edward Kreinar, Benjamin Kreis, Jennifer Ngadiuba, Maurizio Pierini, Ryan Rivera, Nhan Tran, Zhenbin Wu, 2018
https://scholar.google.com/scholar?q=Fast+inference+of+deep+neural+networks+in+FPGAs+for+particle+physics
12. Efficient, reliable and fast high-level triggering using a bonsai boosted decision tree — V. V. Gligorov, M. Williams, 2013
https://scholar.google.com/scholar?q=Efficient%2C+reliable+and+fast+high-level+triggering+using+a+bonsai+boosted+decision+tree
13. Boosted Decision Trees in the Level-1 Muon Endcap Trigger at CMS — CMS Collaboration, 2018
https://scholar.google.com/scholar?q=Boosted+Decision+Trees+in+the+Level-1+Muon+Endcap+Trigger+at+CMS
14. Scalable inference of decision tree ensembles: Flexible design for CPU-FPGA platforms — Muhsen Owaida, Hantian Zhang, Ce Zhang, Gustavo Alonso, 2017
https://scholar.google.com/scholar?q=Scalable+inference+of+decision+tree+ensembles%3A+Flexible+design+for+CPU-FPGA+platforms
15. Machine learning at the energy and intensity frontiers of particle physics — A. Radovic et al., 2018
https://scholar.google.com/scholar?q=Machine+learning+at+the+energy+and+intensity+frontiers+of+particle+physics
16. Low latency transformer inference on FPGAs for physics applications with hls4ml — Zhixing Jiang et al., 2025
https://scholar.google.com/scholar?q=Low+latency+transformer+inference+on+FPGAs+for+physics+applications+with+hls4ml
17. Ultrafast jet classification at the HL-LHC — Patrick Odagiu et al., 2024
https://scholar.google.com/scholar?q=Ultrafast+jet+classification+at+the+HL-LHC
18. AI Post Transformers: FlatAttention for Tile-Based Accelerator Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-flatattention-for-tile-based-accelerator-56e6ca.mp3
19. AI Post Transformers: Computation-Bandwidth-Memory Trade-offs for AI Infrastructure — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-09-computation-bandwidth-memory-trade-offs-a83f2b.mp3
20. AI Post Transformers: TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-25-turboquant-online-vector-quantiz-1967b7.mp3
Interactive Visualization: Fast FPGA BDT Inference for LHC Triggers

This episode explores why LightGBM became a dominant tool for tabular machine learning by unpacking the algorithmic and systems ideas behind its speed. It explains how gradient boosting decision trees work, why split search becomes expensive on massive sparse datasets, and how LightGBM differs from neural-network-style training despite using gradient information. The discussion focuses on two core contributions: Gradient-based One-Side Sampling, which keeps high-gradient examples while subsampling easier ones without badly distorting split-gain estimates, and Exclusive Feature Bundling, which compresses sparse features by grouping columns that rarely activate together. Listeners would find it interesting for its clear account of how classical ideas like histograms, greedy tree growth, and graph coloring were combined into a highly practical system that reshaped real-world applications such as ranking, fraud detection, credit scoring, and forecasting.

Interactive Visualization: Why LightGBM Made Boosted Trees Fast
Sources:
1. Why LightGBM Made Boosted Trees Fast
https://proceedings.neurips.cc/paper_files/paper/2017/file/6449f44a102fde848669bdd9eb6b76fa-Paper.pdf
2. Greedy Function Approximation: A Gradient Boosting Machine — Jerome H. Friedman, 2001
https://scholar.google.com/scholar?q=Greedy+Function+Approximation%3A+A+Gradient+Boosting+Machine
3. XGBoost: A Scalable Tree Boosting System — Tianqi Chen and Carlos Guestrin, 2016
https://scholar.google.com/scholar?q=XGBoost%3A+A+Scalable+Tree+Boosting+System
4. LightGBM: A Highly Efficient Gradient Boosting Decision Tree — Guolin Ke, Qi Meng, Thomas Finley, Taifeng Wang, Wei Chen, Weidong Ma, Qiwei Ye, Tie-Yan Liu, 2017
https://scholar.google.com/scholar?q=LightGBM%3A+A+Highly+Efficient+Gradient+Boosting+Decision+Tree
5. CatBoost: Unbiased Boosting with Categorical Features — Liudmila Prokhorenkova, Gleb Gusev, Aleksandr Vorobev, Anna Veronika Dorogush, Andrey Gulin, 2018
https://scholar.google.com/scholar?q=CatBoost%3A+Unbiased+Boosting+with+Categorical+Features
6. Feature Hashing for Large Scale Multitask Learning — Kilian Weinberger, Anirban Dasgupta, Josh Attenberg, John Langford, Alex Smola, 2009
https://scholar.google.com/scholar?q=Feature+Hashing+for+Large+Scale+Multitask+Learning
7. An Upper Bound for the Chromatic Number of a Graph and Its Application to Timetabling Problems — D. J. A. Welsh and M. B. Powell, 1967
https://scholar.google.com/scholar?q=An+Upper+Bound+for+the+Chromatic+Number+of+a+Graph+and+Its+Application+to+Timetabling+Problems
8. New Methods to Color the Vertices of a Graph — Daniel Brélaz, 1979
https://scholar.google.com/scholar?q=New+Methods+to+Color+the+Vertices+of+a+Graph
9. Worst Case Behavior of Graph Coloring Algorithms — David S. Johnson, 1974
https://scholar.google.com/scholar?q=Worst+Case+Behavior+of+Graph+Coloring+Algorithms
10. A Communication-Efficient Parallel Algorithm for Decision Tree — Qi Meng, Guolin Ke, Taifeng Wang, Wei Chen, Qiwei Ye, Zhi-Ming Ma, Tie-Yan Liu, 2016
https://scholar.google.com/scholar?q=A+Communication-Efficient+Parallel+Algorithm+for+Decision+Tree
11. Stochastic Gradient Boosting — Jerome H. Friedman, 2002
https://scholar.google.com/scholar?q=Stochastic+Gradient+Boosting
12. Parallel Boosted Regression Trees for Web Search Ranking — Stephen Tyree, Kilian Q. Weinberger, Kunal Agrawal, and Jennifer Paykin, 2011
https://scholar.google.com/scholar?q=Parallel+Boosted+Regression+Trees+for+Web+Search+Ranking
13. Best-First Decision Tree Learning — Haijian Shi, 2007
https://scholar.google.com/scholar?q=Best-First+Decision+Tree+Learning
14. GPU-Acceleration for Large-Scale Tree Boosting — Huan Zhang, Si Si, and Cho-Jui Hsieh, 2017
https://scholar.google.com/scholar?q=GPU-Acceleration+for+Large-Scale+Tree+Boosting
15. Implementing machine learning methods with complex survey data: Lessons learned on the impacts of accounting sampling weights in gradient boosting — authors not identified in the provided snippet, recent (2020s)
https://scholar.google.com/scholar?q=Implementing+machine+learning+methods+with+complex+survey+data%3A+Lessons+learned+on+the+impacts+of+accounting+sampling+weights+in+gradient+boosting
16. Explainable boosting algorithms: sparse-group and interaction-aware variable selection in complex data — authors not identified in the provided snippet, recent (2020s)
https://scholar.google.com/scholar?q=Explainable+boosting+algorithms%3A+sparse-group+and+interaction-aware+variable+selection+in+complex+data
17. Multi-objective optimization of performance and interpretability of tabular supervised machine learning models — authors not identified in the provided snippet, recent (2020s)
https://scholar.google.com/scholar?q=Multi-objective+optimization+of+performance+and+interpretability+of+tabular+supervised+machine+learning+models
18. AI Post Transformers: Breiman's Two Cultures of Statistical Modeling — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-24-breimans-two-cultures-of-statistical-mod-71e49f.mp3
Interactive Visualization: Why LightGBM Made Boosted Trees Fast

This episode explores how fpgaConvNet turns CNN inference into a Synchronous Dataflow problem so embedded FPGA accelerators can be designed with analyzable schedules, buffers, and resource tradeoffs instead of ad hoc hardware tuning. It explains why CNN deployment on robots, drones, and cars is constrained as much by data movement, latency, and power as by raw arithmetic, and why FPGAs can outperform embedded GPUs when the hardware is tailored carefully to a model’s structure. The discussion highlights the paper’s central claim that formalizing the mapping problem enables automated design-space exploration across very different CNN topologies, rather than just optimizing a single benchmark. It also examines where that approach is strong and where listeners should be skeptical, including whether the reported GPU speedups are fair and how well the clean SDF abstraction survives real hardware implementation.

Interactive Visualization: Automating CNN Mapping on Embedded FPGAs
Sources:
1. fpgaConvNet: A Toolflow for Mapping Diverse Convolutional Neural Networks on Embedded FPGAs — Stylianos I. Venieris, Christos-Savvas Bouganis, 2017
http://arxiv.org/abs/1711.08740
2. Static Scheduling of Synchronous Data Flow Programs for Digital Signal Processing — Edward A. Lee, David G. Messerschmitt, 1987
https://scholar.google.com/scholar?q=Static+Scheduling+of+Synchronous+Data+Flow+Programs+for+Digital+Signal+Processing
3. Synchronous Data Flow — Edward A. Lee, David G. Messerschmitt, 1987
https://scholar.google.com/scholar?q=Synchronous+Data+Flow
4. Scenario-aware dataflow: modeling, analysis and implementation of dynamic applications — Sander Stuijk, Marc C. W. Geilen, Bart D. Theelen, Twan Basten, 2011
https://scholar.google.com/scholar?q=Scenario-aware+dataflow%3A+modeling%2C+analysis+and+implementation+of+dynamic+applications
5. fpgaConvNet: A Toolflow for Mapping Diverse Convolutional Neural Networks on Embedded FPGAs — Stylianos I. Venieris, Christos-Savvas Bouganis, 2017
https://scholar.google.com/scholar?q=fpgaConvNet%3A+A+Toolflow+for+Mapping+Diverse+Convolutional+Neural+Networks+on+Embedded+FPGAs
6. DNNWeaver: From High-Level Deep Network Models to FPGA Acceleration — Hyoukjun Sharma, Jongse Park, Emmanuel Amaro, Bradley Thwaites, Praneeth Kotha, Anmol Gupta, Joon Kyung Kim, Asit Mishra, and Hsien-Hsin S. Lee, 2016
https://scholar.google.com/scholar?q=DNNWeaver%3A+From+High-Level+Deep+Network+Models+to+FPGA+Acceleration
7. Caffeine: Towards Uniformed Representation and Acceleration for Deep Convolutional Neural Networks — Chen Zhang, Peng Li, Guangyu Sun, Yijin Guan, Bingjun Xiao, and Jason Cong, 2016
https://scholar.google.com/scholar?q=Caffeine%3A+Towards+Uniformed+Representation+and+Acceleration+for+Deep+Convolutional+Neural+Networks
8. FINN: A Framework for Fast, Scalable Binarized Neural Network Inference — Yaman Umuroglu, Nicholas J. Fraser, Giulio Gambardella, Michaela Blott, Philip Leong, Magnus Jahre, and Kees Vissers, 2017
https://scholar.google.com/scholar?q=FINN%3A+A+Framework+for+Fast%2C+Scalable+Binarized+Neural+Network+Inference
9. Optimizing FPGA-based Accelerator Design for Deep Convolutional Neural Networks — Yu Wang, Jiajun Xu, Yanzhi Wang, and Huazhong Yang, 2015
https://scholar.google.com/scholar?q=Optimizing+FPGA-based+Accelerator+Design+for+Deep+Convolutional+Neural+Networks
10. Efficient Processing of Deep Neural Networks: A Tutorial and Survey — Vivienne Sze, Yu-Hsin Chen, Tien-Ju Yang, and Joel S. Emer, 2017
https://scholar.google.com/scholar?q=Efficient+Processing+of+Deep+Neural+Networks%3A+A+Tutorial+and+Survey
11. ViTA: A Vision Transformer Inference Accelerator for Edge Applications — Shashank Nag, Gourav Datta, Souvik Kundu, Nitin Chandrachoodan, Peter A. Beerel, 2023
https://scholar.google.com/scholar?q=ViTA%3A+A+Vision+Transformer+Inference+Accelerator+for+Edge+Applications
12. ME-ViT: A Single-Load Memory-Efficient FPGA Accelerator for Vision Transformers — Kyle Marino, Pengmiao Zhang, Viktor K. Prasanna, 2024
https://scholar.google.com/scholar?q=ME-ViT%3A+A+Single-Load+Memory-Efficient+FPGA+Accelerator+for+Vision+Transformers
13. DRViT: A Dynamic Redundancy-Aware Vision Transformer Accelerator via Algorithm and Architecture Co-Design on FPGA — Xiangfeng Sun, Yuanting Zhang, Qinyu Wang, Xiaofeng Zou, et al., 2025
https://scholar.google.com/scholar?q=DRViT%3A+A+Dynamic+Redundancy-Aware+Vision+Transformer+Accelerator+via+Algorithm+and+Architecture+Co-Design+on+FPGA
14. Realisation of Early-Exit Dynamic Neural Networks on Reconfigurable Hardware — Anastasios Dimitriou, Lei Xun, Jonathon Hare, Geoff V. Merrett, 2024
https://scholar.google.com/scholar?q=Realisation+of+Early-Exit+Dynamic+Neural+Networks+on+Reconfigurable+Hardware
15. Compute-In-Memory on FPGAs for Deep Learning: A Review — Aman Arora, 2025
https://scholar.google.com/scholar?q=Compute-In-Memory+on+FPGAs+for+Deep+Learning%3A+A+Review
16. A Heterogeneous System With Computing in Memory Processing Elements to Accelerate CNN Inference — Jinkai Wang, Youxiang Chen, Zekun Wang, Zhengkun Gu, et al., 2025
https://scholar.google.com/scholar?q=A+Heterogeneous+System+With+Computing+in+Memory+Processing+Elements+to+Accelerate+CNN+Inference
17. AI Post Transformers: FPGA Neural Network Accelerators for Space — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-26-fpga-neural-network-accelerators-for-spa-3087ae.mp3
Interactive Visualization: Automating CNN Mapping on Embedded FPGAs

This episode explores a 2017 EPFL paper on turning a soft-input soft-output decoder kernel into a reusable RFNoC FPGA block using Vivado HLS, with a focus on software-defined radio and forward error correction. It explains why SISO processing sits at the core of turbo decoding, walking through the BCJR/MAP decoding logic, trellis-based forward and backward recursions, and the hardware challenges created by state metrics, memory traffic, and iterative probabilistic updates. The discussion argues that the paper is more convincing as a case study in HLS-based FPGA block construction and RFNoC integration than as proof of a complete, production-ready decoder system. Listeners would find it interesting for its clear look at the tradeoff between modular FPGA design convenience and the stubborn algorithmic complexity that still demands careful hardware thinking.

Sources:
1. RFNoC SISO Processor via High-Level Synthesis
https://www.epfl.ch/labs/lap/wp-content/uploads/2018/11/GuerrieriSep17_DesigningAnRfnocBlockImplementingASisoProcessorUsingHighLevelSynthesis_GNURC17.pdf
2. RFNoC: RF Network-on-Chip — Martin Braun, Jonathan Pendlum, Matt Ettus, 2016
https://scholar.google.com/scholar?q=RFNoC%3A+RF+Network-on-Chip
3. Designing a RFNoC Block implementing a SISO Processor using High-Level Synthesis — Andrea Guerrieri, 2017
https://scholar.google.com/scholar?q=Designing+a+RFNoC+Block+implementing+a+SISO+Processor+using+High-Level+Synthesis
4. Measured Latency Introduced by RF Network-on-Chip (RFNoC) Architecture — Joshua Sunderlin, 2017
https://scholar.google.com/scholar?q=Measured+Latency+Introduced+by+RF+Network-on-Chip+%28RFNoC%29+Architecture
5. Adventures in RFNoC: Lessons Learned From Developing a Real-Time Spectrum Sensing Block — Rylee G. Mattingly, Justin G. Metcalf, 2022
https://scholar.google.com/scholar?q=Adventures+in+RFNoC%3A+Lessons+Learned+From+Developing+a+Real-Time+Spectrum+Sensing+Block
6. The Software Radio Architecture — Joseph Mitola III, 1995
https://scholar.google.com/scholar?q=The+Software+Radio+Architecture
7. State of the art baseband DSP platforms for Software Defined Radio: A survey — Omer Anjum, Tapani Ahonen, Fabio Garzia, Jari Nurmi, Claudio Brunelli, Heikki Berg, 2011
https://scholar.google.com/scholar?q=State+of+the+art+baseband+DSP+platforms+for+Software+Defined+Radio%3A+A+survey
8. Sora: High-Performance Software Radio Using General-Purpose Multi-Core Processors — Kun Tan, He Liu, Jiansong Zhang, Yongguang Zhang, Ji Fang, Geoffrey M. Voelker, 2011
https://scholar.google.com/scholar?q=Sora%3A+High-Performance+Software+Radio+Using+General-Purpose+Multi-Core+Processors
9. OpenRadio: A Programmable Wireless Dataplane — Manu Bansal, Jeffrey Mehlman, Sachin Katti, Philip Levis, 2012
https://scholar.google.com/scholar?q=OpenRadio%3A+A+Programmable+Wireless+Dataplane
10. A Mathematical Theory of Communication — Claude E. Shannon, 1948
https://scholar.google.com/scholar?q=A+Mathematical+Theory+of+Communication
11. Low-Density Parity-Check Codes — Robert G. Gallager, 1962
https://scholar.google.com/scholar?q=Low-Density+Parity-Check+Codes
12. Design of Capacity-Approaching Irregular Low-Density Parity-Check Codes — Thomas J. Richardson, M. Amin Shokrollahi, Rudiger L. Urbanke, 2001
https://scholar.google.com/scholar?q=Design+of+Capacity-Approaching+Irregular+Low-Density+Parity-Check+Codes
13. Channel Coding: The Road to Channel Capacity — G. David Forney Jr., Daniel J. Costello Jr., 2007
https://scholar.google.com/scholar?q=Channel+Coding%3A+The+Road+to+Channel+Capacity
14. Near Shannon Limit Error-Correcting Coding and Decoding: Turbo-Codes — Claude Berrou, Alain Glavieux, Punya Thitimajshima, 1993
https://scholar.google.com/scholar?q=Near+Shannon+Limit+Error-Correcting+Coding+and+Decoding%3A+Turbo-Codes
15. A Comparison of Optimal and Sub-Optimal MAP Decoding Algorithms Operating in the Log Domain — Patrick Robertson, Emmanuelle Villebrun, Peter Hoher, 1995
https://scholar.google.com/scholar?q=A+Comparison+of+Optimal+and+Sub-Optimal+MAP+Decoding+Algorithms+Operating+in+the+Log+Domain
16. Turbo Decoding as an Instance of Pearl's "Belief Propagation" Algorithm — Robert J. McEliece, David J. C. MacKay, Jung-Fu Cheng, 1998
https://scholar.google.com/scholar?q=Turbo+Decoding+as+an+Instance+of+Pearl%27s+%22Belief+Propagation%22+Algorithm
17. A Survey of Three-Dimensional Turbo Codes and Recent Performance Enhancements — Dhouha Kbaier Ben Ismail, Catherine Douillard, Sylvie Kerouedan, 2013
https://scholar.google.com/scholar?q=A+Survey+of+Three-Dimensional+Turbo+Codes+and+Recent+Performance+Enhancements
18. Optimal Decoding of Linear Codes for Minimizing Symbol Error Rate — Lalit R. Bahl, John Cocke, Frederick Jelinek, Josef Raviv, 1974
https://scholar.google.com/scholar?q=Optimal+Decoding+of+Linear+Codes+for+Minimizing+Symbol+Error+Rate
19. From Low-Architectural Expertise up to High-Throughput Non-Binary LDPC Decoders: Optimization Guidelines Using High-Level Synthesis — Nithin George, Kimon Karras, David Novo, Vitor Silva, Paolo Ienne, Gabriel Falcao, 2015
https://scholar.google.com/scholar?q=From+Low-Architectural+Expertise+up+to+High-Throughput+Non-Binary+LDPC+Decoders%3A+Optimization+Guidelines+Using+High-Level+Synthesis
20. Turbo Decoder Architecture for Beyond-4G Applications — Cheng-Chi Wong, Hsie-Chia Chang, 2014
https://scholar.google.com/scholar?q=Turbo+Decoder+Architecture+for+Beyond-4G+Applications
21. A DSP Shared Is a DSP Earned: HLS Task-Level Multi-Pumping for High-Performance Low-Resource Designs — approx. recent FPGA/HLS architecture authors, 2023-2026
https://scholar.google.com/scholar?q=A+DSP+Shared+Is+a+DSP+Earned%3A+HLS+Task-Level+Multi-Pumping+for+High-Performance+Low-Resource+Designs
22. Making Acceleration More Amenable with Novel High-Level Synthesis Techniques for FPGAs — approx. recent FPGA/HLS systems authors, 2023-2026
https://scholar.google.com/scholar?q=Making+Acceleration+More+Amenable+with+Novel+High-Level+Synthesis+Techniques+for+FPGAs
23. High-Level Synthesis for FPGAs: A Hardware Engineer's Perspective — approx. recent survey/review authors, 2023-2026
https://scholar.google.com/scholar?q=High-Level+Synthesis+for+FPGAs%3A+A+Hardware+Engineer%27s+Perspective
24. Implementation of Cognitive Radio Based on SDR and FPGA: Methods, Architectures and Practical Applicability — approx. SDR/FPGA survey authors, recent
https://scholar.google.com/scholar?q=Implementation+of+Cognitive+Radio+Based+on+SDR+and+FPGA%3A+Methods%2C+Architectures+and+Practical+Applicability
25. StreamPU: A DSEL for High Throughput and Low Latency Software-Defined Radio on Multicore CPUs — approx. Cassagne and collaborators, recent
https://scholar.google.com/scholar?q=StreamPU%3A+A+DSEL+for+High+Throughput+and+Low+Latency+Software-Defined+Radio+on+Multicore+CPUs
26. Run-Time Reconfigurable Systems and Algorithms for RFSoC-Based Software Defined Radio Applications — approx. recent RFSoC/SDR authors, recent
https://scholar.google.com/scholar?q=Run-Time+Reconfigurable+Systems+and+Algorithms+for+RFSoC-Based+Software+Defined+Radio+Applications
27. DecodeX: Exploring and Benchmarking of LDPC Decoding Across CPU, GPU, and ASIC Platforms — approx. recent benchmarking authors, recent
https://scholar.google.com/scholar?q=DecodeX%3A+Exploring+and+Benchmarking+of+LDPC+Decoding+Across+CPU%2C+GPU%2C+and+ASIC+Platforms
28. High-Throughput Software-Defined LDPC Encoder and Decoder with x86-Based Data-Level Parallelism — approx. recent communications authors, recent
https://scholar.google.com/scholar?q=High-Throughput+Software-Defined+LDPC+Encoder+and+Decoder+with+x86-Based+Data-Level+Parallelism
Interactive Visualization: RFNoC SISO Processor via High-Level Synthesis

This episode explores a 2025 survey of reinforcement learning as a statement about how the field now organizes itself, covering value-based, policy-based, model-based, multi-agent, offline, and LLM-related RL. It explains core concepts like Markov decision processes, policies, value functions, delayed credit assignment, and the contrast between direct policy optimization and methods that estimate action values before deriving behavior. The discussion highlights why actor-critic methods became so central, how model-based RL uses world models to plan ahead, and why offline RL is difficult when agents must improve from fixed logged data rather than fresh interaction. Listeners would find it interesting because it turns a broad survey into a clear map of where reinforcement learning stands in 2025, including the tensions between elegant theory, unstable training, and the practical compromises that shaped modern RL.

Sources:
1. Reinforcement Learning: An Overview — Kevin Murphy, 2024
http://arxiv.org/abs/2412.05265
2. Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems — Sergey Levine, Aviral Kumar, George Tucker, Justin Fu, 2020
https://scholar.google.com/scholar?q=Offline+Reinforcement+Learning%3A+Tutorial%2C+Review%2C+and+Perspectives+on+Open+Problems
3. D4RL: Datasets for Deep Data-Driven Reinforcement Learning — Justin Fu, Aviral Kumar, Ofir Nachum, George Tucker, Sergey Levine, 2020
https://scholar.google.com/scholar?q=D4RL%3A+Datasets+for+Deep+Data-Driven+Reinforcement+Learning
4. Conservative Q-Learning for Offline Reinforcement Learning — Aviral Kumar, Aurick Zhou, George Tucker, Sergey Levine, 2020
https://scholar.google.com/scholar?q=Conservative+Q-Learning+for+Offline+Reinforcement+Learning
5. Decision Transformer: Reinforcement Learning via Sequence Modeling — Lili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee, Aditya Grover, Michael Laskin, Pieter Abbeel, Aravind Srinivas, Igor Mordatch, 2021
https://scholar.google.com/scholar?q=Decision+Transformer%3A+Reinforcement+Learning+via+Sequence+Modeling
6. Reinforcement Learning: An Introduction — Richard S. Sutton and Andrew G. Barto, 2018
https://scholar.google.com/scholar?q=Reinforcement+Learning%3A+An+Introduction
7. Algorithms for Reinforcement Learning — Csaba Szepesvari, 2010
https://scholar.google.com/scholar?q=Algorithms+for+Reinforcement+Learning
8. Human-level control through deep reinforcement learning — Volodymyr Mnih, Koray Kavukcuoglu, David Silver and others, 2015
https://scholar.google.com/scholar?q=Human-level+control+through+deep+reinforcement+learning
9. Trust Region Policy Optimization — John Schulman, Sergey Levine, Philipp Moritz, Michael Jordan and Pieter Abbeel, 2015
https://scholar.google.com/scholar?q=Trust+Region+Policy+Optimization
10. Proximal Policy Optimization Algorithms — John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford and Oleg Klimov, 2017
https://scholar.google.com/scholar?q=Proximal+Policy+Optimization+Algorithms
11. Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model — Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert and others, 2020
https://scholar.google.com/scholar?q=Mastering+Atari%2C+Go%2C+Chess+and+Shogi+by+Planning+with+a+Learned+Model
12. Fine-Tuning Language Models from Human Preferences — Daniel M. Ziegler, Nisan Stiennon, Jeffrey Wu and others, 2019
https://scholar.google.com/scholar?q=Fine-Tuning+Language+Models+from+Human+Preferences
13. Direct Preference Optimization: Your Language Model is Secretly a Reward Model — Rafael Rafailov, Archit Sharma, Eric Mitchell and others, 2023
https://scholar.google.com/scholar?q=Direct+Preference+Optimization%3A+Your+Language+Model+is+Secretly+a+Reward+Model
14. On the Limited Generalization Capability of the Implicit Reward Model Induced by Direct Preference Optimization — Yong Lin et al., 2024
https://scholar.google.com/scholar?q=On+the+Limited+Generalization+Capability+of+the+Implicit+Reward+Model+Induced+by+Direct+Preference+Optimization
15. Q-learning Decision Transformer: Leveraging Dynamic Programming for Conditional Sequence Modelling in Offline RL — Taku Yamagata, Ahmed Khalil, Raul Santos-Rodriguez, 2023
https://scholar.google.com/scholar?q=Q-learning+Decision+Transformer%3A+Leveraging+Dynamic+Programming+for+Conditional+Sequence+Modelling+in+Offline+RL
16. Reinformer: Max-Return Sequence Modeling for Offline RL — Zifeng Zhuang et al., 2024
https://scholar.google.com/scholar?q=Reinformer%3A+Max-Return+Sequence+Modeling+for+Offline+RL
17. Pre-training Contextualized World Models with In-the-wild Videos for Reinforcement Learning — Jialong Wu, Haoyu Ma, Chaoyi Deng, Mingsheng Long, 2023
https://scholar.google.com/scholar?q=Pre-training+Contextualized+World+Models+with+In-the-wild+Videos+for+Reinforcement+Learning
18. PreLAR: World Model Pre-training with Learnable Action Representation — Lixuan Zhang, Meina Kan, Shiguang Shan, Xilin Chen, 2024
https://scholar.google.com/scholar?q=PreLAR%3A+World+Model+Pre-training+with+Learnable+Action+Representation
19. Ctrl-World: A Controllable Generative World Model for Robot Manipulation — Yanjiang Guo, Lucy Xiaoyang Shi, Jianyu Chen, Chelsea Finn, 2025
https://scholar.google.com/scholar?q=Ctrl-World%3A+A+Controllable+Generative+World+Model+for+Robot+Manipulation
20. Skill Transfer and Discovery for Sim-to-Real Learning: A Representation-Based Viewpoint — Haitong Ma, Zhaolin Ren, Bo Dai, Na Li, 2024
https://scholar.google.com/scholar?q=Skill+Transfer+and+Discovery+for+Sim-to-Real+Learning%3A+A+Representation-Based+Viewpoint
21. AI Post Transformers: DreamerV3 World Models Across 150 Tasks — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-20-dreamerv3-world-models-across-150-tasks-af5edb.mp3
22. AI Post Transformers: Experience-Based Learning Beyond Human Data — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-19-experience-based-learning-beyond-human-d-b0caa4.mp3
23. AI Post Transformers: Learning to Reason with 13 Parameters — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-14-learning-to-reason-with-13-parameters-54c87f.mp3
24. AI Post Transformers: ASI-Evolve for Data, Architectures, and RL — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-05-asi-evolve-for-data-architectures-and-rl-197b2b.mp3
Interactive Visualization: Reinforcement Learning in 2025: An Overview

This episode explores a paper arguing that language models can reason more effectively at test time if they are trained to use divide-and-conquer strategies instead of defaulting to a single linear chain of thought. It explains the core distinction between ordinary step-by-step reasoning and structured decomposition into subproblems, then situates that idea alongside prior work such as Tree of Thoughts, Least-to-Most prompting, self-consistency, and recent reasoning-focused post-training. The discussion highlights the paper’s main claim that current post-training regimes bias models toward linear reasoning habits, which can make naive divide-and-conquer prompting underperform unless the decomposition behavior itself is explicitly trained. A listener would find it interesting because it gets at a central question in modern AI: whether better inference-time scaling comes from simply generating longer reasoning traces, or from teaching models to search, branch, and recombine intermediate results in a more algorithmic way.

Sources:
1. Training LLMs for Divide-and-Conquer Reasoning Elevates Test-Time Scalability — Xiao Liang, Zhong-Zhi Li, Zhenghao Lin, Eric Hancheng Jiang, Hengyuan Zhang, Yelong Shen, Kai-Wei Chang, Ying Nian Wu, Yeyun Gong, Weizhu Chen, 2026
http://arxiv.org/abs/2602.02477
2. Tree of Thoughts: Deliberate Problem Solving with Large Language Models — Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, Karthik Narasimhan, 2023
https://scholar.google.com/scholar?q=Tree+of+Thoughts%3A+Deliberate+Problem+Solving+with+Large+Language+Models
3. Parsel: Algorithmic Reasoning with Language Models by Composing Decompositions — Eric Zelikman, Qian Huang, Gabriel Poesia, Noah Goodman, Nick Haber, 2023
https://scholar.google.com/scholar?q=Parsel%3A+Algorithmic+Reasoning+with+Language+Models+by+Composing+Decompositions
4. Training LLMs for Divide-and-Conquer Reasoning Elevates Test-Time Scalability — Xiao Liang, Zhong-Zhi Li, Zhenghao Lin, Eric Hancheng Jiang, Hengyuan Zhang, Yelong Shen, Kai-Wei Chang, Ying Nian Wu, Yeyun Gong, Weizhu Chen, 2026
https://scholar.google.com/scholar?q=Training+LLMs+for+Divide-and-Conquer+Reasoning+Elevates+Test-Time+Scalability
5. Self-Consistency Improves Chain of Thought Reasoning in Language Models — Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, Denny Zhou, 2022
https://scholar.google.com/scholar?q=Self-Consistency+Improves+Chain+of+Thought+Reasoning+in+Language+Models
6. Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters — Charlie Snell, Jaehoon Lee, Kelvin Xu, Aviral Kumar, 2024
https://scholar.google.com/scholar?q=Scaling+LLM+Test-Time+Compute+Optimally+can+be+More+Effective+than+Scaling+Model+Parameters
7. Learning to Reason with LLMs — OpenAI, 2024
https://scholar.google.com/scholar?q=Learning+to+Reason+with+LLMs
8. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning — Daya Guo, Dejian Yang and the DeepSeek-AI team, 2025
https://scholar.google.com/scholar?q=DeepSeek-R1%3A+Incentivizing+Reasoning+Capability+in+LLMs+via+Reinforcement+Learning
9. Least-to-Most Prompting Enables Complex Reasoning in Large Language Models — Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Claire Cui, Olivier Bousquet, Quoc Le, Ed Chi, 2022
https://scholar.google.com/scholar?q=Least-to-Most+Prompting+Enables+Complex+Reasoning+in+Large+Language+Models
10. Decomposed Prompting: A Modular Approach for Solving Complex Tasks — Tushar Khot, Harsh Trivedi, Matthew Finlayson, Yao Fu, Kyle Richardson, Peter Clark, Ashish Sabharwal, 2022
https://scholar.google.com/scholar?q=Decomposed+Prompting%3A+A+Modular+Approach+for+Solving+Complex+Tasks
11. DeepSeek-Prover-V2: Advancing Formal Mathematical Reasoning via Reinforcement Learning for Subgoal Decomposition — Z. Z. Ren, Zhihong Shao, Junxiao Song, Huajian Xin and the DeepSeek team, 2025
https://scholar.google.com/scholar?q=DeepSeek-Prover-V2%3A+Advancing+Formal+Mathematical+Reasoning+via+Reinforcement+Learning+for+Subgoal+Decomposition
12. Decompose, Analyze and Rethink: Solving Intricate Problems with Human-like Reasoning Cycle — Shangzi Xue, Zhenya Huang, Jiayu Liu, Xin Lin, Yuting Ning, Binbin Jin, Xin Li, Qi Liu, 2024
https://scholar.google.com/scholar?q=Decompose%2C+Analyze+and+Rethink%3A+Solving+Intricate+Problems+with+Human-like+Reasoning+Cycle
13. Seed-Prover: Deep and Broad Reasoning for Automated Theorem Proving — Luoxin Chen, Jinming Gu, Liankai Huang, Wenhao Huang, Zhicheng Jiang, Allan Jie, Xiaoran Jin, Xing Jin, Chenggang Li, Kaijing Ma, Cheng Ren, Jiawei Shen, Wenlei Shi, Tong Sun, He Sun, Jiahui Wang, Siran Wang, Zhihong Wang, Chenrui Wei, Shufa Wei, Yonghui Wu, Yuchen Wu, et al., 2025
https://scholar.google.com/scholar?q=Seed-Prover%3A+Deep+and+Broad+Reasoning+for+Automated+Theorem+Proving
14. DAPO: An Open-Source LLM Reinforcement Learning System at Scale — Qiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan, Xiaochen Zuo, Yu Yue, Weinan Dai, Tiantian Fan, Gaohong Liu, Lingjun Liu, Xin Liu, Haibin Lin, Zhiqi Lin, Bole Ma, Guangming Sheng, Yuxuan Tong, Chi Zhang, Mofan Zhang, Wang Zhang, Hang Zhu, Jinhua Zhu, Jiaze Chen, Jiangjie Chen, Chengyi Wang, Hongli Yu, Yuxuan Song, Xiangpeng Wei, Hao Zhou, Jingjing Liu, Wei-Ying Ma, Ya-Qin Zhang, Lin Yan, Mu Qiao, Yonghui Wu, Mingxuan Wang, 2025
https://scholar.google.com/scholar?q=DAPO%3A+An+Open-Source+LLM+Reinforcement+Learning+System+at+Scale
15. Continuous Chain of Thought Enables Parallel Exploration and Reasoning — Halil Alperen Gozeten et al., 2025
https://scholar.google.com/scholar?q=Continuous+Chain+of+Thought+Enables+Parallel+Exploration+and+Reasoning
16. How to Think Step-by-Step: A Mechanistic Understanding of Chain-of-Thought Reasoning — Subhabrata Dutta et al., 2024
https://scholar.google.com/scholar?q=How+to+Think+Step-by-Step%3A+A+Mechanistic+Understanding+of+Chain-of-Thought+Reasoning
17. Decompose-ToM: Enhancing Theory of Mind Reasoning in Large Language Models through Simulation and Task Decomposition — Sneheel Sarangi et al., 2025
https://scholar.google.com/scholar?q=Decompose-ToM%3A+Enhancing+Theory+of+Mind+Reasoning+in+Large+Language+Models+through+Simulation+and+Task+Decomposition
18. Select-Then-Decompose: From Empirical Analysis to Adaptive Selection Strategy for Task Decomposition in Large Language Models — Shuodi Liu et al., 2025
https://scholar.google.com/scholar?q=Select-Then-Decompose%3A+From+Empirical+Analysis+to+Adaptive+Selection+Strategy+for+Task+Decomposition+in+Large+Language+Models
19. LearNAT: Learning NL2SQL with AST-guided Task Decomposition for Large Language Models — Weibin Liao et al., 2025
https://scholar.google.com/scholar?q=LearNAT%3A+Learning+NL2SQL+with+AST-guided+Task+Decomposition+for+Large+Language+Models
20. Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning — Chang Tian et al., 2025
https://scholar.google.com/scholar?q=Large+Language+Models+Reasoning+Abilities+Under+Non-Ideal+Conditions+After+RL-Fine-Tuning
21. AI Post Transformers: Learning to Reason with 13 Parameters — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-14-learning-to-reason-with-13-parameters-54c87f.mp3
22. AI Post Transformers: IMO-Bench for Robust Mathematical Reasoning — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-imo-bench-for-robust-mathematical-reason-143489.mp3
23. AI Post Transformers: Test-time Scaling for Multi-Agent Collaborative Reasoning — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-22-test-time-scaling-for-multi-agent-collab-082570.mp3
24. AI Post Transformers: Chain-of-Thought Reasoning: A Brittle Mirage? — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/chain-of-thought-reasoning-a-brittle-mirage/
25. AI Post Transformers: NeurIPS 2025: Reinforcement Learning for Reasoning in Large Language Models with One Training Example — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/neurips-2025-reinforcement-learning-for-reasoning-in-large-language-models-with/
26. AI Post Transformers: TraceRL: Reinforcement Learning for Diffusion Language Models — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/tracerl-reinforcement-learning-for-diffusion-language-models/
Interactive Visualization: Training LLMs for Divide-and-Conquer Reasoning

This episode explores PackKV, a method for shrinking the transformer KV cache during long-context inference by combining low-bit quantization with GPU-friendly repacking and lossy compression. It explains why KV cache growth can dominate memory use in large models, using examples where cache size exceeds model weights, and frames the problem as a systems bottleneck driven more by memory traffic than raw computation. The discussion compares PackKV to prior approaches such as KV quantization, token pruning, and offloading to CPU memory, highlighting the paper’s argument that compression is only useful if decompression is tightly integrated into the inference pipeline. A listener would find it interesting because it turns a seemingly low-level optimization into a broader claim about how future long-context LLM performance may depend as much on memory layout and kernel design as on model architecture.

Interactive Visualization: PackKV Lossy Compression for KV Caches
Sources:
1. PackKV: Reducing KV Cache Memory Footprint through LLM-Aware Lossy Compression — Bo Jiang, Taolue Yang, Youyuan Liu, Xubin He, Sheng Di, Sian Jin, 2025
http://arxiv.org/abs/2512.24449
2. Scissorhands: Exploiting the Persistence of Importance Hypothesis for LLM KV Cache Compression at Test Time — Zichang Liu, Aditya Desai, Fangshuo Liao, Victor Xie, Zhaozhuo Xu, Anastasios Kyrillidis, Anshumali Shrivastava, et al., 2023
https://scholar.google.com/scholar?q=Scissorhands%3A+Exploiting+the+Persistence+of+Importance+Hypothesis+for+LLM+KV+Cache+Compression+at+Test+Time
3. H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models — Zhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen, Lianmin Zheng, Zhao Song, Yuandong Tian, Clark Barrett, Zhangyang Wang, Beidi Chen, et al., 2023
https://scholar.google.com/scholar?q=H2O%3A+Heavy-Hitter+Oracle+for+Efficient+Generative+Inference+of+Large+Language+Models
4. KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache — Zirui Liu, Jiayi Yuan, Hongye Jin, Shaochen Zhong, Zhaozhuo Xu, Vladimir Braverman, Beidi Chen, Xia Hu, 2024
https://scholar.google.com/scholar?q=KIVI%3A+A+Tuning-Free+Asymmetric+2bit+Quantization+for+KV+Cache
5. SnapKV: LLM Knows What You are Looking for Before Generation — Yuhong Li, Yingbing Huang, Bowen Yang, Bharat Venkitesh, Acyr Locatelli, Hanchen Ye, Tianle Cai, Patrick Lewis, Deming Chen, 2024
https://scholar.google.com/scholar?q=SnapKV%3A+LLM+Knows+What+You+are+Looking+for+Before+Generation
6. CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving — Yuyang Liu, Haotian Li, Yao Cheng, Siddhant Ray, Yizhou Huang, Qizhen Zhang, Kaixiang Du, Jinyang Yao, Shan Lu, Ganesh Ananthanarayanan et al., 2024
https://scholar.google.com/scholar?q=CacheGen%3A+KV+Cache+Compression+and+Streaming+for+Fast+Large+Language+Model+Serving
7. Q-Hitter: A Better Token Oracle for Efficient LLM Inference via Sparse-Quantized KV Cache — Zhuodong Zhang, Shang Liu, Ruobing Chen, Bhavya Kailkhura, Ben Chen, An Wang, 2024
https://scholar.google.com/scholar?q=Q-Hitter%3A+A+Better+Token+Oracle+for+Efficient+LLM+Inference+via+Sparse-Quantized+KV+Cache
8. PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling — Zefan Cai, Yichi Zhang, Bofei Gao, Yuliang Liu, Tianyu Liu, Keming Lu, Wayne Xiong, Yue Dong, Baobao Chang, Junjie Hu, Wen Xiao, 2025
https://scholar.google.com/scholar?q=PyramidKV%3A+Dynamic+KV+Cache+Compression+based+on+Pyramidal+Information+Funneling
9. Expected Attention: KV Cache Compression by Estimating Attention from Future Queries Distribution — Alessio Devoto, Maximilian Jeblick, Simon Jegou, 2025
https://scholar.google.com/scholar?q=Expected+Attention%3A+KV+Cache+Compression+by+Estimating+Attention+from+Future+Queries+Distribution
10. TurboQuant: Online Vector Quantization with Near-Optimal Distortion — Amir Zandieh, Majid Daliri, Majid Hadian, Vahab Mirrokni, 2025
https://scholar.google.com/scholar?q=TurboQuant%3A+Online+Vector+Quantization+with+Near-Optimal+Distortion
11. KVLink: Accelerating Large Language Models via Efficient KV Cache Reuse — Jingbo Yang et al., 2025
https://scholar.google.com/scholar?q=KVLink%3A+Accelerating+Large+Language+Models+via+Efficient+KV+Cache+Reuse
12. HyperRAG: Enhancing Quality-Efficiency Tradeoffs in Retrieval-Augmented Generation with Reranker KV-Cache Reuse — Yuwei An et al., 2025
https://scholar.google.com/scholar?q=HyperRAG%3A+Enhancing+Quality-Efficiency+Tradeoffs+in+Retrieval-Augmented+Generation+with+Reranker+KV-Cache+Reuse
13. ProphetKV: User-Query-Driven Selective Recomputation for Efficient KV Cache Reuse in Retrieval-Augmented Generation — Shihao Wang et al., 2026
https://scholar.google.com/scholar?q=ProphetKV%3A+User-Query-Driven+Selective+Recomputation+for+Efficient+KV+Cache+Reuse+in+Retrieval-Augmented+Generation
14. ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification — Yefei He et al., 2024
https://scholar.google.com/scholar?q=ZipCache%3A+Accurate+and+Efficient+KV+Cache+Quantization+with+Salient+Token+Identification
15. KVSink: Understanding and Enhancing the Preservation of Attention Sinks in KV Cache Quantization for LLMs — Zunhai Su and Kehong Yuan, 2025
https://scholar.google.com/scholar?q=KVSink%3A+Understanding+and+Enhancing+the+Preservation+of+Attention+Sinks+in+KV+Cache+Quantization+for+LLMs
16. ThinK: Thinner Key Cache by Query-Driven Pruning — Yuhui Xu et al., 2024
https://scholar.google.com/scholar?q=ThinK%3A+Thinner+Key+Cache+by+Query-Driven+Pruning
17. KV-Compress: Paged KV-Cache Compression with Variable Compression Rates per Attention Head — Isaac Rehg, 2024
https://scholar.google.com/scholar?q=KV-Compress%3A+Paged+KV-Cache+Compression+with+Variable+Compression+Rates+per+Attention+Head
18. Paged Attention Meets FlexAttention: Unlocking Long-Context Efficiency in Deployed Inference — Thomas Joshi et al., 2025
https://scholar.google.com/scholar?q=Paged+Attention+Meets+FlexAttention%3A+Unlocking+Long-Context+Efficiency+in+Deployed+Inference
19. AI Post Transformers: TokenDance for Multi-Agent KV Cache Sharing — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-22-tokendance-for-multi-agent-kv-cache-shar-aa9b99.mp3
20. AI Post Transformers: KVSwap for Disk-Aware Long-Context On-Device Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-16-kvswap-for-disk-aware-long-context-on-de-f3c15e.mp3
21. AI Post Transformers: CacheFlow and 3D-Parallel KV Cache Restoration — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-01-cacheflow-and-3d-parallel-kv-cache-resto-8db883.mp3
22. AI Post Transformers: TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-25-turboquant-online-vector-quantiz-1967b7.mp3
23. AI Post Transformers: FengHuang for Rack-Scale LLM Inference Memory — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-12-fenghuang-for-rack-scale-llm-inference-m-62708e.mp3
24. AI Post Transformers: Computation-Bandwidth-Memory Trade-offs for AI Infrastructure — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-09-computation-bandwidth-memory-trade-offs-a83f2b.mp3
Interactive Visualization: PackKV Lossy Compression for KV Caches

This episode explores how the OOMB training system tries to break the memory bottleneck that makes million-token language model training impractical, focusing on why training long contexts is much harder than simply extending inference-time context windows. It explains the paper’s core ideas in plain language, including chunk-recurrent training that recomputes activations during backpropagation, O(1)-style activation memory, and the harder remaining problem of storing and moving the KV cache across extremely long sequences. The discussion also weighs the paper’s central claim with healthy skepticism, asking whether fitting multi-million-token training steps on a single GPU proves genuinely useful long-range learning or mainly demonstrates a strong systems optimization. Listeners would find it interesting because it connects deep learning mechanics, hardware limits, and competing long-context strategies like Ring Attention into a clear debate about what real progress in long-context LLMs should look like.

Sources:
1. Out of the Memory Barrier: A Highly Memory Efficient Training System for LLMs with Million-Token Contexts — Wenhao Li, Daohai Yu, Gen Luo, Yuxin Zhang, Fei Chao, Rongrong Ji, Yifan Wu, Jiaxin Liu, Ziyang Gong, Zimu Liao, 2026
http://arxiv.org/abs/2602.02108
2. Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context — Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V. Le, Ruslan Salakhutdinov, 2019
https://scholar.google.com/scholar?q=Transformer-XL%3A+Attentive+Language+Models+Beyond+a+Fixed-Length+Context
3. Recurrent Memory Transformer — Aydar Bulatov, Yuri Kuratov, Mikhail S. Burtsev, 2022
https://scholar.google.com/scholar?q=Recurrent+Memory+Transformer
4. Ring Attention with Blockwise Transformers for Near-Infinite Context — Hao Liu, Matei Zaharia, Pieter Abbeel, 2023
https://scholar.google.com/scholar?q=Ring+Attention+with+Blockwise+Transformers+for+Near-Infinite+Context
5. Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention — Tsendsuren Munkhdalai, Manaal Faruqui, Siddharth Gopal, 2024
https://scholar.google.com/scholar?q=Leave+No+Context+Behind%3A+Efficient+Infinite+Context+Transformers+with+Infini-attention
6. LongLoRA: Efficient Fine-tuning of Long-Context Large Language Models — Yingfeng Chen et al., 2024
https://scholar.google.com/scholar?q=LongLoRA%3A+Efficient+Fine-tuning+of+Long-Context+Large+Language+Models
7. Efficient Memory Management for Large Language Model Serving with PagedAttention — Woosuk Kwon et al., 2023
https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention
8. Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism — Mohammad Shoeybi et al., 2020
https://scholar.google.com/scholar?q=Megatron-LM%3A+Training+Multi-Billion+Parameter+Language+Models+Using+Model+Parallelism
9. GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints — Joshua Ainslie et al., 2023
https://scholar.google.com/scholar?q=GQA%3A+Training+Generalized+Multi-Query+Transformer+Models+from+Multi-Head+Checkpoints
10. ZeRO-Offload: Democratizing Billion-Scale Model Training — Samyam Rajbhandari et al., 2020
https://scholar.google.com/scholar?q=ZeRO-Offload%3A+Democratizing+Billion-Scale+Model+Training
11. Long Context Compression with Activation Beacon — approx. Liu et al., 2024
https://scholar.google.com/scholar?q=Long+Context+Compression+with+Activation+Beacon
12. Boosting Long-Context Information Seeking via Query-Guided Activation Refilling — approx. unknown from snippet, 2024 or 2025
https://scholar.google.com/scholar?q=Boosting+Long-Context+Information+Seeking+via+Query-Guided+Activation+Refilling
13. Kvlink: Accelerating Large Language Models via Efficient KV Cache Reuse — approx. unknown from snippet, 2024 or 2025
https://scholar.google.com/scholar?q=Kvlink%3A+Accelerating+Large+Language+Models+via+Efficient+KV+Cache+Reuse
14. SuperOffload: Unleashing the Power of Large-Scale LLM Training on Superchips — approx. unknown from snippet, 2024 or 2025
https://scholar.google.com/scholar?q=SuperOffload%3A+Unleashing+the+Power+of+Large-Scale+LLM+Training+on+Superchips
15. SPPO: Efficient Long-Sequence LLM Training via Adaptive Sequence Pipeline Parallel Offloading — approx. unknown from snippet, 2024 or 2025
https://scholar.google.com/scholar?q=SPPO%3A+Efficient+Long-Sequence+LLM+Training+via+Adaptive+Sequence+Pipeline+Parallel+Offloading
16. Keep the Cost Down: A Review on Methods to Optimize LLM's KV-Cache Consumption — approx. unknown from snippet, 2024 or 2025
https://scholar.google.com/scholar?q=Keep+the+Cost+Down%3A+A+Review+on+Methods+to+Optimize+LLM%27s+KV-Cache+Consumption
17. AI Post Transformers: DeepSeek-V4 and Practical Million-Token Context — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-25-deepseek-v4-and-practical-million-token-6f4de1.mp3
18. AI Post Transformers: KVSwap for Disk-Aware Long-Context On-Device Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-16-kvswap-for-disk-aware-long-context-on-de-f3c15e.mp3
19. AI Post Transformers: RetrievalAttention for Long-Context LLM Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-17-retrievalattention-for-long-context-llm-ddf566.mp3
20. AI Post Transformers: CacheFlow and 3D-Parallel KV Cache Restoration — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-01-cacheflow-and-3d-parallel-kv-cache-resto-8db883.mp3
21. AI Post Transformers: Gated Linear Attention for Efficient Long Sequences — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-18-gated-linear-attention-for-efficient-lon-c858ab.mp3
22. AI Post Transformers: Parallelizing DeltaNet Linear Transformers over Sequence Length — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-18-parallelizing-deltanet-linear-transforme-2d0377.mp3
Interactive Visualization: Training Million-Token LLMs Beyond the Memory Barrier

This episode explores how transformers split prediction between knowledge stored in their weights and information inferred from the current prompt, using the paper’s synthetic “bigram world” to make those mechanisms visible. It explains the distinction between global statistical knowledge and true in-context knowledge, then walks through induction heads as a concrete circuit for recalling earlier patterns and continuing them later. The discussion highlights the paper’s main finding that models learn easy dataset-wide averages first, while context-sensitive induction behavior emerges later and requires the right architecture, with two-layer transformers succeeding where one-layer models fail. Listeners would find it interesting because it turns a vague claim about in-context learning into a causal, mechanistic story about how temporary memory may actually form during training.

Sources:
1. Birth of a Transformer: A Memory Viewpoint — Alberto Bietti, Vivien Cabannes, Diane Bouchacourt, Herve Jegou, Leon Bottou, 2023
http://arxiv.org/abs/2306.00802
2. A Mathematical Framework for Transformer Circuits — Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, Chris Olah, 2021
https://scholar.google.com/scholar?q=A+Mathematical+Framework+for+Transformer+Circuits
3. In-context Learning and Induction Heads — Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Scott Johnston, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, Chris Olah, 2022
https://scholar.google.com/scholar?q=In-context+Learning+and+Induction+Heads
4. Birth of a Transformer: A Memory Viewpoint — Alberto Bietti, Vivien Cabannes, Diane Bouchacourt, Herve Jegou, Leon Bottou, 2023
https://scholar.google.com/scholar?q=Birth+of+a+Transformer%3A+A+Memory+Viewpoint
5. What Learning Algorithm Is In-Context Learning? Investigations with Linear Models — Ekin Akyurek, Dale Schuurmans, Jacob Andreas, Tengyu Ma, Denny Zhou, 2023
https://scholar.google.com/scholar?q=What+Learning+Algorithm+Is+In-Context+Learning%3F+Investigations+with+Linear+Models
6. Data Distributional Properties Drive Emergent In-Context Learning in Transformers — Stephanie C. Y. Chan, Adam Santoro, Andrew Lampinen, Jane Wang, Aaditya Singh, Pierre Richemond, Jay McClelland, Felix Hill, 2022
https://scholar.google.com/scholar?q=Data+Distributional+Properties+Drive+Emergent+In-Context+Learning+in+Transformers
7. Transformer Feed-Forward Layers Are Key-Value Memories — Mor Geva, Roei Schuster, Jonathan Berant, Omer Levy, 2021
https://scholar.google.com/scholar?q=Transformer+Feed-Forward+Layers+Are+Key-Value+Memories
8. Dissecting Recall of Factual Associations in Auto-Regressive Language Models — Mor Geva, Jasmijn Bastings, Katja Filippova, Amir Globerson, 2023
https://scholar.google.com/scholar?q=Dissecting+Recall+of+Factual+Associations+in+Auto-Regressive+Language+Models
9. What needs to go right for an induction head? A mechanistic study of in-context learning circuits and their formation — Aaditya K. Singh, Ted Moskovitz, Felix Hill, Stephanie C. Y. Chan, Andrew M. Saxe, 2024
https://scholar.google.com/scholar?q=What+needs+to+go+right+for+an+induction+head%3F+A+mechanistic+study+of+in-context+learning+circuits+and+their+formation
10. Learning to grok: Emergence of in-context learning and skill composition in modular arithmetic tasks — Tianyu He, Darshil Doshi, Aritra Das, Andrey Gromov, 2024
https://scholar.google.com/scholar?q=Learning+to+grok%3A+Emergence+of+in-context+learning+and+skill+composition+in+modular+arithmetic+tasks
11. Selective Induction Heads: How Transformers Select Causal Structures In Context — Francesco D'Angelo, Francesco Croce, Nicolas Flammarion, 2025
https://scholar.google.com/scholar?q=Selective+Induction+Heads%3A+How+Transformers+Select+Causal+Structures+In+Context
12. Induction Head Toxicity Mechanistically Explains Repetition Curse in Large Language Models — Shuxun Wang, Qingyu Yin, Chak Tou Leong, Qiang Zhang, Linyi Yang, 2025
https://scholar.google.com/scholar?q=Induction+Head+Toxicity+Mechanistically+Explains+Repetition+Curse+in+Large+Language+Models
13. AI Post Transformers: In-Place Test-Time Training for Transformers — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-09-in-place-test-time-training-for-transfor-d0b976.mp3
14. AI Post Transformers: Linear Classifier Probes for Intermediate Layers — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-16-linear-classifier-probes-for-intermediat-927ae3.mp3
15. AI Post Transformers: Parallelizing DeltaNet Linear Transformers over Sequence Length — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-18-parallelizing-deltanet-linear-transforme-2d0377.mp3
16. AI Post Transformers: Gated Delta Networks for Long-Context Retrieval — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-17-gated-delta-networks-for-long-context-re-706d85.mp3
Interactive Visualization: How Induction Heads Emerge in Transformers

This episode explores how DeepWalk helped launch modern graph representation learning by turning random walks over a social network into “sentences” and then applying the Skip-Gram ideas behind word2vec to learn dense node embeddings. It explains why that mattered in 2014: instead of relying on heavy spectral or matrix-factorization methods, DeepWalk offered an online, scalable way to learn reusable graph features that worked especially well when labeled data was scarce. The discussion digs into the paper’s main empirical claim that, on social-network benchmarks like BlogCatalog, Flickr, and YouTube, the method substantially improved node classification under sparse-label settings. It is interesting because the conversation goes beyond the headline result and asks what really drove the gains: the language-modeling objective, the community-biased random-walk sampler, or simply a better optimization setup for homophilous graphs.

Sources:
1. DeepWalk and the Rise of Graph Embeddings
https://arxiv.org/pdf/1403.6652
2. Distributed Representations of Words and Phrases and their Compositionality — Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg Corrado, Jeffrey Dean, 2013
https://scholar.google.com/scholar?q=Distributed+Representations+of+Words+and+Phrases+and+their+Compositionality
3. Efficient Estimation of Word Representations in Vector Space — Tomas Mikolov, Kai Chen, Greg Corrado, Jeffrey Dean, 2013
https://scholar.google.com/scholar?q=Efficient+Estimation+of+Word+Representations+in+Vector+Space
4. Learning Latent Social Dimensions for Link Prediction — Lei Tang, Huan Liu, 2011
https://scholar.google.com/scholar?q=Learning+Latent+Social+Dimensions+for+Link+Prediction
5. LINE: Large-scale Information Network Embedding — Jian Tang, Meng Qu, Mingzhe Wang, Ming Zhang, Jun Yan, Qiaozhu Mei, 2015
https://scholar.google.com/scholar?q=LINE%3A+Large-scale+Information+Network+Embedding
6. node2vec: Scalable Feature Learning for Networks — Aditya Grover, Jure Leskovec, 2016
https://scholar.google.com/scholar?q=node2vec%3A+Scalable+Feature+Learning+for+Networks
7. Planetoid: Inductive Representation Learning on Large Graphs — Zhilin Yang, William W. Cohen, Ruslan Salakhutdinov, 2016
https://scholar.google.com/scholar?q=Planetoid%3A+Inductive+Representation+Learning+on+Large+Graphs
8. Node Embedding for Homophilous Graphs with ARGEW: Augmentation of Random walks by Graph Edge Weights — authors not shown in snippet, recent; exact year not shown
https://scholar.google.com/scholar?q=Node+Embedding+for+Homophilous+Graphs+with+ARGEW%3A+Augmentation+of+Random+walks+by+Graph+Edge+Weights
9. Graph neural networks for graphs with heterophily: A survey — authors not shown in snippet, recent; exact year not shown
https://scholar.google.com/scholar?q=Graph+neural+networks+for+graphs+with+heterophily%3A+A+survey
10. Learning attribute and homophily measures through random walks — authors not shown in snippet, recent; exact year not shown
https://scholar.google.com/scholar?q=Learning+attribute+and+homophily+measures+through+random+walks
11. Graph Node Embedding by Neighborhood Prediction Based on Multiview Contrastive Learning — authors not shown in snippet, recent; exact year not shown
https://scholar.google.com/scholar?q=Graph+Node+Embedding+by+Neighborhood+Prediction+Based+on+Multiview+Contrastive+Learning
12. Dynamic graph representation learning with neural networks: A survey — authors not shown in snippet, recent; exact year not shown
https://scholar.google.com/scholar?q=Dynamic+graph+representation+learning+with+neural+networks%3A+A+survey
13. AI Post Transformers: GraphSAGE: Inductive Representation Learning on Large Graphs — Hal Turing & Dr. Ada Shannon, Mon,
https://podcast.do-not-panic.com/episodes/graphsage-inductive-representation-learning-on-large-graphs/
Interactive Visualization: DeepWalk and the Rise of Graph Embeddings

This episode explores how node2vec adapts the word2vec idea to graphs by learning node embeddings from random walks instead of hand-engineered network features. It explains the core technical move in detail: second-order walks controlled by the `p` and `q` parameters, which bias the sampling process toward more local, BFS-like neighborhoods or more exploratory, DFS-like paths. The discussion highlights the paper’s main claim that this tunable notion of context can capture both homophily and structural roles, while also questioning how strongly the experiments actually prove that flexibility versus simply showing better benchmark performance. Listeners would find it interesting for its clear breakdown of why node2vec became influential: it made graph representation learning feel practical, scalable, and easy to use before modern graph neural methods took over.

Interactive Visualization: node2vec and Learning Graph Embeddings
Sources:
1. node2vec: Scalable Feature Learning for Networks — Aditya Grover, Jure Leskovec, 2016
http://arxiv.org/abs/1607.00653
2. DeepWalk: Online Learning of Social Representations — Bryan Perozzi, Rami Al-Rfou, Steven Skiena, 2014
https://scholar.google.com/scholar?q=DeepWalk%3A+Online+Learning+of+Social+Representations
3. LINE: Large-scale Information Network Embedding — Jian Tang, Meng Qu, Mingzhe Wang, Ming Zhang, Jun Yan, Qiaozhu Mei, 2015
https://scholar.google.com/scholar?q=LINE%3A+Large-scale+Information+Network+Embedding
4. node2vec: Scalable Feature Learning for Networks — Aditya Grover, Jure Leskovec, 2016
https://scholar.google.com/scholar?q=node2vec%3A+Scalable+Feature+Learning+for+Networks
5. Inductive Representation Learning on Large Graphs — William L. Hamilton, Rex Ying, Jure Leskovec, 2017
https://scholar.google.com/scholar?q=Inductive+Representation+Learning+on+Large+Graphs
6. The PageRank Citation Ranking: Bringing Order to the Web — Lawrence Page, Sergey Brin, Rajeev Motwani, Terry Winograd, 1999
https://scholar.google.com/scholar?q=The+PageRank+Citation+Ranking%3A+Bringing+Order+to+the+Web
7. Supervised Random Walks: Predicting and Recommending Links in Social Networks — Lars Backstrom, Jure Leskovec, 2011
https://scholar.google.com/scholar?q=Supervised+Random+Walks%3A+Predicting+and+Recommending+Links+in+Social+Networks
8. Pixie: A System for Recommending 3+ Billion Items to 200+ Million Users in Real-Time — Chantat Eksombatchai, Pranav Jindal, Jerry Zitao Liu, Anand Sharma, Charles Sugnet, Mark Ulrich, Jure Leskovec, 2018
https://scholar.google.com/scholar?q=Pixie%3A+A+System+for+Recommending+3%2B+Billion+Items+to+200%2B+Million+Users+in+Real-Time
9. The Link-Prediction Problem for Social Networks — David Liben-Nowell, Jon Kleinberg, 2007
https://scholar.google.com/scholar?q=The+Link-Prediction+Problem+for+Social+Networks
10. Link Prediction Based on Graph Neural Networks — Muhan Zhang, Zhicheng Cui, Marion Neumann, Yixin Chen, 2018
https://scholar.google.com/scholar?q=Link+Prediction+Based+on+Graph+Neural+Networks
11. Revisiting Semi-Supervised Learning with Graph Embeddings — Zhilin Yang, William W. Cohen, Ruslan Salakhutdinov, 2016
https://scholar.google.com/scholar?q=Revisiting+Semi-Supervised+Learning+with+Graph+Embeddings
12. Semi-Supervised Classification with Graph Convolutional Networks — Thomas N. Kipf, Max Welling, 2017
https://scholar.google.com/scholar?q=Semi-Supervised+Classification+with+Graph+Convolutional+Networks
13. word2vec Explained: Deriving Mikolov et al.'s Negative-Sampling Word-Embedding Method — Yoav Goldberg, Omer Levy, 2014
https://scholar.google.com/scholar?q=word2vec+Explained%3A+Deriving+Mikolov+et+al.%27s+Negative-Sampling+Word-Embedding+Method
14. RolX: Structural Role Extraction and Mining in Large Graphs — Keith Henderson, Brian Gallagher, Tina Eliassi-Rad, Hanghang Tong, S. H. Akoglu, Danai Koutra, Christos Faloutsos, Lei Li, 2012
https://scholar.google.com/scholar?q=RolX%3A+Structural+Role+Extraction+and+Mining+in+Large+Graphs
15. Role-aware random walk for network embedding — Hegui Zhang, Gang Kou, Yi Peng, Boyu Zhang, 2024
https://scholar.google.com/scholar?q=Role-aware+random+walk+for+network+embedding
16. WalkLM: A Uniform Language Model Fine-tuning Framework for Attributed Graph Embedding — Yanchao Tan, Zihao Zhou, Hang Lv, Weiming Liu, Carl Yang, 2023
https://scholar.google.com/scholar?q=WalkLM%3A+A+Uniform+Language+Model+Fine-tuning+Framework+for+Attributed+Graph+Embedding
17. INCREASE: Inductive Graph Representation Learning for Spatio-Temporal Kriging — Chuanpan Zheng, Xiaoliang Fan, Cheng Wang, Jianzhong Qi, Chaochao Chen, Longbiao Chen, 2023
https://scholar.google.com/scholar?q=INCREASE%3A+Inductive+Graph+Representation+Learning+for+Spatio-Temporal+Kriging
18. Graph Condensation for Inductive Node Representation Learning — Xinyi Gao, Tong Chen, Yilong Zang, Wentao Zhang, et al., 2023
https://scholar.google.com/scholar?q=Graph+Condensation+for+Inductive+Node+Representation+Learning
19. Distributed Graph Embedding with Information-Oriented Random Walks — Peng Fang, Arijit Khan, Siqiang Luo, Fang Wang, et al., 2023
https://scholar.google.com/scholar?q=Distributed+Graph+Embedding+with+Information-Oriented+Random+Walks
20. AI Post Transformers: GraphSAGE: Inductive Representation Learning on Large Graphs — Hal Turing & Dr. Ada Shannon, Mon,
https://podcast.do-not-panic.com/episodes/graphsage-inductive-representation-learning-on-large-graphs/
Interactive Visualization: node2vec and Learning Graph Embeddings

This episode explores the idea of “self-improving pretraining,” where already post-trained models are used to shape the pretraining of new models rather than waiting to add safety, reasoning, and factuality later. It explains how the approach rewrites training continuations, uses stronger models as judges, and compares original corpus text, teacher-generated suffixes, and learner rollouts to push model preferences upstream into training. The discussion also situates the paper against earlier work like Constitutional AI, STaR, and Quiet-STaR, while debating whether this is a genuine shift in training philosophy or mainly a more aggressive form of distilling a stronger model’s preferences. Listeners would find it interesting because it gets at a central question in modern AI: whether better behavior can be built into a model’s foundations instead of patched on after the fact.

Sources:
1. Self-Improving Pretraining: using post-trained models to pretrain better models — Ellen Xiaoqing Tan, Jack Lanchantin, Shehzaad Dhuliawala, Danwei Li, Thao Nguyen, Jing Xu, Ping Yu, Ilia Kulikov, Sainbayar Sukhbaatar, Jason Weston, Xian Li, Olga Golovneva, 2026
http://arxiv.org/abs/2601.21343
2. Constitutional AI: Harmlessness from AI Feedback — Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, et al., 2022
https://scholar.google.com/scholar?q=Constitutional+AI%3A+Harmlessness+from+AI+Feedback
3. STaR: Self-Taught Reasoner Bootstrapping Reasoning With Reasoning — Eric Zelikman, Yuhuai Wu, Jesse Mu, Noah D. Goodman, and Percy Liang, 2022
https://scholar.google.com/scholar?q=STaR%3A+Self-Taught+Reasoner+Bootstrapping+Reasoning+With+Reasoning
4. Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking — Eric Zelikman, Yuhuai Wu, Noah D. Goodman, and Jakob Foerster, 2024
https://scholar.google.com/scholar?q=Quiet-STaR%3A+Language+Models+Can+Teach+Themselves+to+Think+Before+Speaking
5. Self-Rewarding Language Models — Natasha Shumailov, John Adler, Ming-Wei Chang, Sharan Narang, and Yi Tay, 2024
https://scholar.google.com/scholar?q=Self-Rewarding+Language+Models
6. Confronting Reward Model Overoptimization with Constrained RLHF — Nathan Lambert, Louis Castricato, Leandro von Werra, and Alex Havrilla, 2024
https://scholar.google.com/scholar?q=Confronting+Reward+Model+Overoptimization+with+Constrained+RLHF
7. On the Diversity of Synthetic Data and its Impact on Training Large Language Models — Hao Chen, Abdul Waheed, Xiang Li, Yidong Wang, Jindong Wang, Bhiksha Raj, Marah I. Abdin, 2024
https://scholar.google.com/scholar?q=On+the+Diversity+of+Synthetic+Data+and+its+Impact+on+Training+Large+Language+Models
8. In Their Own Words: Reasoning Traces Tailored for Small Models Make Them Better Reasoners — Jaehoon Kim, Kwangwook Seo, Dongha Lee, 2025
https://scholar.google.com/scholar?q=In+Their+Own+Words%3A+Reasoning+Traces+Tailored+for+Small+Models+Make+Them+Better+Reasoners
9. Rethinking Rubric Generation for Improving LLM Judge and Reward Modeling for Open-ended Tasks — William F. Shen, Xinchi Qiu, Chenxi Whitehouse, Lisa Alazraki, Shashwat Goel, Francesco Barbieri, Timon Willi, Akhil Mathur, Ilias Leontiadis, 2026
https://scholar.google.com/scholar?q=Rethinking+Rubric+Generation+for+Improving+LLM+Judge+and+Reward+Modeling+for+Open-ended+Tasks
10. Ask a Strong LLM Judge when Your Reward Model is Uncertain — Zhenghao Xu, Qin Lu, Qingru Zhang, Liang Qiu, Ilgee Hong, Changlong Yu, Wenlin Yao, Yao Liu, Haoming Jiang, Lihong Li, Hyokun Yun, Tuo Zhao, 2025
https://scholar.google.com/scholar?q=Ask+a+Strong+LLM+Judge+when+Your+Reward+Model+is+Uncertain
11. FOL-Traces: Verified First-Order Logic Reasoning Traces at Scale — Isabelle Lee, Sarah Liaw, Dani Yogatama, 2025
https://scholar.google.com/scholar?q=FOL-Traces%3A+Verified+First-Order+Logic+Reasoning+Traces+at+Scale
12. AI Post Transformers: ASI-Evolve for Data, Architectures, and RL — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-05-asi-evolve-for-data-architectures-and-rl-197b2b.mp3
13. AI Post Transformers: Distilling Multi-Agent Reasoning into a Single LLM — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-22-distilling-multi-agent-reasoning-into-a-143263.mp3
14. AI Post Transformers: Learning to Reason with 13 Parameters — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-14-learning-to-reason-with-13-parameters-54c87f.mp3
15. AI Post Transformers: MASA: Meta-Awareness via Self-Alignment Reinforcement Learning — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/masa-meta-awareness-via-self-alignment-reinforcement-learning/
16. AI Post Transformers: Kimi K2.5 and Visual Agent Swarms — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-24-kimi-k25-and-visual-agent-swarms-7d04d7.mp3
Interactive Visualization: Self-Improving Pretraining With Post-Trained Models

This episode explores selective classification in deep neural networks: adding a post-hoc reject option so a trained model can abstain when its confidence falls below a calibrated threshold. It explains the key concepts of coverage, selective risk, and the risk-coverage tradeoff, arguing that a model should be judged not just by how often it is right, but by how often it chooses to answer. The discussion centers on the paper’s SGR method, which uses a held-out calibration set to choose a threshold that keeps selective risk below a target with high probability under an i.i.d. assumption, and compares softmax response with MC-dropout as confidence scores. Listeners would find it interesting because it gets at a practical question in AI deployment: not whether a model is always confident, but whether it can reliably know when to defer.

Sources:
1. Selective Classification for Deep Neural Networks — Yonatan Geifman, Ran El-Yaniv, 2017
http://arxiv.org/abs/1705.08500
2. SelectiveNet: A Deep Neural Network with an Integrated Reject Option — Yonatan Geifman, Ran El-Yaniv, 2019
http://arxiv.org/abs/1901.09192
3. On Optimum Recognition Error and Reject Tradeoff — C. K. Chow, 1970
https://scholar.google.com/scholar?q=On+Optimum+Recognition+Error+and+Reject+Tradeoff
4. Selective Classification for Deep Neural Networks — Yonatan Geifman and Ran El-Yaniv, 2017
https://scholar.google.com/scholar?q=Selective+Classification+for+Deep+Neural+Networks
5. SelectiveNet: A Deep Neural Network with an Integrated Reject Option — Yonatan Geifman and Ran El-Yaniv, 2019
https://scholar.google.com/scholar?q=SelectiveNet%3A+A+Deep+Neural+Network+with+an+Integrated+Reject+Option
6. Selective Classification via One-Sided Prediction — Aditya Gangrade, Anil Kag, and Venkatesh Saligrama, 2021
https://scholar.google.com/scholar?q=Selective+Classification+via+One-Sided+Prediction
7. Classification with Reject Option — Radu Herbei and Marten H. Wegkamp, 2006
https://scholar.google.com/scholar?q=Classification+with+Reject+Option
8. Learning with Rejection — Corinna Cortes, Giulia DeSalvo, and Mehryar Mohri, 2016
https://scholar.google.com/scholar?q=Learning+with+Rejection
9. Boosting with Abstention — Corinna Cortes, Giulia DeSalvo, and Mehryar Mohri, 2016
https://scholar.google.com/scholar?q=Boosting+with+Abstention
10. Dropout as a Bayesian Approximation: Representing Model Uncertainty in Deep Learning — Yarin Gal and Zoubin Ghahramani, 2016
https://scholar.google.com/scholar?q=Dropout+as+a+Bayesian+Approximation%3A+Representing+Model+Uncertainty+in+Deep+Learning
11. On Calibration of Modern Neural Networks — Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger, 2017
https://scholar.google.com/scholar?q=On+Calibration+of+Modern+Neural+Networks
12. Simple and Scalable Predictive Uncertainty Estimation using Deep Ensembles — Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell, 2017
https://scholar.google.com/scholar?q=Simple+and+Scalable+Predictive+Uncertainty+Estimation+using+Deep+Ensembles
13. Can You Trust Your Model's Uncertainty? Evaluating Predictive Uncertainty Under Dataset Shift — Yaniv Ovadia, Emily Fertig, Jie Ren, Zachary Nado, D. Sculley, Sebastian Nowozin, Joshua Dillon, Balaji Lakshminarayanan, and Jasper Snoek, 2019
https://scholar.google.com/scholar?q=Can+You+Trust+Your+Model%27s+Uncertainty%3F+Evaluating+Predictive+Uncertainty+Under+Dataset+Shift
14. A Novel Characterization of the Population Area Under the Risk Coverage Curve (AURC) and Rates of Finite Sample Estimators — Han Zhou, Jordy Van Landeghem, Teodora Popordanoska, and Matthew B. Blaschko, 2025
https://scholar.google.com/scholar?q=A+Novel+Characterization+of+the+Population+Area+Under+the+Risk+Coverage+Curve+%28AURC%29+and+Rates+of+Finite+Sample+Estimators
15. On the Foundations of Noise-Free Selective Classification — Ran El-Yaniv and Yair Wiener, 2010
https://scholar.google.com/scholar?q=On+the+Foundations+of+Noise-Free+Selective+Classification
16. Deep Residual Learning for Image Recognition — Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun, 2016
https://scholar.google.com/scholar?q=Deep+Residual+Learning+for+Image+Recognition
17. Support Vector Machines with Embedded Reject Option — Giorgio Fumera and Fabio Roli, 2002
https://scholar.google.com/scholar?q=Support+Vector+Machines+with+Embedded+Reject+Option
18. A Deep Neural Network with an Integrated Reject Option — Yonatan Geifman and Ran El-Yaniv, 2019
https://scholar.google.com/scholar?q=A+Deep+Neural+Network+with+an+Integrated+Reject+Option
19. Augmenting the Softmax with Additional Confidence Scores for Improved Selective Classification with Out-of-Distribution Data — Guoxuan Xia, Christos-Savvas Bouganis, 2024
https://scholar.google.com/scholar?q=Augmenting+the+Softmax+with+Additional+Confidence+Scores+for+Improved+Selective+Classification+with+Out-of-Distribution+Data
20. GPify: Leveraging the Combined Strength of Normalizing Flow and Softmax For an Out-of-Distribution aware Confidence Score — Simon Kristoffersson Lind, Rudolph Triebel, Volker Kruger, 2026
https://scholar.google.com/scholar?q=GPify%3A+Leveraging+the+Combined+Strength+of+Normalizing+Flow+and+Softmax+For+an+Out-of-Distribution+aware+Confidence+Score
21. LogitAC: Logit Amplitude Constraints for Confidence Calibration and Out-of-Distribution Detection — Zongjing Cao, Yan Li, Byeong Seok Shin, 2024
https://scholar.google.com/scholar?q=LogitAC%3A+Logit+Amplitude+Constraints+for+Confidence+Calibration+and+Out-of-Distribution+Detection
22. Not all distributional shifts are equal: Fine-grained robust conformal inference — Jiahao Ai, Zhimei Ren, 2024
https://scholar.google.com/scholar?q=Not+all+distributional+shifts+are+equal%3A+Fine-grained+robust+conformal+inference
23. Wasserstein-regularized Conformal Prediction under General Distribution Shift — Rui Xu, Chao Chen, Yue Sun, Parvathinathan Venkitasubramaniam, Sihong Xie, 2025
https://scholar.google.com/scholar?q=Wasserstein-regularized+Conformal+Prediction+under+General+Distribution+Shift
Interactive Visualization: Selective Classification with Deep Neural Networks

This episode explores whether deep sequence models store knowledge as simple associative lookups or as geometric memories that encode broader relational structure. It discusses a recent paper arguing that, after memorizing graph facts in their weights, sequence models can answer multi-hop path queries as if they were making a much shorter move through embedding space, with learned representations resembling graph-embedding methods like node2vec and DeepWalk. The conversation highlights why that matters mechanistically: it suggests some forms of reasoning may be amortized into the model’s parameters during training rather than reconstructed step by step at inference time. Listeners would find it interesting for its sharp debate over what counts as real reasoning versus a clever shortcut, and for its caution about how far results from synthetic graph settings should generalize to large language models in the wild.

Interactive Visualization: Geometric Memory in Deep Sequence Models
Sources:
1. Deep sequence models tend to memorize geometrically; it is unclear why — Shahriar Noroozizadeh, Vaishnavh Nagarajan, Elan Rosenfeld, Sanjiv Kumar, 2025
http://arxiv.org/abs/2510.26745
2. DeepWalk: Online Learning of Social Representations — Bryan Perozzi, Rami Al-Rfou, Steven Skiena, 2014
https://scholar.google.com/scholar?q=DeepWalk%3A+Online+Learning+of+Social+Representations
3. node2vec: Scalable Feature Learning for Networks — Aditya Grover, Jure Leskovec, 2016
https://scholar.google.com/scholar?q=node2vec%3A+Scalable+Feature+Learning+for+Networks
4. Birth of a Transformer: A Memory Viewpoint — Alberto Bietti, Vivien Cabannes, Diane Bouchacourt, Herve Jegou, Leon Bottou, 2023
https://scholar.google.com/scholar?q=Birth+of+a+Transformer%3A+A+Memory+Viewpoint
5. Deep sequence models tend to memorize geometrically; it is unclear why — Shahriar Noroozizadeh, Vaishnavh Nagarajan, Elan Rosenfeld, Sanjiv Kumar, 2025
https://scholar.google.com/scholar?q=Deep+sequence+models+tend+to+memorize+geometrically%3B+it+is+unclear+why
6. The Pitfalls of Next-Token Prediction — Gregor Bachmann, Vaishnavh Nagarajan, 2024
https://scholar.google.com/scholar?q=The+Pitfalls+of+Next-Token+Prediction
7. How Transformers Learn to Plan via Multi-Token Prediction — Jianhao Huang, Zhanpeng Zhou, Renqiu Xia, Baharan Mirzasoleiman, Weijie Su, Wei Huang, 2026
https://scholar.google.com/scholar?q=How+Transformers+Learn+to+Plan+via+Multi-Token+Prediction
8. DeepSeek-V3 Technical Report — DeepSeek-AI and collaborators, 2024
https://scholar.google.com/scholar?q=DeepSeek-V3+Technical+Report
9. Language Models, Graph Searching, and Supervision Adulteration: When More Supervision is Less and How to Make More More — Arvid Frydenlund, 2025
https://scholar.google.com/scholar?q=Language+Models%2C+Graph+Searching%2C+and+Supervision+Adulteration%3A+When+More+Supervision+is+Less+and+How+to+Make+More+More
10. Hopping Too Late: Exploring the Limitations of Large Language Models on Multi-Hop Queries — Eden Biran, Daniela Gottesman, Sohee Yang, Mor Geva, and Amir Globerson, 2024
https://scholar.google.com/scholar?q=Hopping+Too+Late%3A+Exploring+the+Limitations+of+Large+Language+Models+on+Multi-Hop+Queries
11. The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A" — Lukas Berglund, Meg Tong, Maximilian Kaufmann, Mikita Balesni, Asa Cooper Stickland, Tomasz Korbak, and Owain Evans, 2024
https://scholar.google.com/scholar?q=The+Reversal+Curse%3A+LLMs+trained+on+%22A+is+B%22+fail+to+learn+%22B+is+A%22
12. In-Context Denoising with One-Layer Transformers: Connections between Attention and Associative Memory Retrieval — Matthew Smart, Alberto Bietti, Anirvan M. Sengupta, 2025
https://scholar.google.com/scholar?q=In-Context+Denoising+with+One-Layer+Transformers%3A+Connections+between+Attention+and+Associative+Memory+Retrieval
13. In-Context Learning as Conditioned Associative Memory Retrieval — Weimin Wu, Teng-Yun Hsiao, Jerry Yao-Chieh Hu, Wenxin Zhang, Han Liu, 2025
https://scholar.google.com/scholar?q=In-Context+Learning+as+Conditioned+Associative+Memory+Retrieval
14. Position-Aware Relational Transformer for Knowledge Graph Embedding — Guangyao Li, Zequn Sun, Wei Hu, Gong Cheng, et al., 2023
https://scholar.google.com/scholar?q=Position-Aware+Relational+Transformer+for+Knowledge+Graph+Embedding
15. Relational Transformer: Toward Zero-Shot Foundation Models for Relational Data — Rishabh Ranjan, Valter Hudovernik, Mark Znidar, Charilaos Kanatsoulis, Roshan Upendra, Mahmoud Mohammadi, Joe Meyer, Tom Palczewski, Carlos Guestrin, Jure Leskovec, 2025
https://scholar.google.com/scholar?q=Relational+Transformer%3A+Toward+Zero-Shot+Foundation+Models+for+Relational+Data
16. AI Post Transformers: In-Place Test-Time Training for Transformers — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-09-in-place-test-time-training-for-transfor-d0b976.mp3
17. AI Post Transformers: Mamba-3 for Efficient Sequence Modeling — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-16-mamba-3-for-efficient-sequence-modeling-97a22a.mp3
Interactive Visualization: Geometric Memory in Deep Sequence Models

This episode explores whether large language models can genuinely recognize the limits of their own knowledge or whether they have simply learned to sound uncertain in socially acceptable ways. It examines the paper’s idea of “self-knowledge” through the lens of confidence calibration, including the dangerous case where a model does not know an answer but responds with unwarranted confidence. The discussion walks through the SelfAware benchmark, explaining how it pairs unanswerable questions with semantically similar answerable ones and why that design is both insightful and methodologically slippery. Listeners would find it interesting because it gets past simple accuracy scores and asks a more consequential question for AI safety and product reliability: when a model says “I don’t know,” is that real judgment or just polished behavior?

Interactive Visualization: Do Language Models Know Their Limits
Sources:
1. Do Large Language Models Know What They Don't Know? — Zhangyue Yin, Qiushi Sun, Qipeng Guo, Jiawen Wu, Xipeng Qiu, Xuanjing Huang, 2023
http://arxiv.org/abs/2305.18153
2. Finetuned Language Models Are Zero-Shot Learners — Jason Wei, Maarten Bosma, Vincent Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew Dai, Quoc V. Le, 2021
https://scholar.google.com/scholar?q=Finetuned+Language+Models+Are+Zero-Shot+Learners
3. Multitask Prompted Training Enables Zero-Shot Task Generalization — Victor Sanh, Albert Webson, Colin Raffel and many coauthors, 2021
https://scholar.google.com/scholar?q=Multitask+Prompted+Training+Enables+Zero-Shot+Task+Generalization
4. Training Language Models to Follow Instructions with Human Feedback — Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida and many coauthors, 2022
https://scholar.google.com/scholar?q=Training+Language+Models+to+Follow+Instructions+with+Human+Feedback
5. Scaling Instruction-Finetuned Language Models — Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, Xuezhi Wang, Denny Zhou, Quoc V. Le, Jason Wei and many coauthors, 2022
https://scholar.google.com/scholar?q=Scaling+Instruction-Finetuned+Language+Models
6. Self-Instruct: Aligning Language Models with Self-Generated Instructions — Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, Hannaneh Hajishirzi, 2023
https://scholar.google.com/scholar?q=Self-Instruct%3A+Aligning+Language+Models+with+Self-Generated+Instructions
7. Measuring and Improving Factuality in Large Language Models with Calibrated Confidence Scores — Saurav Kadavath, Eric Wallace, Luyu Gao, et al., 2022
https://scholar.google.com/scholar?q=Measuring+and+Improving+Factuality+in+Large+Language+Models+with+Calibrated+Confidence+Scores
8. Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models — Aarohi Srivastava, Jos Rozen, Francesco Tintarev, et al., 2022
https://scholar.google.com/scholar?q=Beyond+the+Imitation+Game%3A+Quantifying+and+extrapolating+the+capabilities+of+language+models
9. Teaching Small Language Models to Reason — Jason Wei, Xuezhi Wang, Dale Schuurmans, et al., 2022
https://scholar.google.com/scholar?q=Teaching+Small+Language+Models+to+Reason
10. Self-Consistency Improves Chain of Thought Reasoning in Language Models — Xuezhi Wang, Jason Wei, Dale Schuurmans, et al., 2022
https://scholar.google.com/scholar?q=Self-Consistency+Improves+Chain+of+Thought+Reasoning+in+Language+Models
11. SimCSE: Simple Contrastive Learning of Sentence Embeddings — Tianyu Gao, Xingcheng Yao, Danqi Chen, 2021
https://scholar.google.com/scholar?q=SimCSE%3A+Simple+Contrastive+Learning+of+Sentence+Embeddings
12. SQuAD 2.0: The Stanford Question Answering Dataset — Pranav Rajpurkar, Robin Jia, Percy Liang, 2018
https://scholar.google.com/scholar?q=SQuAD+2.0%3A+The+Stanford+Question+Answering+Dataset
13. Uncertainty Distillation: Teaching Language Models to Express Semantic Confidence — Sophia Hager et al., 2025
https://scholar.google.com/scholar?q=Uncertainty+Distillation%3A+Teaching+Language+Models+to+Express+Semantic+Confidence
14. Large Language Model Uncertainty Measurement and Calibration for Medical Diagnosis and Treatment — Thomas Savage et al., 2024
https://scholar.google.com/scholar?q=Large+Language+Model+Uncertainty+Measurement+and+Calibration+for+Medical+Diagnosis+and+Treatment
15. Uncertainty Quantification and Confidence Calibration in Large Language Models: A Survey — Xiaoou Liu et al., 2025
https://scholar.google.com/scholar?q=Uncertainty+Quantification+and+Confidence+Calibration+in+Large+Language+Models%3A+A+Survey
16. Unanswerability Evaluation for Retrieval Augmented Generation — Xiangyu Peng, Prafulla Kumar Choubey, Caiming Xiong, Chien-Sheng Wu, 2025
https://scholar.google.com/scholar?q=Unanswerability+Evaluation+for+Retrieval+Augmented+Generation
17. Answerability in Retrieval-Augmented Open-Domain Question Answering — Rustam Abdumalikov, Pasquale Minervini, Yova Kementchedjhieva, 2024
https://scholar.google.com/scholar?q=Answerability+in+Retrieval-Augmented+Open-Domain+Question+Answering
18. The Art of Saying No: Contextual Noncompliance in Language Models — Faeze Brahman et al., 2024
https://scholar.google.com/scholar?q=The+Art+of+Saying+No%3A+Contextual+Noncompliance+in+Language+Models
19. AI Post Transformers: Hallucination to Truth: A Review of Fact-Checking and Factuality Evaluation in Large Language Model — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/hallucination-to-truth-a-review-of-fact-checking-and-factuality-evaluation-in-la/
20. AI Post Transformers: Internal Safety Collapse in Frontier LLMs — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-internal-safety-collapse-in-frontier-llm-8be72f.mp3
21. AI Post Transformers: Self-Search Reinforcement Learning for LLMs — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/self-search-reinforcement-learning-for-llms/
22. AI Post Transformers: Real Context Size and Context Rot — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-07-real-context-size-and-context-rot-56cbb4.mp3
Interactive Visualization: Do Language Models Know Their Limits

This episode explores a speech recognition paper that replaces slow left-to-right transcript generation with a faster draft-and-edit approach, where a speech model produces an initial hypothesis and a bidirectional LLM corrects it in parallel. It explains the tradeoff between CTC-based systems, which are fast but weaker at using linguistic context, and autoregressive decoders, which are more expressive but too slow for low-latency use cases like captioning and meetings. The discussion highlights the paper’s key ideas, including transcript editing with insertion slots, latent alignment inspired by CTC, and the use of LoRA to adapt pretrained language models efficiently. Listeners would find it interesting because it shows a concrete path to pushing ASR onto a better speed-accuracy frontier, with reported gains such as a 27x speedup over an autoregressive baseline while staying competitive on word error rate.

Sources:
1. NLE: Non-autoregressive LLM-based ASR by Transcript Editing — Avihu Dekel, Samuel Thomas, Takashi Fukada, George Saon, 2026
http://arxiv.org/abs/2603.08397
2. Connectionist Temporal Classification: Labelling Unsegmented Sequence Data with Recurrent Neural Networks — Alex Graves, Santiago Fernandez, Faustino Gomez, Jurgen Schmidhuber, 2006
https://scholar.google.com/scholar?q=Connectionist+Temporal+Classification%3A+Labelling+Unsegmented+Sequence+Data+with+Recurrent+Neural+Networks
3. Listen, Attend and Spell — William Chan, Navdeep Jaitly, Quoc V. Le, Oriol Vinyals, 2015
https://scholar.google.com/scholar?q=Listen%2C+Attend+and+Spell
4. Exploring Architectures, Data and Units for Streaming End-to-End Speech Recognition with RNN-Transducer — Hasim Sak, Kanishka Rao, Rohit Prabhavalkar, 2017
https://scholar.google.com/scholar?q=Exploring+Architectures%2C+Data+and+Units+for+Streaming+End-to-End+Speech+Recognition+with+RNN-Transducer
5. Robust Speech Recognition via Large-Scale Weak Supervision — Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, Ilya Sutskever, 2022
https://scholar.google.com/scholar?q=Robust+Speech+Recognition+via+Large-Scale+Weak+Supervision
6. Mask CTC: Non-Autoregressive End-to-End ASR with CTC and Mask Predict — Yosuke Higuchi, Shinji Watanabe, Nanxin Chen, Tetsuji Ogawa, Tetsunori Kobayashi, 2020
https://scholar.google.com/scholar?q=Mask+CTC%3A+Non-Autoregressive+End-to-End+ASR+with+CTC+and+Mask+Predict
7. Align-Refine: Non-autoregressive Speech Recognition via Iterative Realignment — Ethan A. Chi, Julian Salazar, Katrin Kirchhoff, 2021
https://scholar.google.com/scholar?q=Align-Refine%3A+Non-autoregressive+Speech+Recognition+via+Iterative+Realignment
8. A Survey on Non-Autoregressive Generation for Neural Machine Translation and Beyond — Yisheng Xiao, Lijun Wu, Junliang Guo, Juntao Li, Min Zhang, Tao Qin, Tie-Yan Liu, 2022
https://scholar.google.com/scholar?q=A+Survey+on+Non-Autoregressive+Generation+for+Neural+Machine+Translation+and+Beyond
9. Encode, Tag, Realize: High-Precision Text Editing — Eric Malmi, Sebastian Krause, Sascha Rothe, Daniil Mirylenka, Aliaksei Severyn, 2019
https://scholar.google.com/scholar?q=Encode%2C+Tag%2C+Realize%3A+High-Precision+Text+Editing
10. Levenshtein Transformer — Jiatao Gu, Changhan Wang, Junbo Zhao, 2019
https://scholar.google.com/scholar?q=Levenshtein+Transformer
11. GECToR - Grammatical Error Correction: Tag, Not Rewrite — Kostiantyn Omelianchuk, Vitaliy Atrasevych, Artem Chernodub, Oleksandr Skurzhanskyi, 2020
https://scholar.google.com/scholar?q=GECToR+-+Grammatical+Error+Correction%3A+Tag%2C+Not+Rewrite
12. FELIX: Flexible Text Editing Through Tagging and Insertion — Jonathan Mallinson, Aliaksei Severyn, Eric Malmi, Guillermo Garrido, 2020
https://scholar.google.com/scholar?q=FELIX%3A+Flexible+Text+Editing+Through+Tagging+and+Insertion
13. LoRA: Low-Rank Adaptation of Large Language Models — Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, 2021
https://scholar.google.com/scholar?q=LoRA%3A+Low-Rank+Adaptation+of+Large+Language+Models
14. Token-and-Duration Transducer — Dmitriy Genzel, Yanzhang He, et al., 2023
https://scholar.google.com/scholar?q=Token-and-Duration+Transducer
15. Mask-Predict: Parallel Decoding of Conditional Masked Language Models — Marjan Ghazvininejad, Omer Levy, Yinhan Liu, Luke Zettlemoyer, 2019
https://scholar.google.com/scholar?q=Mask-Predict%3A+Parallel+Decoding+of+Conditional+Masked+Language+Models
16. QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache — Rishabh Tiwari, Haocheng Xi, Aditya Tomar, et al., 2025
https://scholar.google.com/scholar?q=QuantSpec%3A+Self-Speculative+Decoding+with+Hierarchical+Quantized+KV+Cache
17. SpeCache: Speculative Key-Value Caching for Efficient Generation of LLMs — Shibo Jie, Yehui Tang, Kai Han, Zhi-Hong Deng, Jing Han, 2025
https://scholar.google.com/scholar?q=SpeCache%3A+Speculative+Key-Value+Caching+for+Efficient+Generation+of+LLMs
18. SoftCorrect: Error Correction with Soft Detection for Automatic Speech Recognition — Yichong Leng, Xu Tan, Wenjie Liu, Kaitao Song, Rui Wang, Xiang-Yang Li, Tao Qin, Edward Lin, Tie-Yan Liu, 2022
https://scholar.google.com/scholar?q=SoftCorrect%3A+Error+Correction+with+Soft+Detection+for+Automatic+Speech+Recognition
19. Semantic Segmentation with Bidirectional Language Models Improves Long-form ASR — W. Ronny Huang, Hao Zhang, Shankar Kumar, Shuo-Yiin Chang, Tara N. Sainath, 2023
https://scholar.google.com/scholar?q=Semantic+Segmentation+with+Bidirectional+Language+Models+Improves+Long-form+ASR
20. Bi-directional Context-Enhanced Speech Large Language Models for Multilingual Conversational ASR — Yizhou Peng, Hexin Liu, Eng Siong Chng, 2025
https://scholar.google.com/scholar?q=Bi-directional+Context-Enhanced+Speech+Large+Language+Models+for+Multilingual+Conversational+ASR
21. AI Post Transformers: Qwen3.5-Omni Thinker-Talker for Omnimodal Streaming — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-21-qwen35-omni-thinker-talker-for-omnimodal-36b26c.mp3
22. AI Post Transformers: Batch-Aware Expert Routing for Faster MoE Decoding — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-batch-aware-expert-routing-for-faster-mo-683ab6.mp3
Interactive Visualization: Fast Speech Recognition by Transcript Editing

This episode explores CacheFlow, a systems approach to speeding up long-context LLM serving by restoring transformer KV caches more intelligently. It explains the tradeoffs between recomputing prior attention state, loading it from storage, or combining both, and argues that the real user-facing bottleneck is now time-to-first-token rather than raw generation speed. The discussion focuses on CacheFlow’s main idea: a batch-aware scheduler that splits restoration across recomputation and I/O at token, layer, and GPU levels to reduce wasted work under contention. Listeners would find it interesting because it shows how practical transformer serving is increasingly shaped by runtime scheduling, cache movement, and latency engineering rather than new model architectures.

Sources:
1. CacheFlow and 3D-Parallel KV Cache Restoration
https://arxiv.org/pdf/2604.25080
2. Efficient Memory Management for Large Language Model Serving with PagedAttention — Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, Ion Stoica, 2023
https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention
3. Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention — Bin Gao, Zhuomin He, Puru Sharma, Qingxuan Kang, Djordje Jevdjic, Junbo Deng, Xingkun Yang, Zhou Yu, Pengfei Zuo, 2024
https://scholar.google.com/scholar?q=Cost-Efficient+Large+Language+Model+Serving+for+Multi-turn+Conversations+with+CachedAttention
4. Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving — Ruoyu Qin, Zheming Li, Weiran He, Mingxing Zhang, Yongwei Wu, Weimin Zheng, Xinran Xu, 2024
https://scholar.google.com/scholar?q=Mooncake%3A+A+KVCache-centric+Disaggregated+Architecture+for+LLM+Serving
5. Fast State Restoration in LLM Serving with HCache — Shiwei Gao, Youmin Chen, Jiwu Shu, 2024
https://scholar.google.com/scholar?q=Fast+State+Restoration+in+LLM+Serving+with+HCache
6. LMCache: An Efficient KV Cache Layer for Enterprise-Scale LLM Inference — Yihua Cheng, Yuhan Liu, Jiayi Yao, Yuwei An, Xiaokun Chen, Shaoting Feng, Yuyang Huang, Samuel Shen, Kuntai Du, Junchen Jiang, 2025
https://scholar.google.com/scholar?q=LMCache%3A+An+Efficient+KV+Cache+Layer+for+Enterprise-Scale+LLM+Inference
7. DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving — Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, Hao Zhang, 2024
https://scholar.google.com/scholar?q=DistServe%3A+Disaggregating+Prefill+and+Decoding+for+Goodput-optimized+Large+Language+Model+Serving
8. Mooncake: Trading More Storage for Less Computation — A KVCache-centric Architecture for Serving LLM Chatbot — Ruoyu Qin, Zheming Li, Weiran He, Jialei Cui, Feng Ren, Mingxing Zhang, Yongwei Wu, Weimin Zheng, Xinran Xu, 2025
https://scholar.google.com/scholar?q=Mooncake%3A+Trading+More+Storage+for+Less+Computation+%E2%80%94+A+KVCache-centric+Architecture+for+Serving+LLM+Chatbot
9. DéjàVu: KV-cache Streaming for Fast, Fault-tolerant Generative LLM Serving — Foteini Strati, Sara Mcallister, Amar Phanishayee, Jakub Tarnawski, Ana Klimovic, 2024
https://scholar.google.com/scholar?q=D%C3%A9j%C3%A0Vu%3A+KV-cache+Streaming+for+Fast%2C+Fault-tolerant+Generative+LLM+Serving
10. KV Prediction for Improved Time to First Token — Maxwell Horton, Qingqing Cao, Chenfan Sun, Yanzi Jin, Sachin Mehta, Mohammad Rastegari, Moin Nabi, 2025
https://scholar.google.com/scholar?q=KV+Prediction+for+Improved+Time+to+First+Token
11. KVSharer: Efficient Inference via Layer-Wise Dissimilar KV Cache Sharing — Yifei Yang et al., 2024
https://scholar.google.com/scholar?q=KVSharer%3A+Efficient+Inference+via+Layer-Wise+Dissimilar+KV+Cache+Sharing
12. CommonKV: Compressing KV Cache with Cross-layer Parameter Sharing — Yixuan Wang et al., 2025
https://scholar.google.com/scholar?q=CommonKV%3A+Compressing+KV+Cache+with+Cross-layer+Parameter+Sharing
13. Lossless KV Cache Compression to 2% — Zhen Yang et al., 2024
https://scholar.google.com/scholar?q=Lossless+KV+Cache+Compression+to+2%25
14. XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference — Weizhuo Li et al., 2024
https://scholar.google.com/scholar?q=XKV%3A+Personalized+KV+Cache+Memory+Reduction+for+Long-Context+LLM+Inference
15. TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding — Hanshi Sun et al., 2024
https://scholar.google.com/scholar?q=TriForce%3A+Lossless+Acceleration+of+Long+Sequence+Generation+with+Hierarchical+Speculative+Decoding
16. MagicDec: Breaking the Latency-Throughput Tradeoff for Long Context Generation with Speculative Decoding — Jian Chen, Vashisth Tiwari, Ranajoy Sadhukhan, Zhuoming Chen, Jinyuan Shi, Ian En-Hsu Yen, Beidi Chen, 2024
https://scholar.google.com/scholar?q=MagicDec%3A+Breaking+the+Latency-Throughput+Tradeoff+for+Long+Context+Generation+with+Speculative+Decoding
17. SpeCache: Speculative Key-Value Caching for Efficient Generation of LLMs — Shibo Jie et al., 2025
https://scholar.google.com/scholar?q=SpeCache%3A+Speculative+Key-Value+Caching+for+Efficient+Generation+of+LLMs
18. Prefill-Decode Aggregation or Disaggregation? Unifying Both for Goodput-Optimized LLM Serving — Chao Wang, Pengfei Zuo, Zhangyu Chen, Yunkai Liang, Zhou Yu, Ming-Chang Yang, 2025
https://scholar.google.com/scholar?q=Prefill-Decode+Aggregation+or+Disaggregation%3F+Unifying+Both+for+Goodput-Optimized+LLM+Serving
19. AI Post Transformers: ContiguousKV for Faster LLM Prefill KV Reuse — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-20-contiguouskv-for-faster-llm-prefill-kv-r-59f545.mp3
20. AI Post Transformers: ScoutAttention for Efficient KV Cache Offloading — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-24-scoutattention-for-efficient-kv-cache-of-b26699.mp3
21. AI Post Transformers: Prefill-as-a-Service for Cross-Datacenter KV Cache — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-19-prefill-as-a-service-for-cross-datacente-7560be.mp3
22. AI Post Transformers: FengHuang for Rack-Scale LLM Inference Memory — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-12-fenghuang-for-rack-scale-llm-inference-m-62708e.mp3
23. AI Post Transformers: KV Cache TTL for Multi-Turn Agent Scheduling — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-09-kv-cache-ttl-for-multi-turn-agent-schedu-996bf1.mp3
24. AI Post Transformers: Splitwise: Phase-Split LLM Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-26-splitwise-phase-split-llm-inference-e8945b.mp3
25. AI Post Transformers: KVSwap for Disk-Aware Long-Context On-Device Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-16-kvswap-for-disk-aware-long-context-on-de-f3c15e.mp3
Interactive Visualization: CacheFlow and 3D-Parallel KV Cache Restoration

This episode explores DeltaKV, a method for reducing the huge GPU memory burden of KV caches in long-context language model inference without simply discarding old tokens. It contrasts three strategies for handling long contexts: token eviction, dynamic sparse attention, and true compression, arguing that the cache contains structured redundancy that can be exploited rather than treated as disposable overhead. The discussion highlights DeltaKV’s core idea of keeping a small uncompressed reference set and storing other cache entries as compressed residuals relative to similar past states, drawing an analogy to delta encoding or version control. Listeners would find it interesting because it connects transformer internals, systems constraints, and practical serving performance, including claims of cutting memory to 29 percent of baseline and reaching up to 2x throughput with supporting infrastructure like Sparse-vLLM.

Sources:
1. DeltaKV: Residual-Based KV Cache Compression via Long-Range Similarity — Jitai Hao, Qiang Huang, Yaowei Wang, Min Zhang, Jun Yu, 2026
http://arxiv.org/abs/2602.08005
2. Generalization through Memorization: Nearest Neighbor Language Models — Urvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer, Mike Lewis, 2019
https://scholar.google.com/scholar?q=Generalization+through+Memorization%3A+Nearest+Neighbor+Language+Models
3. Reformer: The Efficient Transformer — Nikita Kitaev, Lukasz Kaiser, Anselm Levskaya, 2020
https://scholar.google.com/scholar?q=Reformer%3A+The+Efficient+Transformer
4. Improving Language Models by Retrieving from Trillions of Tokens — Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann and many others, 2021
https://scholar.google.com/scholar?q=Improving+Language+Models+by+Retrieving+from+Trillions+of+Tokens
5. Memorizing Transformers — Yuhuai Wu, Markus N. Rabe, DeLesley Hutchins, Christian Szegedy, 2022
https://scholar.google.com/scholar?q=Memorizing+Transformers
6. CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving — Yuhan Liu, Hanchen Li, Yihua Cheng and others, 2023
https://scholar.google.com/scholar?q=CacheGen%3A+KV+Cache+Compression+and+Streaming+for+Fast+Large+Language+Model+Serving
7. Palu: Compressing KV-Cache with Low-Rank Projection — Chi-Chih Chang, Wei-Cheng Lin, Chien-Yu Lin and others, 2024
https://scholar.google.com/scholar?q=Palu%3A+Compressing+KV-Cache+with+Low-Rank+Projection
8. Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries — Junhyuck Kim, Jongho Park, Jaewoong Cho, Dimitris Papailiopoulos, 2025
https://scholar.google.com/scholar?q=Lexico%3A+Extreme+KV+Cache+Compression+via+Sparse+Coding+over+Universal+Dictionaries
9. Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models — Alina Shutova, Vladimir Malinovskii, Vage Egiazarian and others, 2025
https://scholar.google.com/scholar?q=Cache+Me+If+You+Must%3A+Adaptive+Key-Value+Quantization+for+Large+Language+Models
10. OmniKV: Dynamic Context Selection for Efficient Long-Context LLMs — Jitai Hao, Yuke Zhu, Tian Wang, Jun Yu, Xin Xin, Bo Zheng, Zhaochun Ren, Sheng Guo, 2025
https://scholar.google.com/scholar?q=OmniKV%3A+Dynamic+Context+Selection+for+Efficient+Long-Context+LLMs
11. QUEST: Query-Aware Sparsity for Efficient Long-Context LLM Inference — Jiaming Tang, Yilong Zhao, Kan Zhu, Guangxuan Xiao, Baris Kasikci, Song Han, 2024
https://scholar.google.com/scholar?q=QUEST%3A+Query-Aware+Sparsity+for+Efficient+Long-Context+LLM+Inference
12. Palu: KV-Cache Compression with Low-Rank Projection — Chi-Chih Chang, Wei-Cheng Lin, Chien-Yu Lin, Chong-Yan Chen, Yu-Fang Hu, Pei-Shuo Wang, Ning-Chi Huang, Luis Ceze, Mohamed S. Abdelfattah, Kai-Chiang Wu, 2025
https://scholar.google.com/scholar?q=Palu%3A+KV-Cache+Compression+with+Low-Rank+Projection
13. SCBench: A KV Cache-Centric Analysis of Long-Context Methods — Yucheng Li, Huiqiang Jiang, Qianhui Wu, Xufang Luo, Surin Ahn, Chengruidong Zhang, Amir H. Abdi, Dongsheng Li, Jianfeng Gao, Yuqing Yang, Lili Qiu, 2025
https://scholar.google.com/scholar?q=SCBench%3A+A+KV+Cache-Centric+Analysis+of+Long-Context+Methods
14. The Pitfalls of KV Cache Compression — Alex Chen, Renato Geh, Aditya Grover, Guy Van den Broeck, Daniel Israel, 2025
https://scholar.google.com/scholar?q=The+Pitfalls+of+KV+Cache+Compression
15. KVzip: Query-Agnostic KV Cache Compression with Context Reconstruction — Jang-Hyun Kim, Jinuk Kim, Sangwoo Kwon, Jae W. Lee, Sangdoo Yun, Hyun Oh Song, 2025
https://scholar.google.com/scholar?q=KVzip%3A+Query-Agnostic+KV+Cache+Compression+with+Context+Reconstruction
16. Kvlink: Accelerating Large Language Models via Efficient KV Cache Reuse — approx. recent LLM systems authors, 2025/2026
https://scholar.google.com/scholar?q=Kvlink%3A+Accelerating+Large+Language+Models+via+Efficient+KV+Cache+Reuse
17. HyperRAG: Enhancing Quality-Efficiency Tradeoffs in Retrieval-Augmented Generation with Reranker KV-Cache Reuse — approx. recent RAG systems authors, 2025/2026
https://scholar.google.com/scholar?q=HyperRAG%3A+Enhancing+Quality-Efficiency+Tradeoffs+in+Retrieval-Augmented+Generation+with+Reranker+KV-Cache+Reuse
18. ProphetKV: User-Query-Driven Selective Recomputation for Efficient KV Cache Reuse in Retrieval-Augmented Generation — approx. recent RAG/serving authors, 2025/2026
https://scholar.google.com/scholar?q=ProphetKV%3A+User-Query-Driven+Selective+Recomputation+for+Efficient+KV+Cache+Reuse+in+Retrieval-Augmented+Generation
19. KVSink: Understanding and Enhancing the Preservation of Attention Sinks in KV Cache Quantization for LLMs — approx. recent KV quantization authors, 2025/2026
https://scholar.google.com/scholar?q=KVSink%3A+Understanding+and+Enhancing+the+Preservation+of+Attention+Sinks+in+KV+Cache+Quantization+for+LLMs
20. LLMs Know What to Drop: Self-Attention Guided KV Cache Eviction for Efficient Long-Context Inference — approx. recent long-context inference authors, 2025/2026
https://scholar.google.com/scholar?q=LLMs+Know+What+to+Drop%3A+Self-Attention+Guided+KV+Cache+Eviction+for+Efficient+Long-Context+Inference
21. Task-KV: Task-Aware KV Cache Optimization via Semantic Differentiation of Attention Heads — approx. recent attention/KV optimization authors, 2025/2026
https://scholar.google.com/scholar?q=Task-KV%3A+Task-Aware+KV+Cache+Optimization+via+Semantic+Differentiation+of+Attention+Heads
22. WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More — approx. recent quantization authors, 2025/2026
https://scholar.google.com/scholar?q=WKVQuant%3A+Quantizing+Weight+and+Key%2FValue+Cache+for+Large+Language+Models+Gains+More
23. AlignedKV: Reducing Memory Access of KV-Cache with Precision-Aligned Quantization — approx. recent KV systems/quantization authors, 2025/2026
https://scholar.google.com/scholar?q=AlignedKV%3A+Reducing+Memory+Access+of+KV-Cache+with+Precision-Aligned+Quantization
24. Making Every Head Count: Sparse Attention Without the Speed-Performance Trade-off — approx. recent sparse attention authors, 2025/2026
https://scholar.google.com/scholar?q=Making+Every+Head+Count%3A+Sparse+Attention+Without+the+Speed-Performance+Trade-off
25. Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention — approx. recent sparse systems authors, 2025/2026
https://scholar.google.com/scholar?q=Native+Sparse+Attention%3A+Hardware-Aligned+and+Natively+Trainable+Sparse+Attention
26. AI Post Transformers: Stochastic KV Routing for Cache Sharing — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-29-stochastic-kv-routing-for-cache-sharing-5fef63.mp3
27. AI Post Transformers: ContiguousKV for Faster LLM Prefill KV Reuse — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-20-contiguouskv-for-faster-llm-prefill-kv-r-59f545.mp3
28. AI Post Transformers: KVSwap for Disk-Aware Long-Context On-Device Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-16-kvswap-for-disk-aware-long-context-on-de-f3c15e.mp3
29. AI Post Transformers: Native Sparse Attention: Efficient Long-Context LLMs — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/native-sparse-attention-efficient-long-context-llms/
30. AI Post Transformers: Memory Sparse Attention for 100M-Token Scaling — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-07-memory-sparse-attention-for-100m-token-s-377cff.mp3
31. AI Post Transformers: AWQ: On-Device LLM Compression and Acceleration — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/awq-on-device-llm-compression-and-acceleration/

This episode explores a review of mechanistic interpretability for transformer language models, focusing on how researchers study internal features, circuits, and claims of universality across models. It explains the core toolkit behind the field, including linear probes, hidden-state analysis, intervention methods, vocabulary projection, and sparse autoencoders, while grounding those ideas in transformer anatomy such as attention heads, MLPs, and the residual stream. The discussion highlights a central tension in the literature: finding information encoded in activations is not the same as proving that information causally drives model behavior, and the episode repeatedly questions where interpretability claims may be overstated. Listeners would find it interesting because it offers a concrete map of a fast-growing area of AI research while also giving a careful critique of the field’s assumptions, evidence, and real-world usefulness.

Sources:
1. A Practical Review of Mechanistic Interpretability for Transformer-Based Language Models — Daking Rai, Yilun Zhou, Shi Feng, Abulhair Saparov, Ziyu Yao, 2024
http://arxiv.org/abs/2407.02646
2. Towards Monosemanticity: Decomposing Language Models With Dictionary Learning — Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Chris Olah, and collaborators, 2023
https://scholar.google.com/scholar?q=Towards+Monosemanticity%3A+Decomposing+Language+Models+With+Dictionary+Learning
3. Sparse Autoencoders Find Highly Interpretable Features in Language Models — Hoagy Cunningham, Aidan Ewart, Logan Riggs, Robert Huben, Lee Sharkey, 2023
https://scholar.google.com/scholar?q=Sparse+Autoencoders+Find+Highly+Interpretable+Features+in+Language+Models
4. Scaling and Evaluating Sparse Autoencoders — Leo Gao, Tom Dupré la Tour, Henk Tillman, Gabriel Goh, Rajan Troll, Alec Radford, Ilya Sutskever, Jan Leike, Jeffrey Wu, 2024
https://scholar.google.com/scholar?q=Scaling+and+Evaluating+Sparse+Autoencoders
5. Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse Autoencoders — Senthooran Rajamanoharan, Tom Lieberum, Nicolas Sonnerat, Arthur Conmy, Vikrant Varma, Janos Kramar, Neel Nanda, 2024
https://scholar.google.com/scholar?q=Jumping+Ahead%3A+Improving+Reconstruction+Fidelity+with+JumpReLU+Sparse+Autoencoders
6. A Mathematical Framework for Transformer Circuits — Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, Chris Olah, 2021
https://scholar.google.com/scholar?q=A+Mathematical+Framework+for+Transformer+Circuits
7. Towards Best Practices of Activation Patching in Language Models: Metrics and Methods — Fred Zhang, Neel Nanda, 2024
https://scholar.google.com/scholar?q=Towards+Best+Practices+of+Activation+Patching+in+Language+Models%3A+Metrics+and+Methods
8. Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet — Adly Templeton, Tom Conerly, Jonathan Marcus, Jack Lindsey, Trenton Bricken, Brian Chen, Adam Pearce, Craig Citro, Emmanuel Ameisen, Andy Jones, Hoagy Cunningham, Nicholas L. Turner, Callum McDougall, Monte MacDiarmid, C. Daniel Freeman, Theodore R. Sumers, Edward Rees, Joshua Batson, Adam Jermyn, Shan Carter, Chris Olah, Tom Henighan, 2024
https://scholar.google.com/scholar?q=Scaling+Monosemanticity%3A+Extracting+Interpretable+Features+from+Claude+3+Sonnet
9. Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models — Samuel Marks, Can Rager, Eric J. Michaud, Yonatan Belinkov, David Bau, Aaron Mueller, 2025
https://scholar.google.com/scholar?q=Sparse+Feature+Circuits%3A+Discovering+and+Editing+Interpretable+Causal+Graphs+in+Language+Models
10. On the Theoretical Understanding of Identifiable Sparse Autoencoders and Beyond — Jingyi Cui, Qi Zhang, Yifei Wang, Yisen Wang, 2025
https://scholar.google.com/scholar?q=On+the+Theoretical+Understanding+of+Identifiable+Sparse+Autoencoders+and+Beyond
11. Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words — Gouki Minegishi, Hiroki Furuta, Yusuke Iwasawa, Yutaka Matsuo, 2025
https://scholar.google.com/scholar?q=Rethinking+Evaluation+of+Sparse+Autoencoders+through+the+Representation+of+Polysemous+Words
12. Causal Head Gating: A Framework for Interpreting Roles of Attention Heads in Transformers — Andrew Nam, Henry Conklin, Yukang Yang, Thomas Griffiths, Jonathan Cohen, Sarah-Jane Leslie, 2025
https://scholar.google.com/scholar?q=Causal+Head+Gating%3A+A+Framework+for+Interpreting+Roles+of+Attention+Heads+in+Transformers
13. Quantifying LLM Attention-Head Stability: Implications for Circuit Universality — Karan Bali, Jack Stanley, Praneet Suresh, Danilo Bzdok, 2026
https://scholar.google.com/scholar?q=Quantifying+LLM+Attention-Head+Stability%3A+Implications+for+Circuit+Universality
14. AI Post Transformers: Mechanistic interpretability: Decoding the AI's Inner Logic: Circuits and Sparse Features — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/mechanistic-interpretability-decoding-the-ais-inner-logic-circuits-and-sparse-fe/
15. AI Post Transformers: Advancing Mechanistic Interpretability with Sparse Autoencoders — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/advancing-mechanistic-interpretability-with-sparse-autoencoders/
16. AI Post Transformers: Neural Chameleons and Evading Activation Monitors — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-14-neural-chameleons-and-evading-activation-bc470e.mp3
17. AI Post Transformers: Linear Classifier Probes for Intermediate Layers — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-16-linear-classifier-probes-for-intermediat-927ae3.mp3
18. AI Post Transformers: Internal Safety Collapse in Frontier LLMs — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-internal-safety-collapse-in-frontier-llm-8be72f.mp3

This episode explores a 2026 paper on recursive multi-agent systems that asks whether AI collaboration can scale better by replacing text-based agent communication with shared latent-state updates. It explains the difference between agent topology, assigned roles, and communication channels, and argues that natural language may be a costly bottleneck compared with the richer internal representations models use during reasoning. The discussion connects this idea to earlier multi-agent chat frameworks, classic transformer architectures, and recent test-time compute work on latent recurrent refinement. Listeners would find it interesting because it frames a sharp debate between today’s practical, inspectable text-first agent workflows and a more trainable, neural-network-like approach that could change how complex AI teams are built.

Sources:
1. Recursive Multi-Agent Systems — Xiyuan Yang, Jiaru Zou, Rui Pan, Ruizhong Qiu, Pan Lu, Shizhe Diao, Jindong Jiang, Hanghang Tong, Tong Zhang, Markus J. Buehler, Jingrui He, James Zou, 2026
http://arxiv.org/abs/2604.25917
2. CAMEL: Communicative Agents for "Mind" Exploration of Large Scale Language Model Society — Guohao Li, Hasan Abed Al Kader Hammoud, Hani Itani, Dmitrii Khizbullin, Bernard Ghanem, 2023
https://scholar.google.com/scholar?q=CAMEL%3A+Communicative+Agents+for+%22Mind%22+Exploration+of+Large+Scale+Language+Model+Society
3. AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation — Qingyun Wu, Gagan Bansal, Jieyu Zhang, Yiran Wu, Shaokun Zhang, Erkang Zhu, Beibin Li, Li Jiang, Xiaoyun Zhang, Ahmed Hassan Awadallah, Ryen W. White, Doug Burger, Chi Wang, 2023
https://scholar.google.com/scholar?q=AutoGen%3A+Enabling+Next-Gen+LLM+Applications+via+Multi-Agent+Conversation
4. MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework — Sirui Hong, Xiawu Zheng, Jonathan Chen, Yuheng Cheng, Jinlin Wang, Ceyao Zhang, Zili Wang, Steven Ka Shing Yau, Zijuan Lin, Liyang Zhou, Chenyu Ran, Lingfeng Xiao, Chenglin Wu, Jürgen Schmidhuber, 2023
https://scholar.google.com/scholar?q=MetaGPT%3A+Meta+Programming+for+A+Multi-Agent+Collaborative+Framework
5. Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach — Jonas Geiping, Sean McLeish, Neel Jain, John Kirchenbauer, Siddharth Singh, Brian R. Bartoldson, Bhavya Kailkhura, Abhinav Bhatele, Tom Goldstein, 2025
https://scholar.google.com/scholar?q=Scaling+up+Test-Time+Compute+with+Latent+Reasoning%3A+A+Recurrent+Depth+Approach
6. Learning Multiagent Communication with Backpropagation — Sainbayar Sukhbaatar, Arthur Szlam, Rob Fergus, 2016
https://scholar.google.com/scholar?q=Learning+Multiagent+Communication+with+Backpropagation
7. Learning to Communicate with Deep Multi-Agent Reinforcement Learning — Jakob Foerster, Ioannis Alexandros Assael, Nando de Freitas, Shimon Whiteson, 2016
https://scholar.google.com/scholar?q=Learning+to+Communicate+with+Deep+Multi-Agent+Reinforcement+Learning
8. Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments — Ryan Lowe, Yi Wu, Aviv Tamar, Jean Harb, Pieter Abbeel, Igor Mordatch, 2017
https://scholar.google.com/scholar?q=Multi-Agent+Actor-Critic+for+Mixed+Cooperative-Competitive+Environments
9. QMIX: Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement Learning — Tabish Rashid, Mikayel Samvelyan, Christian Schroeder de Witt, Gregory Farquhar, Jakob Foerster, Shimon Whiteson, 2018
https://scholar.google.com/scholar?q=QMIX%3A+Monotonic+Value+Function+Factorisation+for+Deep+Multi-Agent+Reinforcement+Learning
10. TextGrad: Automatic "Differentiation" via Text — Mert Yuksekgonul, Federico Bianchi, Joseph Boen, Sheng Liu, Zhi Huang, Carlos Guestrin, James Zou, 2024
https://scholar.google.com/scholar?q=TextGrad%3A+Automatic+%22Differentiation%22+via+Text
11. Mixture-of-Agents Enhances Large Language Model Capabilities — Junlin Wang, Jue Wang, Ben Athiwaratkun, Ce Zhang, James Zou, 2024
https://scholar.google.com/scholar?q=Mixture-of-Agents+Enhances+Large+Language+Model+Capabilities
12. ReDel: A Toolkit for LLM-Powered Recursive Multi-Agent Systems — Andrew Zhu, Liam Dugan, Chris Callison-Burch, 2024
https://scholar.google.com/scholar?q=ReDel%3A+A+Toolkit+for+LLM-Powered+Recursive+Multi-Agent+Systems
13. Why Do Multi-Agent LLM Systems Fail? — Mert Cemri, Melissa Z. Pan, Shuyi Yang, Lakshya A. Agrawal, Bhavya Chopra, Rishabh Tiwari, Kurt Keutzer, Aditya Parameswaran, Dan Klein, Kannan Ramchandran, Matei Zaharia, Joseph E. Gonzalez, Ion Stoica, 2025
https://scholar.google.com/scholar?q=Why+Do+Multi-Agent+LLM+Systems+Fail%3F
14. Scaling Latent Reasoning via Looped Language Models — Rui-Jie Zhu, Zixuan Wang, Kai Hua, Tianyu Zhang, Ziniu Li, Haoran Que, Boyi Wei, Zixin Wen, Fan Yin, He Xing, Lu Li, Jiajun Shi, Kaijing Ma, Shanda Li, Taylor Kergan, Andrew Smith, Xingwei Qu, Mude Hui, Bohong Wu, Qiyang Min, Hongzhi Huang, Xun Zhou, Wei Ye, Jiaheng Liu, Jian Yang, Yunfeng Shi, Chenghua Lin, Enduo Zhao, Tianle Cai, Ge Zhang, Wenhao Huang, Yoshua Bengio, Jason Eshraghian, 2025
https://scholar.google.com/scholar?q=Scaling+Latent+Reasoning+via+Looped+Language+Models
15. Enabling Agents to Communicate Entirely in Latent Space — Zhuoyun Du et al., 2025
https://scholar.google.com/scholar?q=Enabling+Agents+to+Communicate+Entirely+in+Latent+Space
16. Latent Collaboration in Multi-Agent Systems — Jiaru Zou, Xiyuan Yang et al., 2025
https://scholar.google.com/scholar?q=Latent+Collaboration+in+Multi-Agent+Systems
17. Reasoning Models Know When They're Right: Probing Hidden States for Self-Verification — Anqi Zhang et al., 2025
https://scholar.google.com/scholar?q=Reasoning+Models+Know+When+They%27re+Right%3A+Probing+Hidden+States+for+Self-Verification
18. Do Language Models Use Their Depth Efficiently? — Robert Csordas, Christopher D. Manning, Christopher Potts, 2025
https://scholar.google.com/scholar?q=Do+Language+Models+Use+Their+Depth+Efficiently%3F
19. Exploring Depth Generalization in Large Language Models for Solving Recursive Logic Tasks — Zhiyuan He, 2025
https://scholar.google.com/scholar?q=Exploring+Depth+Generalization+in+Large+Language+Models+for+Solving+Recursive+Logic+Tasks
20. Minimizing Response Latency in LLM-Based Agent Systems: A Comprehensive Survey — Gyeongmuk Park, Seonghyeon Lee, Yeonsu Park, 2026
https://scholar.google.com/scholar?q=Minimizing+Response+Latency+in+LLM-Based+Agent+Systems%3A+A+Comprehensive+Survey
21. AI Post Transformers: TUMIX Multi-Agent Test-Time Scaling with Tools — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-22-tumix-multi-agent-test-time-scaling-with-40671c.mp3
22. AI Post Transformers: Neural Computers as Learned Latent Runtimes — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-11-neural-computers-as-learned-latent-runti-9fa282.mp3
23. AI Post Transformers: VL-JEPA for Vision-Language Semantic Prediction — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-12-vl-jepa-for-vision-language-semantic-pre-69c9f4.mp3
24. AI Post Transformers: Neural Chameleons and Evading Activation Monitors — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-14-neural-chameleons-and-evading-activation-bc470e.mp3
25. AI Post Transformers: Efficient KV Cache Sharing for Multi-LoRA Agents — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-22-efficient-kv-cache-sharing-for-multi-lor-afda05.mp3
26. AI Post Transformers: TokenDance for Multi-Agent KV Cache Sharing — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-22-tokendance-for-multi-agent-kv-cache-shar-aa9b99.mp3

This episode explores whether learned discrete representations actually improve reinforcement learning, world modeling, and continual adaptation compared with standard continuous latent spaces. It explains how vector-quantized codebook latents, sparse binary-style features, and older ideas like tile coding relate to modern world models, and why the real advantage may come from reduced interference rather than discreteness alone. The discussion centers on three evaluation settings: predicting future dynamics in latent space, improving downstream control in model-free RL, and helping agents adapt to shifting tasks without forgetting earlier behavior. Listeners would find it interesting because it cuts through the “discrete vs. continuous” hype and turns the paper into a sharper engineering question about which representation bottlenecks produce more stable, reusable abstractions under changing conditions.

Sources:
1. Harnessing Discrete Representations For Continual Reinforcement Learning — Edan Meyer, Adam White, Marlos C. Machado, 2023
http://arxiv.org/abs/2312.01203
2. Neural Discrete Representation Learning — Aaron van den Oord, Oriol Vinyals, Koray Kavukcuoglu, 2017
https://papers.nips.cc/paper/2017/hash/7a98af17e63a0ac09ce2e96d03992fbc-Abstract.html
3. Mastering Atari with Discrete World Models — Danijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy Ba, 2021
https://openreview.net/forum?id=0oabwyZbOu
4. Transformers are Sample-Efficient World Models — Vincent Micheli, Eloi Alonso, Francois Fleuret, 2023
https://openreview.net/forum?id=vhFu1Acb0xb
5. Harnessing Discrete Representations for Continual Reinforcement Learning — Edan Jacob Meyer, Adam White, Marlos C. Machado, 2024
https://openreview.net/forum?id=tCXURNlAZ3
6. Mastering Diverse Domains through World Models — Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, Timothy P. Lillicrap, 2023
https://scholar.google.com/scholar?q=Mastering+Diverse+Domains+through+World+Models
7. Fuzzy Tiling Activations: A Simple Approach to Learning Sparse Representations Online — Yangchen Pan, Kirby Banman, Martha White, 2021
https://scholar.google.com/scholar?q=Fuzzy+Tiling+Activations%3A+A+Simple+Approach+to+Learning+Sparse+Representations+Online
8. Investigating the Properties of Neural Network Representations in Reinforcement Learning — Han Wang, Erfan Miahi, Martha White, Marlos C. Machado, Zaheer Abbas, Raksha Kumaraswamy, Vincent Liu, Adam White, 2022
https://scholar.google.com/scholar?q=Investigating+the+Properties+of+Neural+Network+Representations+in+Reinforcement+Learning
9. Smaller World Models for Reinforcement Learning — Jan Robine, Tobias Uelwer, Stefan Harmeling, 2021
https://scholar.google.com/scholar?q=Smaller+World+Models+for+Reinforcement+Learning
10. Continual Learning as Computationally Constrained Reinforcement Learning — Saurabh Kumar, Henrik Marklund, Ashish Rao, Yifan Zhu, Hong Jun Jeon, Yueyang Liu, Benjamin Van Roy, 2023
https://scholar.google.com/scholar?q=Continual+Learning+as+Computationally+Constrained+Reinforcement+Learning
11. Efficient World Models with Context-Aware Tokenization — Vincent Micheli, Eloi Alonso, Francois Fleuret, 2024
https://scholar.google.com/scholar?q=Efficient+World+Models+with+Context-Aware+Tokenization
12. AdaWorld: Learning Adaptable World Models with Latent Actions — authors not identifiable from the snippet, recent
https://scholar.google.com/scholar?q=AdaWorld%3A+Learning+Adaptable+World+Models+with+Latent+Actions
13. A Survey of Continual Reinforcement Learning — authors not identifiable from the snippet, recent
https://scholar.google.com/scholar?q=A+Survey+of+Continual+Reinforcement+Learning
14. Stable Continual Reinforcement Learning via Diffusion-Based Trajectory Replay — authors not identifiable from the snippet, recent
https://scholar.google.com/scholar?q=Stable+Continual+Reinforcement+Learning+via+Diffusion-Based+Trajectory+Replay
15. AI Post Transformers: DreamerV3 World Models Across 150 Tasks — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-20-dreamerv3-world-models-across-150-tasks-af5edb.mp3
16. AI Post Transformers: TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-25-turboquant-online-vector-quantiz-1967b7.mp3
17. AI Post Transformers: Latent Space as a New Computational Paradigm — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-05-latent-space-as-a-new-computational-para-810f39.mp3

This episode explores whether language models can express uncertainty in natural language in a way that is actually calibrated and useful, rather than merely sounding cautious or confident. It focuses on the paper’s central distinction between token uncertainty and epistemic uncertainty, arguing that next-token probabilities are a poor proxy for whether a model truly knows an answer. The discussion situates this idea alongside earlier work on Bayesian approximations, dataset shift, and newer hallucination-detection methods such as semantic entropy, all pointing to the same challenge: uncertainty should attach to claims, not just strings. A listener would find it interesting because it connects a seemingly simple design choice, having models state confidence in words, to the much larger problem of building AI systems that can warn users when they are likely guessing.

Sources:
1. Teaching Models to Express Their Uncertainty in Words — Stephanie Lin, Jacob Hilton, Owain Evans, 2022
http://arxiv.org/abs/2205.14334
2. Aleatoric and Epistemic Uncertainty in Machine Learning: An Introduction to Concepts and Methods — Eyke Hullermeier, Willem Waegeman, 2021
https://scholar.google.com/scholar?q=Aleatoric+and+Epistemic+Uncertainty+in+Machine+Learning%3A+An+Introduction+to+Concepts+and+Methods
3. Dropout as a Bayesian Approximation: Representing Model Uncertainty in Deep Learning — Yarin Gal, Zoubin Ghahramani, 2016
https://scholar.google.com/scholar?q=Dropout+as+a+Bayesian+Approximation%3A+Representing+Model+Uncertainty+in+Deep+Learning
4. Can You Trust Your Model's Uncertainty? Evaluating Predictive Uncertainty Under Dataset Shift — Yaniv Ovadia, Emily Fertig, Jie Ren, Zachary Nado, D. Sculley, Sebastian Nowozin, Joshua V. Dillon, Balaji Lakshminarayanan, Jasper Snoek, 2019
https://scholar.google.com/scholar?q=Can+You+Trust+Your+Model%27s+Uncertainty%3F+Evaluating+Predictive+Uncertainty+Under+Dataset+Shift
5. Detecting Hallucinations in Large Language Models Using Semantic Entropy — Sebastian Farquhar, Jannik Kossen, Lorenz Kuhn, and colleagues including Yarin Gal, 2024
https://scholar.google.com/scholar?q=Detecting+Hallucinations+in+Large+Language+Models+Using+Semantic+Entropy
6. Language Models (Mostly) Know What They Know — Katherine Kadavath, Roberta Raileanu, Akul Arora, et al., 2022
https://scholar.google.com/scholar?q=Language+Models+%28Mostly%29+Know+What+They+Know
7. On Calibration of Modern Neural Networks — Chuan Guo, Geoff Pleiss, Yu Sun, Kilian Q. Weinberger, 2017
https://scholar.google.com/scholar?q=On+Calibration+of+Modern+Neural+Networks
8. Can Language Models Be Too Honest? The Curious Case of Adaptive Honesty — Owain Evans, Jacob Hilton, Lukas Heim, et al., 2021
https://scholar.google.com/scholar?q=Can+Language+Models+Be+Too+Honest%3F+The+Curious+Case+of+Adaptive+Honesty
9. TruthfulQA: Measuring How Models Mimic Human Falsehoods — Stephanie Lin, Jacob Hilton, Owain Evans, 2021
https://scholar.google.com/scholar?q=TruthfulQA%3A+Measuring+How+Models+Mimic+Human+Falsehoods
10. Measuring Calibration in Deep Learning — Mahdi Pakdaman Naeini, Gregory Cooper, Milos Hauskrecht, 2015
https://scholar.google.com/scholar?q=Measuring+Calibration+in+Deep+Learning
11. Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation — Sebastian Kuhn, Yarin Gal, et al., 2023
https://scholar.google.com/scholar?q=Semantic+Uncertainty%3A+Linguistic+Invariances+for+Uncertainty+Estimation+in+Natural+Language+Generation
12. SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models — Potsawee Manakul, Adian Liusie, Mark Gales, 2023
https://scholar.google.com/scholar?q=SelfCheckGPT%3A+Zero-Resource+Black-Box+Hallucination+Detection+for+Generative+Large+Language+Models
13. Are Large Language Models More Honest in Their Probabilistic or Verbalized Confidence? — Shiyu Ni, Keping Bi, Lulu Yu, Jiafeng Guo, 2024
https://scholar.google.com/scholar?q=Are+Large+Language+Models+More+Honest+in+Their+Probabilistic+or+Verbalized+Confidence%3F
14. Calibrating Verbalized Probabilities for Large Language Models — Cheng Wang, Gyuri Szarvas, Georges Balazs, Pavel Danchenko, Patrick Ernst, 2024
https://scholar.google.com/scholar?q=Calibrating+Verbalized+Probabilities+for+Large+Language+Models
15. Know the Unknown: An Uncertainty-Sensitive Method for LLM Instruction Tuning — Jiaqi Li, Yixuan Tang, Yi Yang, 2025
https://scholar.google.com/scholar?q=Know+the+Unknown%3A+An+Uncertainty-Sensitive+Method+for+LLM+Instruction+Tuning
16. From Calibration to Collaboration: LLM Uncertainty Quantification Should Be More Human-Centered — Siddartha Devic, Tejas Srinivasan, Jesse Thomason, Willie Neiswanger, Vatsal Sharan, 2025
https://scholar.google.com/scholar?q=From+Calibration+to+Collaboration%3A+LLM+Uncertainty+Quantification+Should+Be+More+Human-Centered
17. Closing the Confidence-Faithfulness Gap in Large Language Models — Miranda Muqing Miao, Lyle Ungar, 2026
https://scholar.google.com/scholar?q=Closing+the+Confidence-Faithfulness+Gap+in+Large+Language+Models
18. AI Post Transformers: CLUE: Hidden-State Clustering for Non-parametric Verification — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/clue-hidden-state-clustering-for-non-parametric-verification/
19. AI Post Transformers: Neural Chameleons and Evading Activation Monitors — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-14-neural-chameleons-and-evading-activation-bc470e.mp3
20. AI Post Transformers: Latent Space as a New Computational Paradigm — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-05-latent-space-as-a-new-computational-para-810f39.mp3
Interactive Visualization: Teaching Language Models to Verbalize Uncertainty

This episode explores ChartNet, a 1.5 million-sample multimodal dataset designed to improve how vision-language models read and reason about charts. It explains why chart understanding is harder than OCR or captioning alone, because models must connect visual marks, axes, legends, numerical values, and language-based reasoning with high precision. The discussion places ChartNet in the context of earlier benchmarks like DVQA, PlotQA, ChartQA, and UniChart, arguing that past datasets were too small or too narrow to teach robust chart comprehension. It also examines ChartNet’s code-guided pipeline, where models reconstruct plotting code from seed charts, generate structurally varied new examples, and align each chart with images, code, tables, summaries, and QA, making the episode interesting for listeners who want to understand whether scale and multimodal alignment can produce more reliable chart-reading AI.

Sources:
1. ChartNet for Robust Multimodal Chart Understanding
https://arxiv.org/pdf/2603.27064
2. DVQA: Understanding Data Visualizations via Question Answering — Kushal Kafle, Brian Price, Scott Cohen, Christopher Kanan, 2018
https://scholar.google.com/scholar?q=DVQA%3A+Understanding+Data+Visualizations+via+Question+Answering
3. PlotQA: Reasoning over Scientific Plots — Nitesh Methani, Pritha Ganguly, Mitesh M. Khapra, Pratyush Kumar, 2019
https://scholar.google.com/scholar?q=PlotQA%3A+Reasoning+over+Scientific+Plots
4. ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning — Ahmed Masry, Do Xuan Long, Jia Qing Tan, Shafiq Joty, Enamul Hoque, 2022
https://scholar.google.com/scholar?q=ChartQA%3A+A+Benchmark+for+Question+Answering+about+Charts+with+Visual+and+Logical+Reasoning
5. UniChart: A Universal Vision-language Pretrained Model for Chart Comprehension and Reasoning — Ahmed Masry, Parsa Kavehzadeh, Xuan Long Do, Enamul Hoque, Shafiq Joty, 2023
https://scholar.google.com/scholar?q=UniChart%3A+A+Universal+Vision-language+Pretrained+Model+for+Chart+Comprehension+and+Reasoning
6. TinyChart: Efficient Chart Understanding with Visual Token Merging and Program-of-Thoughts Learning — L. Zhang, A. Hu, H. Xu, M. Yan, Y. Xu, Q. Jin, J. Zhang, and F. Huang, 2024
https://scholar.google.com/scholar?q=TinyChart%3A+Efficient+Chart+Understanding+with+Visual+Token+Merging+and+Program-of-Thoughts+Learning
7. ChartQAPro: A More Diverse and Challenging Benchmark for Chart Question Answering — A. Masry, M. S. Islam, M. Ahmed, A. Bajaj, F. Kabir, A. Kartha, M. T. R. Laskar, M. Rahman, S. Rahman, M. Shahmohammadi, et al., 2025
https://scholar.google.com/scholar?q=ChartQAPro%3A+A+More+Diverse+and+Challenging+Benchmark+for+Chart+Question+Answering
8. EvoChart: A Benchmark and a Self-Training Approach Towards Real-World Chart Understanding — M. Huang, H. Lai, X. Zhang, W. Wu, J. Ma, L. Zhang, and J. Liu, 2025
https://scholar.google.com/scholar?q=EvoChart%3A+A+Benchmark+and+a+Self-Training+Approach+Towards+Real-World+Chart+Understanding
9. ChartMimic: Evaluating LMM's Cross-Modal Reasoning Capability via Chart-to-Code Generation — C. Yang, C. Shi, Y. Liu, B. Shui, J. Wang, M. Jing, L. Xu, X. Zhu, S. Li, Y. Zhang, G. Liu, X. Nie, D. Cai, and Y. Yang, 2025
https://scholar.google.com/scholar?q=ChartMimic%3A+Evaluating+LMM%27s+Cross-Modal+Reasoning+Capability+via+Chart-to-Code+Generation
10. OpenCQA: Open-Ended Question Answering with Charts — S. Kantharaj, X. L. Do, R. T. K. Leong, J. Q. Tan, E. Hoque, and S. Joty, 2022
https://scholar.google.com/scholar?q=OpenCQA%3A+Open-Ended+Question+Answering+with+Charts
11. Effective Training Data Synthesis for Improving MLLM Chart Understanding — approximate; unknown from snippet, 2024/2025
https://scholar.google.com/scholar?q=Effective+Training+Data+Synthesis+for+Improving+MLLM+Chart+Understanding
12. From Charts to Code: A Hierarchical Benchmark for Multimodal Models — approximate; unknown from snippet, 2024/2025
https://scholar.google.com/scholar?q=From+Charts+to+Code%3A+A+Hierarchical+Benchmark+for+Multimodal+Models
13. GRAFT: GRaPH and Table Reasoning for Textual Alignment — approximate; unknown from snippet, 2024/2025
https://scholar.google.com/scholar?q=GRAFT%3A+GRaPH+and+Table+Reasoning+for+Textual+Alignment
14. ChartQA-X: Generating Explanations for Visual Chart Reasoning — approximate; unknown from snippet, 2024/2025
https://scholar.google.com/scholar?q=ChartQA-X%3A+Generating+Explanations+for+Visual+Chart+Reasoning
15. AI Post Transformers: Procgen Benchmark: Measuring Generalization in Reinforcement Learning — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/procgen-benchmark-measuring-generalization-in-reinforcement-learning/
16. AI Post Transformers: The Endless Gym: Training Terminal Agents — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/the-endless-gym-training-terminal-agents/
17. AI Post Transformers: Evaluating Large Language Models Trained on Code — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/evaluating-large-language-models-trained-on-code/
18. AI Post Transformers: VL-JEPA for Vision-Language Semantic Prediction — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-12-vl-jepa-for-vision-language-semantic-pre-69c9f4.mp3
19. AI Post Transformers: IMO-Bench for Robust Mathematical Reasoning — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-imo-bench-for-robust-mathematical-reason-143489.mp3
Interactive Visualization: ChartNet for Robust Multimodal Chart Understanding

This episode explores whether large language models can estimate their chances of success before acting, update those estimates during a task, and use them to decide when to abstain from costly work. It explains why that matters for agentic systems in coding and software environments, where overconfidence can lead to wasted effort, unsafe actions, or expensive mistakes, and connects the paper to earlier work on calibration, uncertainty, and selective abstention. The discussion highlights the paper’s focus on three settings, including single-step coding, sequential decisions with feedback, and multi-step software engineering, while also stressing the distinction between raw capability, calibration, and rational decision-making. Listeners would find it interesting because it treats self-assessment not as a philosophical question, but as a practical requirement for building AI systems that know when not to act.

Interactive Visualization: Can LLMs Judge Their Own Capabilities?
Sources:
1. Do Large Language Models Know What They Are Capable Of? — Casey O. Barkan, Sid Black, Oliver Sourbut, 2025
http://arxiv.org/abs/2512.24661
2. Language Models (Mostly) Know What They Know — Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Ethan Perez, Nicholas Joseph, and collaborators, 2022
https://scholar.google.com/scholar?q=Language+Models+%28Mostly%29+Know+What+They+Know
3. Do Large Language Models Know What They Don’t Know? — Zhangyue Yin, Qiushi Sun, Qipeng Guo, Jiawen Wu, Xipeng Qiu, Xuanjing Huang, 2023
https://scholar.google.com/scholar?q=Do+Large+Language+Models+Know+What+They+Don%E2%80%99t+Know%3F
4. Towards Fully Exploiting LLM Internal States to Enhance Knowledge Boundary Perception — Shiyu Ni, Keping Bi, Jiafeng Guo, Lulu Yu, Baolong Bi, Xueqi Cheng, 2025
https://scholar.google.com/scholar?q=Towards+Fully+Exploiting+LLM+Internal+States+to+Enhance+Knowledge+Boundary+Perception
5. What Large Language Models Know and What People Think They Know — Mark Steyvers, Heliodoro Tejeda, Aakriti Kumar, Catarina Belem, Lukas W. Mayer, Padhraic Smyth, and collaborators, 2025
https://scholar.google.com/scholar?q=What+Large+Language+Models+Know+and+What+People+Think+They+Know
6. Teaching Models to Express Their Uncertainty in Words — Stephanie Lin, Jacob Hilton, Owain Evans, 2022
https://scholar.google.com/scholar?q=Teaching+Models+to+Express+Their+Uncertainty+in+Words
7. A Survey of Confidence Estimation and Calibration in Large Language Models — Jiahui Geng, Fengyu Cai, Yuxia Wang, Heinz Koeppl, Preslav Nakov, Iryna Gurevych, 2024
https://scholar.google.com/scholar?q=A+Survey+of+Confidence+Estimation+and+Calibration+in+Large+Language+Models
8. Calibrating Language Models with Adaptive Temperature Scaling — Johnathan Xie, Annie S. Chen, Yoonho Lee, Eric Mitchell, Chelsea Finn, 2024
https://scholar.google.com/scholar?q=Calibrating+Language+Models+with+Adaptive+Temperature+Scaling
9. Calibrating the Confidence of Large Language Models by Eliciting Fidelity — Mozhi Zhang, Mianqiu Huang, Rundong Shi, Linsen Guo, Chong Peng, Peng Yan, Yaqian Zhou, Xipeng Qiu, 2024
https://scholar.google.com/scholar?q=Calibrating+the+Confidence+of+Large+Language+Models+by+Eliciting+Fidelity
10. Selective Classification for Deep Neural Networks — Yonatan Geifman, Ran El-Yaniv, 2017
https://scholar.google.com/scholar?q=Selective+Classification+for+Deep+Neural+Networks
11. SelectiveNet: A Deep Neural Network with an Integrated Reject Option — Yonatan Geifman, Ran El-Yaniv, 2019
https://scholar.google.com/scholar?q=SelectiveNet%3A+A+Deep+Neural+Network+with+an+Integrated+Reject+Option
12. Selective-LAMA: Selective Prediction for Confidence-Aware Evaluation of Language Models — Hiyori Yoshikawa, Naoaki Okazaki, 2023
https://scholar.google.com/scholar?q=Selective-LAMA%3A+Selective+Prediction+for+Confidence-Aware+Evaluation+of+Language+Models
13. Confidence-Driven Multi-Scale Model Selection for Cost-Efficient Inference — Bo-Wei Chen, Chung-Chi Chen, An-Zi Yen, 2026
https://scholar.google.com/scholar?q=Confidence-Driven+Multi-Scale+Model+Selection+for+Cost-Efficient+Inference
14. Quantifying Uncert-AI-nty: Testing the Accuracy of LLMs' Confidence Judgments — Trent N. Cash, Daniel M. Oppenheimer, Sara Christie, Mira Devgan, 2025
https://scholar.google.com/scholar?q=Quantifying+Uncert-AI-nty%3A+Testing+the+Accuracy+of+LLMs%27+Confidence+Judgments
15. Credence Calibration Game? Calibrating Large Language Models Through Structured Play — Ke Fang, Tianyi Zhao, Lu Cheng, 2025
https://scholar.google.com/scholar?q=Credence+Calibration+Game%3F+Calibrating+Large+Language+Models+Through+Structured+Play
16. Calibration and Correctness of Language Models for Code — Claudio Spiess, David Gros, Kunal Suresh Pai, Michael Pradel, Md Rafiqul Islam Rabin, Amin Alipour, Susmit Jha, Prem Devanbu, Toufique Ahmed, 2025
https://scholar.google.com/scholar?q=Calibration+and+Correctness+of+Language+Models+for+Code
17. Large Language Models Must Be Taught to Know What They Don't Know — Sanyam Kapoor, Nate Gruver, Manley Roberts, Katie Collins, Arka Pal, Umang Bhatt, Adrian Weller, Samuel Dooley, Micah Goldblum, Andrew G. Wilson, 2024
https://scholar.google.com/scholar?q=Large+Language+Models+Must+Be+Taught+to+Know+What+They+Don%27t+Know
18. SMARTCAL: An Approach to Self-Aware Tool-Use Evaluation and Calibration — Yuanhao Shen, Xiaodan Zhu, Lei Chen, 2024
https://scholar.google.com/scholar?q=SMARTCAL%3A+An+Approach+to+Self-Aware+Tool-Use+Evaluation+and+Calibration
19. Calibration Is Not Enough: Evaluating Confidence Estimation Under Language Variations — Yuxi Xia, Dennis Ulmer, Terra Blevins, Yihong Liu, Hinrich Schutze, Benjamin Roth, 2026
https://scholar.google.com/scholar?q=Calibration+Is+Not+Enough%3A+Evaluating+Confidence+Estimation+Under+Language+Variations
20. CoT-UQ: Improving Response-wise Uncertainty Quantification in LLMs with Chain-of-Thought — Boxuan Zhang, Ruqi Zhang, 2025
https://scholar.google.com/scholar?q=CoT-UQ%3A+Improving+Response-wise+Uncertainty+Quantification+in+LLMs+with+Chain-of-Thought
21. Structured Uncertainty guided Clarification for LLM Agents — Manan Suri, Puneet Mathur, Nedim Lipka, Franck Dernoncourt, Ryan A. Rossi, Dinesh Manocha, 2025
https://scholar.google.com/scholar?q=Structured+Uncertainty+guided+Clarification+for+LLM+Agents
22. CAR-bench: Evaluating the Consistency and Limit-Awareness of LLM Agents under Real-World Uncertainty — Johannes Kirmayr, Lukas Stappen, Elisabeth Andre, 2026
https://scholar.google.com/scholar?q=CAR-bench%3A+Evaluating+the+Consistency+and+Limit-Awareness+of+LLM+Agents+under+Real-World+Uncertainty
23. Improving Interactive In-Context Learning from Natural Language Feedback — Martin Klissarov, Jonathan Cook, Diego Antognini, Hao Sun, Jingling Li, Natasha Jaques, Claudiu Musat, Edward Grefenstette, 2026
https://scholar.google.com/scholar?q=Improving+Interactive+In-Context+Learning+from+Natural+Language+Feedback
24. AI Post Transformers: Benchmarking Test-Time Scaling for General LLM Agents — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-22-benchmarking-test-time-scaling-for-gener-8f14f9.mp3
25. AI Post Transformers: Program Synthesis with Large Language Models — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-20-program-synthesis-with-large-language-mo-b962ec.mp3
26. AI Post Transformers: SkillsBench for Evaluating Agent Skills — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-14-skillsbench-for-evaluating-agent-skills-58bb1e.mp3
27. AI Post Transformers: Stabilizing Efficient Reasoning with Step-Level Advantage Selection — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-29-stabilizing-efficient-reasoning-with-ste-1e589d.mp3
28. AI Post Transformers: Neural Chameleons and Evading Activation Monitors — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-14-neural-chameleons-and-evading-activation-bc470e.mp3
Interactive Visualization: Can LLMs Judge Their Own Capabilities?

This episode explores the 2017 ConvS2S paper from Facebook AI Research, which argued that sequence-to-sequence models for machine translation did not need recurrence and could instead use fully convolutional encoder-decoder networks with attention. It explains the seq2seq and neural machine translation setup, clarifies that the paper replaces recurrent computation rather than attention, and breaks down the architecture’s core ideas: stacked convolutions, expanding receptive fields, positional embeddings, residual connections, and gated linear units. The discussion highlights why the approach was provocative at the time: it challenged the LSTM-based consensus, reported stronger BLEU scores, and promised much faster, more parallelizable decoding on GPUs. Listeners would find it interesting as a key pre-transformer moment that shows how researchers were already rethinking sequential modeling, while also surfacing the tradeoff between efficiency and limited context windows for long-range dependencies.

Sources:
1. Convolutional Sequence to Sequence Learning — Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, Yann N. Dauphin, 2017
http://arxiv.org/abs/1705.03122
2. Convolutional Sequence to Sequence Learning — Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, Yann N. Dauphin, 2017
https://scholar.google.com/scholar?q=Convolutional+Sequence+to+Sequence+Learning
3. ByteNet: Generating High-Resolution Discrete Sequences with Multiscale Dilation — Nal Kalchbrenner, Lasse Espeholt, Karen Simonyan, Aaron van den Oord, Alex Graves, Koray Kavukcuoglu, 2016
https://scholar.google.com/scholar?q=ByteNet%3A+Generating+High-Resolution+Discrete+Sequences+with+Multiscale+Dilation
4. Language Modeling with Gated Convolutional Networks — Yann N. Dauphin, Angela Fan, Michael Auli, David Grangier, 2017
https://scholar.google.com/scholar?q=Language+Modeling+with+Gated+Convolutional+Networks
5. WaveNet: A Generative Model for Raw Audio — Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, et al., 2016
https://scholar.google.com/scholar?q=WaveNet%3A+A+Generative+Model+for+Raw+Audio
6. Sequence to Sequence Learning with Neural Networks — Ilya Sutskever, Oriol Vinyals, Quoc V. Le, 2014
https://scholar.google.com/scholar?q=Sequence+to+Sequence+Learning+with+Neural+Networks
7. Neural Machine Translation by Jointly Learning to Align and Translate — Dzmitry Bahdanau, Kyunghyun Cho, Yoshua Bengio, 2014
https://scholar.google.com/scholar?q=Neural+Machine+Translation+by+Jointly+Learning+to+Align+and+Translate
8. Effective Approaches to Attention-based Neural Machine Translation — Minh-Thang Luong, Hieu Pham, Christopher D. Manning, 2015
https://scholar.google.com/scholar?q=Effective+Approaches+to+Attention-based+Neural+Machine+Translation
9. A Survey on Deep Learning for Neural Machine Translation — Antonio Toral, Víctor M. Sánchez-Cartagena, et al. (representative survey literature varies by edition), 2018
https://scholar.google.com/scholar?q=A+Survey+on+Deep+Learning+for+Neural+Machine+Translation
10. Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation — Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V. Le, Mohammad Norouzi, et al., 2016
https://scholar.google.com/scholar?q=Google%27s+Neural+Machine+Translation+System%3A+Bridging+the+Gap+between+Human+and+Machine+Translation
11. Neural Machine Translation of Rare Words with Subword Units — Rico Sennrich, Barry Haddow, Alexandra Birch, 2016
https://scholar.google.com/scholar?q=Neural+Machine+Translation+of+Rare+Words+with+Subword+Units
12. Attention Is All You Need — Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, Illia Polosukhin, 2017
https://scholar.google.com/scholar?q=Attention+Is+All+You+Need
13. Neural Machine Translation and Sequence-to-sequence Models: A Tutorial — Graham Neubig, 2017
https://scholar.google.com/scholar?q=Neural+Machine+Translation+and+Sequence-to-sequence+Models%3A+A+Tutorial
14. Neural Machine Translation in Linear Time — Jonas Gehring, Michael Auli, David Grangier, Yann N. Dauphin, 2016
https://scholar.google.com/scholar?q=Neural+Machine+Translation+in+Linear+Time
15. ByteNet: Neural Machine Translation in Linear Time — Nal Kalchbrenner, Lasse Espeholt, Karen Simonyan, Aaron van den Oord, Alex Graves, Koray Kavukcuoglu, 2016
https://scholar.google.com/scholar?q=ByteNet%3A+Neural+Machine+Translation+in+Linear+Time
16. Quasi-Recurrent Neural Networks — James Bradbury, Stephen Merity, Caiming Xiong, Richard Socher, 2016
https://scholar.google.com/scholar?q=Quasi-Recurrent+Neural+Networks
17. Convolutional Neural Network for Machine Translation — Fandong Meng, Zhengdong Lu, Hang Li, Qun Liu, 2015
https://scholar.google.com/scholar?q=Convolutional+Neural+Network+for+Machine+Translation
18. Resurrecting Recurrent Neural Networks for Long Sequences — approx. Orvieto et al., 2023
https://scholar.google.com/scholar?q=Resurrecting+Recurrent+Neural+Networks+for+Long+Sequences
19. A Comprehensive Survey on Long Context Language Modeling — approx. recent survey authors, exact authorship to verify, 2024
https://scholar.google.com/scholar?q=A+Comprehensive+Survey+on+Long+Context+Language+Modeling
20. Linear Recurrent Models for Robust and Interpretable Long-context Modeling — approx. recent long-context modeling authors, exact authorship to verify, 2024
https://scholar.google.com/scholar?q=Linear+Recurrent+Models+for+Robust+and+Interpretable+Long-context+Modeling
21. Cross-layer Attention Sharing for Pre-trained Large Language Models — approx. recent LLM systems authors, exact authorship to verify, 2024
https://scholar.google.com/scholar?q=Cross-layer+Attention+Sharing+for+Pre-trained+Large+Language+Models
22. Reducing Transformer Key-Value Cache Size with Cross-Layer Attention — approx. recent efficiency-focused LLM authors, exact authorship to verify, 2024
https://scholar.google.com/scholar?q=Reducing+Transformer+Key-Value+Cache+Size+with+Cross-Layer+Attention
23. AI Post Transformers: RoPE — Hal Turing & Dr. Ada Shannon, Thu,
https://podcast.do-not-panic.com/episodes/rope/
24. AI Post Transformers: Apple's Speculative Streaming: Fast LLM Inference without Auxiliary Models — Hal Turing & Dr. Ada Shannon, Sat,
https://podcast.do-not-panic.com/episodes/apples-speculative-streaming-fast-llm-inference-without-auxiliary-models/
25. AI Post Transformers: Speculative Decoding in Real vLLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-speculative-decoding-in-real-vllm-servin-6f4e2b.mp3
26. AI Post Transformers: FlatAttention for Tile-Based Accelerator Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-flatattention-for-tile-based-accelerator-56e6ca.mp3

This episode explores the 2014–2015 breakthrough paper that introduced attention to neural machine translation, framing it as a solution to a specific flaw in early encoder-decoder models: forcing an entire source sentence into one fixed-length vector. It explains how pre-attention RNN-based seq2seq systems struggled on long or complex sentences, and how Bahdanau et al.’s “soft alignment” let the decoder focus on different source words at each generation step instead of relying on a single compressed summary. Along the way, it situates the paper against phrase-based statistical translation and earlier LSTM/GRU seq2seq models, showing why attention was more than a performance tweak—it was a durable structural idea. Listeners would find it interesting for its clear account of what was actually broken before attention, what changed technically, and why this paper became a foundational step toward modern language models.

Sources:
1. Neural Machine Translation by Jointly Learning to Align and Translate — Dzmitry Bahdanau, Kyunghyun Cho, Yoshua Bengio, 2014
http://arxiv.org/abs/1409.0473
2. Sequence to Sequence Learning with Neural Networks — Ilya Sutskever, Oriol Vinyals, Quoc V. Le, 2014
https://scholar.google.com/scholar?q=Sequence+to+Sequence+Learning+with+Neural+Networks
3. Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation — Kyunghyun Cho, Bart van Merrienboer, Dzmitry Bahdanau, Yoshua Bengio, 2014
https://scholar.google.com/scholar?q=Learning+Phrase+Representations+using+RNN+Encoder-Decoder+for+Statistical+Machine+Translation
4. A Neural Network for Machine Translation, at Production Scale — Nal Kalchbrenner, Phil Blunsom, 2013
https://scholar.google.com/scholar?q=A+Neural+Network+for+Machine+Translation%2C+at+Production+Scale
5. On the Properties of Neural Machine Translation: Encoder-Decoder Approaches — Kyunghyun Cho, Bart van Merrienboer, Çağlar Gülçehre, Fethi Bougares, Holger Schwenk, Yoshua Bengio, 2014
https://scholar.google.com/scholar?q=On+the+Properties+of+Neural+Machine+Translation%3A+Encoder-Decoder+Approaches
6. Moses: Open Source Toolkit for Statistical Machine Translation — Philipp Koehn, Hieu Hoang, Alexandra Birch, Chris Callison-Burch, Marcello Federico, Nicola Bertoldi, Brooke Cowan, Wade Shen, Christine Moran, Richard Zens, et al., 2007
https://scholar.google.com/scholar?q=Moses%3A+Open+Source+Toolkit+for+Statistical+Machine+Translation
7. Neural Machine Translation by Jointly Learning to Align and Translate — Dzmitry Bahdanau, Kyunghyun Cho, Yoshua Bengio, 2014
https://scholar.google.com/scholar?q=Neural+Machine+Translation+by+Jointly+Learning+to+Align+and+Translate
8. Effective Approaches to Attention-based Neural Machine Translation — Minh-Thang Luong, Hieu Pham, Christopher D. Manning, 2015
https://scholar.google.com/scholar?q=Effective+Approaches+to+Attention-based+Neural+Machine+Translation
9. Neural Machine Translation of Rare Words with Subword Units — Rico Sennrich, Barry Haddow, Alexandra Birch, 2016
https://scholar.google.com/scholar?q=Neural+Machine+Translation+of+Rare+Words+with+Subword+Units
10. Looking for a Needle in a Haystack: A Comprehensive Study of Hallucinations in Neural Machine Translation — approx. Guerreiro et al., 2023
https://scholar.google.com/scholar?q=Looking+for+a+Needle+in+a+Haystack%3A+A+Comprehensive+Study+of+Hallucinations+in+Neural+Machine+Translation
11. Key, Value, Compress: A Systematic Exploration of KV Cache Compression Techniques — approx. recent systems/LLM authors, 2024
https://scholar.google.com/scholar?q=Key%2C+Value%2C+Compress%3A+A+Systematic+Exploration+of+KV+Cache+Compression+Techniques
12. DeltaKV: Residual-Based KV Cache Compression via Long-Range Similarity — approx. recent systems/LLM authors, 2024
https://scholar.google.com/scholar?q=DeltaKV%3A+Residual-Based+KV+Cache+Compression+via+Long-Range+Similarity
13. A Survey on Large Language Model Acceleration Based on KV Cache Management — approx. recent survey authors, 2024
https://scholar.google.com/scholar?q=A+Survey+on+Large+Language+Model+Acceleration+Based+on+KV+Cache+Management
14. AI Post Transformers: Memory Sparse Attention for 100M-Token Scaling — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-07-memory-sparse-attention-for-100m-token-s-377cff.mp3
15. AI Post Transformers: Mamba-3 for Efficient Sequence Modeling — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-16-mamba-3-for-efficient-sequence-modeling-97a22a.mp3
16. AI Post Transformers: Latent Space as a New Computational Paradigm — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-05-latent-space-as-a-new-computational-para-810f39.mp3
17. AI Post Transformers: Doc-to-LoRA: Internalizing Context as LoRA — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-29-doc-to-lora-internalizing-context-as-lor-8dd5ec.mp3

This episode explores the 2014 seq2seq paper by Sutskever and colleagues as a turning point in machine translation, asking what problem it actually solved and why it mattered at the time. It explains how deep LSTM encoder-decoder models reframed translation as end-to-end learning of \(p(\text{output}|\text{input})\), replacing hand-built phrase tables, alignment models, and decoding heuristics with a single learned system. The discussion highlights both the breakthrough and the limitation: compressing an entire source sentence into one fixed-length vector was elegant but created a severe bottleneck, which later made attention mechanisms so important. Listeners would find it interesting because the episode separates the paper’s real contribution from its mythology, situating it against phrase-based SMT, earlier encoder-decoder work, and the practical role of LSTMs and beam search in making the approach viable.

Sources:
1. Sequence to Sequence Learning with Neural Networks — Ilya Sutskever, Oriol Vinyals, Quoc V. Le, 2014
http://arxiv.org/abs/1409.3215
2. Sequence to Sequence Learning with Neural Networks — Ilya Sutskever, Oriol Vinyals, Quoc V. Le, 2014
https://scholar.google.com/scholar?q=Sequence+to+Sequence+Learning+with+Neural+Networks
3. Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation — Kyunghyun Cho, Bart van Merrienboer, Dzmitry Bahdanau, Yoshua Bengio and collaborators, 2014
https://scholar.google.com/scholar?q=Learning+Phrase+Representations+using+RNN+Encoder-Decoder+for+Statistical+Machine+Translation
4. Neural Machine Translation by Jointly Learning to Align and Translate — Dzmitry Bahdanau, Kyunghyun Cho, Yoshua Bengio, 2014
https://scholar.google.com/scholar?q=Neural+Machine+Translation+by+Jointly+Learning+to+Align+and+Translate
5. A Survey of the Usages of Deep Learning for Natural Language Processing — Tom Young, Devamanyu Hazarika, Soujanya Poria, Erik Cambria, 2018
https://scholar.google.com/scholar?q=A+Survey+of+the+Usages+of+Deep+Learning+for+Natural+Language+Processing
6. Long Short-Term Memory — Sepp Hochreiter, Jürgen Schmidhuber, 1997
https://scholar.google.com/scholar?q=Long+Short-Term+Memory
7. Learning to Forget: Continual Prediction with LSTM — Felix A. Gers, Jürgen Schmidhuber, Fred Cummins, 2000
https://scholar.google.com/scholar?q=Learning+to+Forget%3A+Continual+Prediction+with+LSTM
8. Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling — Junyoung Chung, Caglar Gulcehre, Kyunghyun Cho, Yoshua Bengio, 2014
https://scholar.google.com/scholar?q=Empirical+Evaluation+of+Gated+Recurrent+Neural+Networks+on+Sequence+Modeling
9. An Empirical Exploration of Recurrent Network Architectures — Razvan Pascanu, Tomas Mikolov, Yoshua Bengio, 2013
https://scholar.google.com/scholar?q=An+Empirical+Exploration+of+Recurrent+Network+Architectures
10. Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation — Yonghui Wu and many Google Brain collaborators, 2016
https://scholar.google.com/scholar?q=Google%27s+Neural+Machine+Translation+System%3A+Bridging+the+Gap+between+Human+and+Machine+Translation
11. A Call for Clarity in Reporting BLEU Scores — Matt Post, 2018
https://scholar.google.com/scholar?q=A+Call+for+Clarity+in+Reporting+BLEU+Scores
12. On the Properties of Neural Machine Translation: Encoder-Decoder Approaches — Kyunghyun Cho and collaborators, 2014
https://scholar.google.com/scholar?q=On+the+Properties+of+Neural+Machine+Translation%3A+Encoder-Decoder+Approaches
13. Beam Search Strategies for Neural Machine Translation — Markus Freitag, Yaser Al-Onaizan, 2017
https://scholar.google.com/scholar?q=Beam+Search+Strategies+for+Neural+Machine+Translation
14. A Systematic Comparison of Search Algorithms for Neural Machine Translation — Various later comparison studies, e.g. work by Felix Stahlberg and collaborators, 2018
https://scholar.google.com/scholar?q=A+Systematic+Comparison+of+Search+Algorithms+for+Neural+Machine+Translation
15. Skip-Thought Vectors — Ryan Kiros, Yukun Zhu, Russ R. Salakhutdinov, Richard Zemel, Raquel Urtasun, Antonio Torralba, Sanja Fidler, 2015
https://scholar.google.com/scholar?q=Skip-Thought+Vectors
16. Supervised Learning of Universal Sentence Representations from Natural Language Inference Data — Alexis Conneau, Douwe Kiela, Holger Schwenk, Loic Barrault, Antoine Bordes, 2017
https://scholar.google.com/scholar?q=Supervised+Learning+of+Universal+Sentence+Representations+from+Natural+Language+Inference+Data
17. Universal Sentence Encoder — Daniel Cer, Yinfei Yang, Sheng-yi Kong and collaborators, 2018
https://scholar.google.com/scholar?q=Universal+Sentence+Encoder
18. A Neural Network for Machine Translation, at Production Scale — Nal Kalchbrenner, Edward Grefenstette, Phil Blunsom, 2014
https://scholar.google.com/scholar?q=A+Neural+Network+for+Machine+Translation%2C+at+Production+Scale
19. Sequence Transduction with Recurrent Neural Networks — Alex Graves, 2012
https://scholar.google.com/scholar?q=Sequence+Transduction+with+Recurrent+Neural+Networks
20. Connectionist Temporal Classification: Labelling Unsegmented Sequence Data with Recurrent Neural Networks — Alex Graves, Santiago Fernandez, Faustino Gomez, Jürgen Schmidhuber, 2006
https://scholar.google.com/scholar?q=Connectionist+Temporal+Classification%3A+Labelling+Unsegmented+Sequence+Data+with+Recurrent+Neural+Networks
21. Statistical Machine Translation: The Mathematics of SMT — Philipp Koehn, 2010
https://scholar.google.com/scholar?q=Statistical+Machine+Translation%3A+The+Mathematics+of+SMT
22. Recurrent Neural Network Based Language Model — Tomáš Mikolov, Martin Karafiát, Lukáš Burget, Jan Černocký, Sanjeev Khudanpur, 2010
https://scholar.google.com/scholar?q=Recurrent+Neural+Network+Based+Language+Model
23. Effective Approaches to Attention-based Neural Machine Translation — Luong, Pham, Manning, 2015
https://scholar.google.com/scholar?q=Effective+Approaches+to+Attention-based+Neural+Machine+Translation
24. Learning Transductions and Alignments with RNN Seq2seq Models — approx. Britz, Goldie, Luong, Le, or related Google authors, 2017
https://scholar.google.com/scholar?q=Learning+Transductions+and+Alignments+with+RNN+Seq2seq+Models
25. Neural Machine Translation of Rare Words with Subword Units — Sennrich, Haddow, Birch, 2016
https://scholar.google.com/scholar?q=Neural+Machine+Translation+of+Rare+Words+with+Subword+Units
26. On Using Very Large Target Vocabulary for Neural Machine Translation — Jean, Cho, Memisevic, Bengio, 2015
https://scholar.google.com/scholar?q=On+Using+Very+Large+Target+Vocabulary+for+Neural+Machine+Translation
27. Attention Is All You Need — Vaswani et al., 2017
https://scholar.google.com/scholar?q=Attention+Is+All+You+Need
28. AI Post Transformers: Language Models are Injective and Hence Invertible — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-21-language-models-are-injective-an-7545e0.mp3
29. AI Post Transformers: Experimental Comparison of Agentic and Enhanced RAG — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-14-experimental-comparison-of-agentic-and-e-37d8bc.mp3

This episode explores how the 2021 paper “Hopfield Networks Is All You Need” reframes transformer attention as a modern continuous Hopfield network, connecting attention, associative memory, and similarity-based retrieval under one mathematical lens. It explains the core idea of content-addressable memory, contrasts classical binary Hopfield networks with newer differentiable versions, and shows why attention can be understood not just as weighted averaging but as an energy-based retrieval process with fixed points and attractor states. The discussion highlights the paper’s major claims: one-step retrieval, exponential storage capacity, low retrieval error under assumptions, and distinct retrieval regimes such as global averaging and subset averaging. Listeners interested in AI theory will find it compelling because it offers a concrete, less mystical interpretation of transformer heads and suggests practical memory-layer designs grounded in formal guarantees.

Sources:
1. Hopfield Networks is All You Need — Hubert Ramsauer, Bernhard Schäfl, Johannes Lehner, Philipp Seidl, Michael Widrich, Thomas Adler, Lukas Gruber, Markus Holzleitner, Milena Pavlović, Geir Kjetil Sandve, Victor Greiff, David Kreil, Michael Kopp, Günter Klambauer, Johannes Brandstetter, Sepp Hochreiter, 2020
http://arxiv.org/abs/2008.02217
2. https://isl.stanford.edu/~cover/papers/transIT/0021cove.pdf
https://isl.stanford.edu/~cover/papers/transIT/0021cove.pdf
3. Neural Networks and Physical Systems with Emergent Collective Computational Abilities — John J. Hopfield, 1982
https://scholar.google.com/scholar?q=Neural+Networks+and+Physical+Systems+with+Emergent+Collective+Computational+Abilities
4. A Neural Network with Locality-Sensitive Hashed Dynamics — Dmitry Krotov, John J. Hopfield, 2016
https://scholar.google.com/scholar?q=A+Neural+Network+with+Locality-Sensitive+Hashed+Dynamics
5. Hopfield Networks is All You Need — Hubert Ramsauer, Bernhard Schäfl, Johannes Lehner, Philipp Seidl, Michael Widrich, Thomas Adler, Lukas Gruber, Markus Holzleitner, Milena Pavlović, Geir Kjetil Sandve, Victor Greiff, David Kreil, Michael Kopp, Günter Klambauer, Johannes Brandstetter, Sepp Hochreiter, 2020
https://scholar.google.com/scholar?q=Hopfield+Networks+is+All+You+Need
6. Dense Associative Memory for Pattern Recognition — Dmitry Krotov, John J. Hopfield, 2020
https://scholar.google.com/scholar?q=Dense+Associative+Memory+for+Pattern+Recognition
7. Neural Turing Machines — Alex Graves, Greg Wayne, Ivo Danihelka, 2014
https://scholar.google.com/scholar?q=Neural+Turing+Machines
8. Memory Networks — Jason Weston, Sumit Chopra, Antoine Bordes, 2014
https://scholar.google.com/scholar?q=Memory+Networks
9. End-To-End Memory Networks — Sainbayar Sukhbaatar, Arthur Szlam, Jason Weston, Rob Fergus, 2015
https://scholar.google.com/scholar?q=End-To-End+Memory+Networks
10. Hybrid computing using a neural network with dynamic external memory — Alex Graves, Greg Wayne, Malcolm Reynolds, Tim Harley, Ivo Danihelka, Agnieszka Grabska-Barwińska, Sergio Gómez Colmenarejo, Edward Grefenstette, Tiago Ramalho, John Agapiou, Adrià Puigdomènech Badia, Karl Moritz Hermann, Yori Zwols, Georg Ostrovski, Adam Cain, Helen King, Christopher Summerfield, Phil Blunsom, Koray Kavukcuoglu, Demis Hassabis, 2016
https://scholar.google.com/scholar?q=Hybrid+computing+using+a+neural+network+with+dynamic+external+memory
11. Attention Is All You Need — Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, Illia Polosukhin, 2017
https://scholar.google.com/scholar?q=Attention+Is+All+You+Need
12. Dense Associative Memory is Robust to Adversarial Inputs — Dmitry Krotov, John J. Hopfield, 2018
https://scholar.google.com/scholar?q=Dense+Associative+Memory+is+Robust+to+Adversarial+Inputs
13. A Robust Exponential Associative Memory with Fixed Point Analysis — Mert Demircigil, Judith Heusel, Matthias Löwe, Sven Upgang, Franck Vermet, 2017
https://scholar.google.com/scholar?q=A+Robust+Exponential+Associative+Memory+with+Fixed+Point+Analysis
14. Set Transformer: A Framework for Attention-based Permutation-Invariant Neural Networks — Juho Lee, Yoonho Lee, Jungtaek Kim, Adam Kosiorek, Seungjin Choi, Yee Whye Teh, 2019
https://scholar.google.com/scholar?q=Set+Transformer%3A+A+Framework+for+Attention-based+Permutation-Invariant+Neural+Networks
15. Deep Sets — Manzil Zaheer, Satwik Kottur, Siamak Ravanbakhsh, Barnabas Poczos, Ruslan Salakhutdinov, Alexander Smola, 2017
https://scholar.google.com/scholar?q=Deep+Sets
16. Modern Hopfield Networks and Attention for Immune Repertoire Classification — Michael Widrich, Bernhard Schäfl, Milena Pavlović, Hubert Ramsauer, Lukas Gruber, Markus Holzleitner, Geir Kjetil Sandve, Victor Greiff, Sepp Hochreiter, et al., 2020
https://scholar.google.com/scholar?q=Modern+Hopfield+Networks+and+Attention+for+Immune+Repertoire+Classification
17. Attention Heads of Large Language Models — authors unclear from snippet, likely 2024-2025
https://scholar.google.com/scholar?q=Attention+Heads+of+Large+Language+Models
18. Causal Head Gating: A Framework for Interpreting Roles of Attention Heads in Transformers — authors unclear from snippet, likely 2024-2025
https://scholar.google.com/scholar?q=Causal+Head+Gating%3A+A+Framework+for+Interpreting+Roles+of+Attention+Heads+in+Transformers
19. Mechanistic Interpretability of Fine-Tuned Vision Transformers on Distorted Images: Decoding Attention Head Behavior for Transparent and Trustworthy AI — authors unclear from snippet, likely 2024-2025
https://scholar.google.com/scholar?q=Mechanistic+Interpretability+of+Fine-Tuned+Vision+Transformers+on+Distorted+Images%3A+Decoding+Attention+Head+Behavior+for+Transparent+and+Trustworthy+AI
20. Iterative Sparse Attention for Long-Sequence Recommendation — authors unclear from snippet, likely recent
https://scholar.google.com/scholar?q=Iterative+Sparse+Attention+for+Long-Sequence+Recommendation
21. An Evolved Universal Transformer Memory — authors unclear from snippet, likely recent
https://scholar.google.com/scholar?q=An+Evolved+Universal+Transformer+Memory
22. AI Post Transformers: In-Place Test-Time Training for Transformers — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-09-in-place-test-time-training-for-transfor-d0b976.mp3
23. AI Post Transformers: RoPE — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/rope/
24. AI Post Transformers: Jet-Nemotron and PostNAS for Faster Long Context — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-24-jet-nemotron-and-postnas-for-faster-long-436381.mp3
25. AI Post Transformers: Experimental Comparison of Agentic and Enhanced RAG — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-14-experimental-comparison-of-agentic-and-e-37d8bc.mp3
26. AI Post Transformers: Doc-to-LoRA: Internalizing Context as LoRA — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-29-doc-to-lora-internalizing-context-as-lor-8dd5ec.mp3

Browse by month

Sep 2026 Aug 2026 Jul 2026 Jun 2026 May 2026 Apr 2026 Mar 2026 Feb 2026 Jan 2026 Dec 2025 Nov 2025 Oct 2025 Sep 2025 Aug 2025