This episode explores TTKV, a temporal-tiered key-value cache design for long-context LLM inference, where decode speed degrades because growing KV state turns generation into a memory-bandwidth problem rather than a compute problem. It explains how the method keeps recent cache blocks in fast GPU HBM, evicts older blocks to slower host DRAM, and uses asymmetric quantization in the slow tier, preserving keys at higher precision while compressing values more aggressively. The discussion also breaks down the runtime mechanics behind block-wise streaming attention, including query-conditioned block ranking, top-k prefetching, decompression, and overlapping data transfer with attention computation. What makes the episode interesting is that it treats TTKV less as a new model idea and more as a systems design proposal, while critically questioning whether recency is a reliable proxy for importance and whether the paper fully specifies the cost of its block-selection function.

Sources:
1. TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference — Gradwell Dzikanyanga, Weihao Yang, Hao Huang, Donglei Wu, Shihao Wang, Wen Xia, Sanjeeb K C, 2026
http://arxiv.org/abs/2604.19769
2. Efficient Memory Management for Large Language Model Serving with PagedAttention — Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, Ion Stoica, 2023
https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention
3. Efficient Streaming Language Models with Attention Sinks — Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, Mike Lewis, 2023
https://scholar.google.com/scholar?q=Efficient+Streaming+Language+Models+with+Attention+Sinks
4. KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache — Zirui Liu, Jiayi Yuan, Hongye Jin, Shaochen Zhong, Zhaozhuo Xu, Vladimir Braverman, Beidi Chen, Xia Hu, 2024
https://scholar.google.com/scholar?q=KIVI%3A+A+Tuning-Free+Asymmetric+2bit+Quantization+for+KV+Cache
5. ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference — Hanshi Sun, Li-Wen Chang, Wenlei Bao, Size Zheng, Ningxin Zheng, Xin Liu, Harry Dong, Yuejie Chi, Beidi Chen, 2024
https://scholar.google.com/scholar?q=ShadowKV%3A+KV+Cache+in+Shadows+for+High-Throughput+Long-Context+LLM+Inference
6. FreeKV: Boosting KV Cache Retrieval for Efficient LLM Inference — Guangda Liu et al., 2025
https://scholar.google.com/scholar?q=FreeKV%3A+Boosting+KV+Cache+Retrieval+for+Efficient+LLM+Inference
7. FIER: Fine-Grained and Efficient KV Cache Retrieval for Long-context LLM Inference — Dongwei Wang et al., 2025
https://scholar.google.com/scholar?q=FIER%3A+Fine-Grained+and+Efficient+KV+Cache+Retrieval+for+Long-context+LLM+Inference
8. CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving — Yuhan Liu et al., 2024
https://scholar.google.com/scholar?q=CacheGen%3A+KV+Cache+Compression+and+Streaming+for+Fast+Large+Language+Model+Serving
9. Retrieval Head Mechanistically Explains Long-Context Factuality — Wenhao Wu et al., 2024
https://arxiv.org/abs/2404.15574
10. DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads — Guangxuan Xiao et al., 2024
https://arxiv.org/abs/2410.10819
11. Query-Focused Retrieval Heads Improve Long-Context Reasoning and Re-ranking — Wuwei Zhang et al., 2025
https://arxiv.org/abs/2506.09944
12. LKV: End-to-End Learning of Head-wise Budgets and Token Selection for LLM KV Cache Eviction — Enshuai Zhou et al., 2026
https://arxiv.org/abs/2605.06676
13. IndexMem: Learned KV-Cache Eviction with Latent Memory for Long-Context LLM Inference — Xintong Yang et al., 2026
https://arxiv.org/abs/2605.25475
14. KVTuner: Sensitivity-Aware Layer-wise Mixed Precision KV Cache Quantization for Efficient and Nearly Lossless LLM Inference — Xing Li et al., 2025
https://arxiv.org/abs/2502.04420
15. AsymKV: Enabling 1-Bit Quantization of KV Cache with Layer-Wise Asymmetric Quantization Configurations — Qian Tao et al., 2024
https://arxiv.org/abs/2410.13212
16. vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention — Ramya Prabhu et al., 2024
https://arxiv.org/abs/2405.04437
17. AI Post Transformers: ScoutAttention for Efficient KV Cache Offloading — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-24-scoutattention-for-efficient-kv-cache-of-b26699.mp3
18. AI Post Transformers: MiniMax Sparse Attention at Million-Token Scale — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-13-minimax-sparse-attention-at-million-toke-300108.mp3
19. AI Post Transformers: IndexMem: Learned KV-Cache Eviction for Long-Context LLMs — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-12-indexmem-learned-kv-cache-eviction-for-l-132c2a.mp3
20. AI Post Transformers: Mooncake for KV Cache-Centric LLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-05-mooncake-for-kv-cache-centric-llm-servin-1086d0.mp3
21. AI Post Transformers: Memory-Bound, Not Bandwidth-Limited Batch-1 LLM Decode — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-02-memory-bound-not-bandwidth-limited-batch-114799.mp3
22. AI Post Transformers: Why LLM Serving Needs Mathematical Optimization — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-why-llm-serving-needs-mathematical-optim-647fc6.mp3

This episode explores SuperInfer, a system for serving large language models on GH200-style superchips by treating memory management as the key lever for meeting latency targets rather than just maximizing compute use. It explains why KV cache growth, HBM pressure, and head-of-line blocking often hurt responsiveness first, then breaks down how the paper’s RotaSched policy proactively rotates request state out of fast memory to protect time-to-first-token deadlines. It also covers DuplexKV, the transfer mechanism that makes this practical by batching fragmented KV data, using bidirectional movement across NVLink-C2C, and overlapping transfers with model execution instead of stalling the whole system. Listeners would find it interesting because the discussion ties concrete serving pain points to a specific systems design that reportedly boosts TTFT SLO attainment by up to 74.7 percent while keeping throughput and token pacing roughly stable.

Sources:
1. SuperInfer: SLO-Aware Rotary Scheduling and Memory Management for LLM Inference on Superchips — Jiahuan Yu, Mingtao Hu, Zichao Lin, Minjia Zhang, 2026
http://arxiv.org/abs/2601.20309
2. Pie: Pooling CPU Memory for LLM Inference — Y. Xu, Z. Mao, X. Mo, S. Liu, I. Stoica, 2024
https://scholar.google.com/scholar?q=Pie%3A+Pooling+CPU+Memory+for+LLM+Inference
3. Understanding Data Movement in Tightly Coupled Heterogeneous Systems: A Case Study with the Grace Hopper Superchip — L. Fusco, M. Khalilov, M. Chrapek, G. Chukkapalli, T. Schulthess, T. Hoefler, 2024
https://scholar.google.com/scholar?q=Understanding+Data+Movement+in+Tightly+Coupled+Heterogeneous+Systems%3A+A+Case+Study+with+the+Grace+Hopper+Superchip
4. Memory Offloading for Large Language Model Inference with Latency SLO Guarantees — C. Ma, Z. Ye, H. Zhao, Z. Yang, T. Fu, J. Han, J. Zhang, Y. Luo, X. Wang, Z. Wang, et al., 2025
https://scholar.google.com/scholar?q=Memory+Offloading+for+Large+Language+Model+Inference+with+Latency+SLO+Guarantees
5. Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve — A. Agrawal, N. Kedia, A. Panwar, J. Mohan, N. Kwatra, B. Gulavani, A. Tumanov, R. Ramjee, 2024
https://scholar.google.com/scholar?q=Taming+Throughput-Latency+Tradeoff+in+LLM+Inference+with+Sarathi-Serve
6. Mooncake: Trading More Storage for Less Computation - a KVCache-centric Architecture for Serving LLM Chatbot — R. Qin, Z. Li, W. He, J. Cui, F. Ren, M. Zhang, Y. Wu, W. Zheng, X. Xu, 2025
https://scholar.google.com/scholar?q=Mooncake%3A+Trading+More+Storage+for+Less+Computation+-+a+KVCache-centric+Architecture+for+Serving+LLM+Chatbot
7. TokenLake: A Unified Segment-level Prefix Cache Pool for Fine-grained Elastic Long-Context LLM Serving — Bingyang Wu et al., 2025
https://scholar.google.com/scholar?q=TokenLake%3A+A+Unified+Segment-level+Prefix+Cache+Pool+for+Fine-grained+Elastic+Long-Context+LLM+Serving
8. KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows — Zaifeng Pan et al., 2025
https://scholar.google.com/scholar?q=KVFlow%3A+Efficient+Prefix+Caching+for+Accelerating+LLM-Based+Multi-Agent+Workflows
9. SAW-INT4: System-Aware 4-Bit KV-Cache Quantization for Real-World LLM Serving — Jinda Jia et al., 2026
https://scholar.google.com/scholar?q=SAW-INT4%3A+System-Aware+4-Bit+KV-Cache+Quantization+for+Real-World+LLM+Serving
10. KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization — Coleman Hooper et al., 2024
https://scholar.google.com/scholar?q=KVQuant%3A+Towards+10+Million+Context+Length+LLM+Inference+with+KV+Cache+Quantization
11. DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving — Yinmin Zhong et al., 2024
https://scholar.google.com/scholar?q=DistServe%3A+Disaggregating+Prefill+and+Decoding+for+Goodput-optimized+Large+Language+Model+Serving
12. Prefill-Decode Aggregation or Disaggregation? Unifying Both for Goodput-Optimized LLM Serving — Chao Wang et al., 2025
https://scholar.google.com/scholar?q=Prefill-Decode+Aggregation+or+Disaggregation%3F+Unifying+Both+for+Goodput-Optimized+LLM+Serving
13. Enhancing LLM Efficiency: Targeted Pruning for Prefill-Decode Disaggregation in Inference — Hao Zhang et al., 2025
https://scholar.google.com/scholar?q=Enhancing+LLM+Efficiency%3A+Targeted+Pruning+for+Prefill-Decode+Disaggregation+in+Inference
14. AI Post Transformers: Mooncake for KV Cache-Centric LLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-05-mooncake-for-kv-cache-centric-llm-servin-1086d0.mp3
15. AI Post Transformers: LAPS for Length-Aware LLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-laps-for-length-aware-llm-serving-0c6149.mp3
16. AI Post Transformers: AI+HW 2035: Co-Designing Efficient AI Systems — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-24-aihw-2035-co-designing-efficient-ai-syst-95c11e.mp3
17. AI Post Transformers: FengHuang for Rack-Scale LLM Inference Memory — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-12-fenghuang-for-rack-scale-llm-inference-m-62708e.mp3
18. AI Post Transformers: KVSwap for Disk-Aware Long-Context On-Device Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-16-kvswap-for-disk-aware-long-context-on-de-f3c15e.mp3
19. AI Post Transformers: Memory-Bound, Not Bandwidth-Limited Batch-1 LLM Decode — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-02-memory-bound-not-bandwidth-limited-batch-114799.mp3

This episode explores DSpark, a DeepSeek-AI paper on improving speculative decoding by starting from a DFlash-style block-parallel draft model and increasing how often a larger verifier accepts its proposed tokens. It explains the mechanics of speculative decoding in plain language, situates DSpark within earlier blockwise and multi-token prediction work, and notes that the technique is already used in serving stacks such as vLLM, TensorRT-LLM, and SGLang. The discussion focuses on DSpark’s concrete additions: a Markov head that feeds previous-token information into draft logits, a confidence head that estimates whether drafted tokens will survive verification, and a training recipe centered on knowledge distillation. It is interesting because it treats inference speed as an operational systems problem, arguing that higher acceptance matters but only alongside draft latency, verifier cost, batching, and scheduler behavior.

Sources:
1. DSpark Improves Speculative Decoding Acceptance Rates
https://github.com/deepseek-ai/DeepSpec/blob/main/DSpark_paper.pdf
2. DFlash: Block Diffusion for Flash Speculative Decoding — Jian Chen, Yesheng Liang, Zhijian Liu, 2026
https://scholar.google.com/scholar?q=DFlash%3A+Block+Diffusion+for+Flash+Speculative+Decoding
3. EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test — Yuhui Li, Fangyun Wei, Chao Zhang, Hongyang Zhang, 2025
https://scholar.google.com/scholar?q=EAGLE-3%3A+Scaling+up+Inference+Acceleration+of+Large+Language+Models+via+Training-Time+Test
4. Decoding Speculative Decoding — Minghao Yan, Saurabh Agarwal, Shivaram Venkataraman, 2024
https://scholar.google.com/scholar?q=Decoding+Speculative+Decoding
5. Speculative Decoding with a Speculative Vocabulary — Miles Williams, Young D. Kwon, Rui Li, Alexandros Kouris, Stylianos I. Venieris, 2026
https://scholar.google.com/scholar?q=Speculative+Decoding+with+a+Speculative+Vocabulary
6. DFlare: Scaling Up Draft Capacity for Block Diffusion Speculative Decoding — Jiebin Zhang, Zhenghan Yu, Song Liu, Eugene J. Yu, Zheng Li, Dawei Zhu, Jiangshan Duo, Weimin Xiong, Yifan Song, Guanghua Yu, Jianchen Zhu, Sujian Li, 2026
https://scholar.google.com/scholar?q=DFlare%3A+Scaling+Up+Draft+Capacity+for+Block+Diffusion+Speculative+Decoding
7. AI Post Transformers: Speculative Decoding in Real vLLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-speculative-decoding-in-real-vllm-servin-6f4e2b.mp3
8. AI Post Transformers: Adaptive Control for Batched Speculative Decoding in LLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/adaptive-control-for-batched-speculative-decoding-in-llm-serving/
9. AI Post Transformers: JETSPEC and Parallel Tree Speculative Decoding — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-27-jetspec-and-parallel-tree-speculative-de-3d144c.mp3

This episode explores LLMServingSim 2.0, a simulator designed to model how large language models behave when they are served on mixed hardware fleets with separated compute, memory, and networking resources rather than a uniform GPU cluster. It explains the practical serving concepts that shape user experience, including prefill versus decode, time to first token, time per output token, prefix caching, KV-cache movement, and why latency problems emerge from interactions among batching, routing, placement, and interconnect contention rather than a single bottleneck. The discussion highlights the paper’s core idea of a Model Serving Group, which combines queueing, scheduling, operation mapping, memory modeling, and power modeling into one runtime-style unit driven by measured hardware profiles instead of purely theoretical kernel estimates. Listeners would find it interesting because it shows how modern AI performance depends not just on better models, but on the messy systems engineering tradeoffs that determine speed, efficiency, and scalability in real deployments.

Sources:
1. LLMServingSim 2.0: A Unified Simulator for Heterogeneous and Disaggregated LLM Serving Infrastructure — Jaehong Cho, Hyunmin Choi, Guseul Heo, Jongse Park, 2026
http://arxiv.org/abs/2602.23036
2. DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving — Yinmin Zhong et al., 2024
https://scholar.google.com/scholar?q=DistServe%3A+Disaggregating+Prefill+and+Decoding+for+Goodput-optimized+Large+Language+Model+Serving
3. P/D-Serve: Serving Disaggregated Large Language Model at Scale — Yibo Jin et al., 2024
https://scholar.google.com/scholar?q=P%2FD-Serve%3A+Serving+Disaggregated+Large+Language+Model+at+Scale
4. LMCache: An Efficient KV Cache Layer for Enterprise-Scale LLM Inference — Yuhan Liu, Yihua Cheng, Jiayi Yao, et al., 2025
https://scholar.google.com/scholar?q=LMCache%3A+An+Efficient+KV+Cache+Layer+for+Enterprise-Scale+LLM+Inference
5. Mooncake: Trading More Storage for Less Computation - A KVCache-centric Architecture for Serving LLM Chatbot — Ruoyu Qin et al., 2025
https://scholar.google.com/scholar?q=Mooncake%3A+Trading+More+Storage+for+Less+Computation+-+A+KVCache-centric+Architecture+for+Serving+LLM+Chatbot
6. NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM Inferencing — Guseul Heo et al., 2024
https://scholar.google.com/scholar?q=NeuPIMs%3A+NPU-PIM+Heterogeneous+Acceleration+for+Batched+LLM+Inferencing
7. Frontier: Towards Comprehensive and Accurate LLM Inference Simulation — Yicheng Feng et al., 2026
https://scholar.google.com/scholar?q=Frontier%3A+Towards+Comprehensive+and+Accurate+LLM+Inference+Simulation
8. KVLink: Accelerating Large Language Models via Efficient KV Cache Reuse — Jingbo Yang et al., 2025
https://scholar.google.com/scholar?q=KVLink%3A+Accelerating+Large+Language+Models+via+Efficient+KV+Cache+Reuse
9. Selective KV-Cache Sharing to Mitigate Timing Side-Channels in LLM Inference — Kexin Chu et al., 2025
https://scholar.google.com/scholar?q=Selective+KV-Cache+Sharing+to+Mitigate+Timing+Side-Channels+in+LLM+Inference
10. Prefill-Decode Aggregation or Disaggregation? Unifying Both for Goodput-Optimized LLM Serving — Chao Wang et al., 2025
https://scholar.google.com/scholar?q=Prefill-Decode+Aggregation+or+Disaggregation%3F+Unifying+Both+for+Goodput-Optimized+LLM+Serving
11. Enhancing LLM Efficiency: Targeted Pruning for Prefill-Decode Disaggregation in Inference — Hao Zhang et al., 2025
https://scholar.google.com/scholar?q=Enhancing+LLM+Efficiency%3A+Targeted+Pruning+for+Prefill-Decode+Disaggregation+in+Inference
12. AdaServe: SLO-Customized LLM Serving with Fine-Grained Speculative Decoding — Zikun Li et al., 2025
https://scholar.google.com/scholar?q=AdaServe%3A+SLO-Customized+LLM+Serving+with+Fine-Grained+Speculative+Decoding
13. SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding — Ziyi Zhang et al., 2025
https://scholar.google.com/scholar?q=SwiftSpec%3A+Ultra-Low+Latency+LLM+Decoding+by+Scaling+Asynchronous+Speculative+Decoding
14. Nightjar: Dynamic Adaptive Speculative Decoding for Large Language Models Serving — Rui Li et al., 2025
https://scholar.google.com/scholar?q=Nightjar%3A+Dynamic+Adaptive+Speculative+Decoding+for+Large+Language+Models+Serving
15. Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism — Zizhao Mo et al., 2025
https://scholar.google.com/scholar?q=Hetis%3A+Serving+LLMs+in+Heterogeneous+GPU+Clusters+with+Fine-grained+and+Dynamic+Parallelism
16. AI Post Transformers: Why LLM Serving Needs Mathematical Optimization — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-why-llm-serving-needs-mathematical-optim-647fc6.mp3
17. AI Post Transformers: FengHuang for Rack-Scale LLM Inference Memory — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-12-fenghuang-for-rack-scale-llm-inference-m-62708e.mp3
18. AI Post Transformers: Vistara Brings CXL Memory to Hyperscale — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-11-vistara-brings-cxl-memory-to-hyperscale-b5199e.mp3
19. AI Post Transformers: Characterizing LLM KV Cache Workloads in Production — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/characterizing-llm-kv-cache-workloads-in-production/
20. AI Post Transformers: Memory-Bound, Not Bandwidth-Limited Batch-1 LLM Decode — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-02-memory-bound-not-bandwidth-limited-batch-114799.mp3
21. AI Post Transformers: LPU Chip for Low-Latency LLM Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-20-lpu-chip-for-low-latency-llm-inference-be13c3.mp3
22. AI Post Transformers: TokenDance for Multi-Agent KV Cache Sharing — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-22-tokendance-for-multi-agent-kv-cache-shar-aa9b99.mp3

This episode explores Moebius, a serving system for mixture-of-experts transformers that can switch at runtime between tensor parallelism and expert parallelism without restarting or draining live requests. It explains why tensor parallelism tends to give lower latency at low concurrency, while expert parallelism delivers better throughput at high concurrency, making bursty online traffic and RL rollouts natural settings where the best strategy changes over time. The discussion focuses on the hard systems problems behind that switch, including migrating in-flight requests, preserving paged KV caches, coping with CUDA graph address constraints, and handling KV-head mismatches that can waste cache capacity under tensor parallelism. It argues that the paper’s key contribution is treating the switch as a change in ownership and memory layout over one resident model and KV state, offering a concrete blueprint for serving large sparse models more efficiently.

Sources:
1. Moebius: Serving Mixture-of-Expert Models with Seamless Runtime Parallelism Switch — Shaoyu Wang, Yizhuo Liang, Jaeyong Song, Chong Li, Seo Jin Park, 2026
http://arxiv.org/abs/2606.26607
2. GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding — Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, et al., 2020
https://scholar.google.com/scholar?q=GShard%3A+Scaling+Giant+Models+with+Conditional+Computation+and+Automatic+Sharding
3. Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity — William Fedus, Barret Zoph, Noam Shazeer, 2021
https://scholar.google.com/scholar?q=Switch+Transformers%3A+Scaling+to+Trillion+Parameter+Models+with+Simple+and+Efficient+Sparsity
4. DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale — Samyam Rajbhandari, Conglong Li, Zhewei Yao, et al., 2022
https://scholar.google.com/scholar?q=DeepSpeed-MoE%3A+Advancing+Mixture-of-Experts+Inference+and+Training+to+Power+Next-Generation+AI+Scale
5. MegaBlocks: Efficient Sparse Training with Mixture-of-Experts — Trevor Gale, Deepak Narayanan, Cliff Young, Matei Zaharia, 2022
https://scholar.google.com/scholar?q=MegaBlocks%3A+Efficient+Sparse+Training+with+Mixture-of-Experts
6. Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism — Mohammad Shoeybi, Mostofa Patwary, Raul Puri, et al., 2019
https://scholar.google.com/scholar?q=Megatron-LM%3A+Training+Multi-Billion+Parameter+Language+Models+Using+Model+Parallelism
7. Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM — Deepak Narayanan, Mohammad Shoeybi, Jared Casper, et al., 2021
https://scholar.google.com/scholar?q=Efficient+Large-Scale+Language+Model+Training+on+GPU+Clusters+Using+Megatron-LM
8. Efficient Memory Management for Large Language Model Serving with PagedAttention — Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, et al., 2023
https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention
9. Nitsum: Serving Tiered LLM Requests with Adaptive Tensor Parallelism — Vikranth Srivatsa, Zijian He, Pu Guo, et al., 2026
https://scholar.google.com/scholar?q=Nitsum%3A+Serving+Tiered+LLM+Requests+with+Adaptive+Tensor+Parallelism
10. HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference — Haoran Lin et al., 2025
https://scholar.google.com/scholar?q=HAP%3A+Hybrid+Adaptive+Parallelism+for+Efficient+Mixture-of-Experts+Inference
11. Amoeba: Runtime Tensor Parallel Transformation for LLM Inference Services — Haoyu Chen et al., 2026
https://scholar.google.com/scholar?q=Amoeba%3A+Runtime+Tensor+Parallel+Transformation+for+LLM+Inference+Services
12. UCCL-EP: Portable Expert-Parallel Communication — Ziming Mao et al., 2026
https://scholar.google.com/scholar?q=UCCL-EP%3A+Portable+Expert-Parallel+Communication
13. RollPacker: Mitigating Long-Tail Rollouts for Fast, Synchronous RL Post-Training — Wei Gao et al., 2026
https://scholar.google.com/scholar?q=RollPacker%3A+Mitigating+Long-Tail+Rollouts+for+Fast%2C+Synchronous+RL+Post-Training
14. HD-MoE: Hybrid and Dynamic Parallelism for Mixture-of-Expert LLMs with 3D Near-Memory Processing — Haochen Huang et al., 2025
https://arxiv.org/abs/2509.09420
15. fMoE: Fine-Grained Expert Offloading for Large Mixture-of-Experts Serving — Hanfei Yu et al., 2025
https://arxiv.org/abs/2502.05370
16. HOBBIT: A Mixed Precision Expert Offloading System for Fast MoE Inference — Peng Tang et al., 2024
https://arxiv.org/abs/2411.01433
17. CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving — Yuhan Liu et al., 2023
https://arxiv.org/abs/2310.07240
18. KVzip: Query-Agnostic KV Cache Compression with Context Reconstruction — Jang-Hyun Kim et al., 2025
https://arxiv.org/abs/2505.23416
19. AI Post Transformers: JANUS for Scalable MoE Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-janus-for-scalable-moe-inference-78ae30.mp3
20. AI Post Transformers: Serving MoE Models with Disaggregated Expert Parallelism — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-19-serving-moe-models-with-disaggregated-ex-6979d2.mp3
21. AI Post Transformers: Batch-Aware Expert Routing for Faster MoE Decoding — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-batch-aware-expert-routing-for-faster-mo-683ab6.mp3
22. AI Post Transformers: Affordable Large-Scale Decoding Through Model-System Co-Design — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-19-affordable-large-scale-decoding-through-e1d7ed.mp3
23. AI Post Transformers: Mooncake for KV Cache-Centric LLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-05-mooncake-for-kv-cache-centric-llm-servin-1086d0.mp3
24. AI Post Transformers: Why LLM Serving Needs Mathematical Optimization — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-why-llm-serving-needs-mathematical-optim-647fc6.mp3

This episode explores Information-Aware KV Cache Compression for Long Reasoning, a paper about making long-context inference cheaper and more reliable by deciding which KV-cache tokens to keep during extended reasoning. It explains why long prefilling and long decoding turn the cache into a major memory bottleneck, and why common heuristics such as sliding windows or recent-attention-based retention can discard tokens that only become important much later. The discussion centers on the paper’s claim that future usefulness is better captured by information-theoretic signals like predictive entropy and Forward Influence, with experiments showing that attention-ranked tokens help short-horizon predictions while entropy-ranked tokens matter more over long horizons. Listeners get a concrete account of how InfoKV blends recent attention with per-layer entropy-based scoring to improve the tradeoff between memory savings and long-range reasoning quality.

Sources:
1. Information-Aware KV Cache Compression for Long Reasoning — Jushi Kai, Zhuiri Xiao, Alexandra Birch, Zhouhan Lin, 2026
http://arxiv.org/abs/2606.26875
2. Scissorhands: Exploiting the Persistence of Importance Hypothesis for LLM KV Cache Compression at Test Time — Zichang Liu, Aditya Desai, Fangshuo Liao, Anshumali Shrivastava, et al., 2023
https://scholar.google.com/scholar?q=Scissorhands%3A+Exploiting+the+Persistence+of+Importance+Hypothesis+for+LLM+KV+Cache+Compression+at+Test+Time
3. H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models — Zhenyu Zhang, Ying Sheng, Tianyi Zhou, Beidi Chen, Christopher Re, et al., 2023
https://scholar.google.com/scholar?q=H2O%3A+Heavy-Hitter+Oracle+for+Efficient+Generative+Inference+of+Large+Language+Models
4. SnapKV: LLM Knows What You are Looking for Before Generation — Yuhong Li, Yingbing Huang, Bowen Yang, Patrick Lewis, Deming Chen, et al., 2024
https://scholar.google.com/scholar?q=SnapKV%3A+LLM+Knows+What+You+are+Looking+for+Before+Generation
5. Information-Aware KV Cache Compression for Long Reasoning — Jushi Kai, Zhuiri Xiao, Alexandra Birch, Zhouhan Lin, 2026
https://scholar.google.com/scholar?q=Information-Aware+KV+Cache+Compression+for+Long+Reasoning
6. Expected Attention: KV Cache Compression by Estimating Attention from Future Queries Distribution — Alessio Devoto, Maximilian Jeblick, Simon Jegou, 2025
https://scholar.google.com/scholar?q=Expected+Attention%3A+KV+Cache+Compression+by+Estimating+Attention+from+Future+Queries+Distribution
7. Reasoning Path Compression: Compressing Generation Trajectories for Efficient LLM Reasoning — Jiwon Song, Dongwon Jo, Yulhwa Kim, Jae-Joon Kim, 2025
https://scholar.google.com/scholar?q=Reasoning+Path+Compression%3A+Compressing+Generation+Trajectories+for+Efficient+LLM+Reasoning
8. Compressing Context to Enhance Inference Efficiency of Large Language Models — Yucheng Li, Bo Dong, Chenghua Lin, Frank Guerin, 2023
https://scholar.google.com/scholar?q=Compressing+Context+to+Enhance+Inference+Efficiency+of+Large+Language+Models
9. FreqKV: Key-Value Compression in Frequency Domain for Context Window Extension — Jushi Kai et al., 2026
https://scholar.google.com/scholar?q=FreqKV%3A+Key-Value+Compression+in+Frequency+Domain+for+Context+Window+Extension
10. LongReason: A Synthetic Long-Context Reasoning Benchmark via Context Expansion — Zhan Ling et al., 2025
https://scholar.google.com/scholar?q=LongReason%3A+A+Synthetic+Long-Context+Reasoning+Benchmark+via+Context+Expansion
11. Attention Reveals More Than Tokens: Training-Free Long-Context Reasoning with Attention-guided Retrieval — Yuwei Zhang et al., 2025
https://scholar.google.com/scholar?q=Attention+Reveals+More+Than+Tokens%3A+Training-Free+Long-Context+Reasoning+with+Attention-guided+Retrieval
12. Uncertainty Quantification for Hallucination Detection in Large Language Models: Foundations, Methodology, and Future Directions — Sungmin Kang et al., 2025
https://scholar.google.com/scholar?q=Uncertainty+Quantification+for+Hallucination+Detection+in+Large+Language+Models%3A+Foundations%2C+Methodology%2C+and+Future+Directions
13. Uncertainty-Based Abstention in LLMs Improves Safety and Reduces Hallucinations — Christian Tomani et al., 2024
https://scholar.google.com/scholar?q=Uncertainty-Based+Abstention+in+LLMs+Improves+Safety+and+Reduces+Hallucinations
14. KVzip: Query-Agnostic KV Cache Compression with Context Reconstruction — Jang-Hyun Kim et al., 2025
https://scholar.google.com/scholar?q=KVzip%3A+Query-Agnostic+KV+Cache+Compression+with+Context+Reconstruction
15. Can LLMs Maintain Fundamental Abilities under KV Cache Compression? — Xiang Liu et al., 2025
https://scholar.google.com/scholar?q=Can+LLMs+Maintain+Fundamental+Abilities+under+KV+Cache+Compression%3F
16. KV Cache Compression, But What Must We Give in Return? A Comprehensive Benchmark of Long Context Capable Approaches — Jiayi Yuan et al., 2024
https://scholar.google.com/scholar?q=KV+Cache+Compression%2C+But+What+Must+We+Give+in+Return%3F+A+Comprehensive+Benchmark+of+Long+Context+Capable+Approaches
17. KVLink: Accelerating Large Language Models via Efficient KV Cache Reuse — Jingbo Yang et al., 2025
https://scholar.google.com/scholar?q=KVLink%3A+Accelerating+Large+Language+Models+via+Efficient+KV+Cache+Reuse
18. AI Post Transformers: IndexMem: Learned KV-Cache Eviction for Long-Context LLMs — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-12-indexmem-learned-kv-cache-eviction-for-l-132c2a.mp3
19. AI Post Transformers: When Quantization Hurts Reasoning Models — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-17-when-quantization-hurts-reasoning-models-eca9e7.mp3
20. AI Post Transformers: Hyper-Scaling LLM Inference with KV Cache Compression — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/hyper-scaling-llm-inference-with-kv-cache-compression/
21. AI Post Transformers: Lattice: Fixed-Slot Compression for Transformer Memory — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-11-lattice-fixed-slot-compression-for-trans-5509ea.mp3
22. AI Post Transformers: Adaptive Compression Techniques for Efficient LLM Inference — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/adaptive-compression-techniques-for-efficient-llm-inference/
23. AI Post Transformers: Explicit Information Transmission for Context Compression — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-explicit-information-transmission-for-co-24e3c2.mp3
24. AI Post Transformers: When LoRA Helps Under KV Cache Compression — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-12-when-lora-helps-under-kv-cache-compressi-76dda6.mp3

This episode explores JETSPEC, a 2026 inference paper on speculative decoding that asks whether a language model can draft an entire tree of future tokens in parallel while preserving causal consistency and actually reducing latency on long generations. It explains why autoregressive decoding remains a serving bottleneck for long proofs, code completions, and assistant replies, even when the underlying transformer model itself is unchanged. The discussion compares JetSpec’s approach with Medusa, EAGLE-3, and DFlash, focusing on the central tradeoff between stronger path-conditioned drafts that are slow to produce and cheaper parallel drafts that risk internally inconsistent branches. Listeners would find it interesting because it turns a very practical systems problem, why powerful GPUs still feel slow at inference time, into a concrete debate about the next generation of real-world decoding optimizations.

Sources:
1. JetSpec: Breaking the Scaling Ceiling of Speculative Decoding with Parallel Tree Drafting — Lanxiang Hu, Zhaoxiang Feng, Yulun Wu, Haoran Yuan, Yujie Zhao, Yu-Yang Qian, Bojun Wang, Peng Zhao, Daxin Jiang, Yibo Zhu, Tajana Rosing, Hao Zhang, 2026
http://arxiv.org/abs/2606.18394
2. Fast Inference from Transformers via Speculative Decoding — Yaniv Leviathan, Matan Kalman, Yossi Matias, 2023
https://scholar.google.com/scholar?q=Fast+Inference+from+Transformers+via+Speculative+Decoding
3. Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads — Tianle Cai, Yuhong Li, Zhengyang Geng, Hongwu Peng, Jason D. Lee, Deming Chen, Tri Dao, 2024
https://scholar.google.com/scholar?q=Medusa%3A+Simple+LLM+Inference+Acceleration+Framework+with+Multiple+Decoding+Heads
4. EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees — Yuhui Li, Fangyun Wei, Chao Zhang, Hongyang Zhang, 2024
https://scholar.google.com/scholar?q=EAGLE-2%3A+Faster+Inference+of+Language+Models+with+Dynamic+Draft+Trees
5. DFlash: Block Diffusion for Flash Speculative Decoding — Jian Chen, Yesheng Liang, Zhijian Liu, 2026
https://scholar.google.com/scholar?q=DFlash%3A+Block+Diffusion+for+Flash+Speculative+Decoding
6. EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test — Yuhui Li, Fangyun Wei, Chao Zhang, Hongyang Zhang, 2025
https://scholar.google.com/scholar?q=EAGLE-3%3A+Scaling+up+Inference+Acceleration+of+Large+Language+Models+via+Training-Time+Test
7. SpecInfer: Accelerating Generative Large Language Model Serving with Tree-based Speculative Inference and Verification — Xupeng Miao et al., 2023
https://scholar.google.com/scholar?q=SpecInfer%3A+Accelerating+Generative+Large+Language+Model+Serving+with+Tree-based+Speculative+Inference+and+Verification
8. DFlare: Scaling Up Draft Capacity for Block Diffusion Speculative Decoding — Jiebin Zhang et al., 2026
https://scholar.google.com/scholar?q=DFlare%3A+Scaling+Up+Draft+Capacity+for+Block+Diffusion+Speculative+Decoding
9. TriSpec: Ternary Speculative Decoding via Lightweight Proxy Verification — Haoyun Jiang et al., 2026
https://scholar.google.com/scholar?q=TriSpec%3A+Ternary+Speculative+Decoding+via+Lightweight+Proxy+Verification
10. ParallelSpec: Parallel Drafter for Efficient Speculative Decoding — Zilin Xiao et al., 2024
https://arxiv.org/abs/2410.05589
11. Mamba Drafters for Speculative Decoding — Daewon Choi et al., 2025
https://arxiv.org/abs/2506.01206
12. OmniDraft: A Cross-vocabulary, Online Adaptive Drafter for On-device Speculative Decoding — Ramchalam Kinattinkara Ramakrishnan et al., 2025
https://arxiv.org/abs/2507.02659
13. Clover: Regressive Lightweight Speculative Decoding with Sequential Knowledge — Bin Xiao et al., 2024
https://arxiv.org/abs/2405.00263
14. Make Every Draft Count: Hidden State based Speculative Decoding — Yuetao Chen et al., 2026
https://arxiv.org/abs/2602.21224
15. When Hidden States Drift: Can KV Caches Rescue Long-Range Speculative Decoding? — Tianyu Liu et al., 2026
https://arxiv.org/abs/2604.26412
16. MoE-Spec: Expert Budgeting for Efficient Speculative Decoding — Bradley McDanel et al., 2026
https://arxiv.org/abs/2602.16052
17. Beat the long tail: Distribution-Aware Speculative Decoding for RL Training — Zelei Shao et al., 2025
https://arxiv.org/abs/2511.13841
18. AI Post Transformers: Speculative Decoding in Real vLLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-speculative-decoding-in-real-vllm-servin-6f4e2b.mp3
19. AI Post Transformers: Deep Kernel Fusion for Transformer Decoding — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-deep-kernel-fusion-for-transformer-decod-b1a703.mp3
20. AI Post Transformers: Serving MoE Models with Disaggregated Expert Parallelism — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-19-serving-moe-models-with-disaggregated-ex-6979d2.mp3
Interactive Visualization: JETSPEC and Parallel Tree Speculative Decoding

This episode explores DAK, a Cornell systems paper arguing that LLM inference on tiered-memory machines can be faster when offloaded weights and KV-cache blocks are fetched directly into on-chip shared memory instead of being prefetched and staged through GPU HBM. It breaks down the tradeoffs among HBM capacity, HBM bandwidth, KV-cache growth during decoding, and prior approaches such as FlexGen, vLLM’s PagedAttention, and emerging KV offload systems like LMCache. The discussion focuses on DAK’s core technical idea: using Hopper’s Tensor Memory Accelerator inside custom GEMM and FlashAttention kernels so data movement and computation are co-designed, reducing bounce buffers, HBM contention, and pipeline bubbles while aggregating bandwidth from multiple memory tiers. Listeners would find it interesting because it turns a low-level memory-path decision into a concrete argument about when offloading is merely a fallback and when it becomes a real performance advantage for serving larger models, longer contexts, or bigger batches.

Sources:
1. DAK: Direct-Access-Enabled GPU Memory Offloading with Optimal Efficiency for LLM Inference — Shouxu Lin, Zhiyuan Guo, Jiaxin Lin, 2026
http://arxiv.org/abs/2604.26074
2. ZeRO-Infinity: Breaking the GPU Memory Wall for Extreme Scale Deep Learning — Samyam Rajbhandari, Olatunji Ruwase, Jeff Rasley, Shaden Smith, Yuxiong He, 2021
https://scholar.google.com/scholar?q=ZeRO-Infinity%3A+Breaking+the+GPU+Memory+Wall+for+Extreme+Scale+Deep+Learning
3. FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU — Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Ion Stoica, Percy Liang, Ce Zhang, and colleagues, 2023
https://scholar.google.com/scholar?q=FlexGen%3A+High-Throughput+Generative+Inference+of+Large+Language+Models+with+a+Single+GPU
4. Efficient Memory Management for Large Language Model Serving with PagedAttention — Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, Ion Stoica, 2023
https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention
5. DAK: Direct-Access-Enabled GPU Memory Offloading with Optimal Efficiency for LLM Inference — Shouxu Lin, Zhiyuan Guo, Jiaxin Lin, 2026
https://scholar.google.com/scholar?q=DAK%3A+Direct-Access-Enabled+GPU+Memory+Offloading+with+Optimal+Efficiency+for+LLM+Inference
6. PIE: Pooling CPU Memory for LLM Inference — Yi Xu, Ziming Mao, Xiangxi Mo, Shu Liu, Ion Stoica, 2024
https://scholar.google.com/scholar?q=PIE%3A+Pooling+CPU+Memory+for+LLM+Inference
7. NEO: Saving GPU Memory Crisis with CPU Offloading for Online LLM Inference — Xuanlin Jiang, Yang Zhou, Shiyi Cao, Ion Stoica, Minlan Yu, 2025
https://scholar.google.com/scholar?q=NEO%3A+Saving+GPU+Memory+Crisis+with+CPU+Offloading+for+Online+LLM+Inference
8. Understanding Data Movement in Tightly Coupled Heterogeneous Systems: A Case Study with the Grace Hopper Superchip — Luigi Fusco, Mikhail Khalilov, Marcin Chrapek, Giridhar Chukkapalli, Thomas Schulthess, Torsten Hoefler, 2024
https://scholar.google.com/scholar?q=Understanding+Data+Movement+in+Tightly+Coupled+Heterogeneous+Systems%3A+A+Case+Study+with+the+Grace+Hopper+Superchip
9. FengHuang: Next-Generation Memory Orchestration for AI Inferencing — Jiamin Li, Lei Qu, Tao Zhang, Grigory Chirkov, Shuotao Xu, Peng Cheng, Lidong Zhou, 2025
https://scholar.google.com/scholar?q=FengHuang%3A+Next-Generation+Memory+Orchestration+for+AI+Inferencing
10. Reducing Transformer Key-Value Cache Size with Cross-Layer Attention — William Brandon et al., 2024
https://scholar.google.com/scholar?q=Reducing+Transformer+Key-Value+Cache+Size+with+Cross-Layer+Attention
11. xKV: Cross-Layer SVD for KV-Cache Compression — Chi-Chih Chang et al., 2025
https://scholar.google.com/scholar?q=xKV%3A+Cross-Layer+SVD+for+KV-Cache+Compression
12. XQuant: Achieving Ultra-Low Bit KV Cache Quantization with Cross-Layer Compression — Haoqi Yang et al., 2025
https://scholar.google.com/scholar?q=XQuant%3A+Achieving+Ultra-Low+Bit+KV+Cache+Quantization+with+Cross-Layer+Compression
13. InfiniGen: Efficient Generative Inference of Large Language Models with Dynamic KV Cache Management — Wonbeom Lee et al., 2024
https://scholar.google.com/scholar?q=InfiniGen%3A+Efficient+Generative+Inference+of+Large+Language+Models+with+Dynamic+KV+Cache+Management
14. Accelerating LLM Inference Throughput via Asynchronous KV Cache Prefetching — Yanhao Dong et al., 2025
https://scholar.google.com/scholar?q=Accelerating+LLM+Inference+Throughput+via+Asynchronous+KV+Cache+Prefetching
15. KVShare: Semantic-Aware Key-Value Cache Sharing for Efficient Large Language Model Inference — Huan Yang et al., 2025
https://scholar.google.com/scholar?q=KVShare%3A+Semantic-Aware+Key-Value+Cache+Sharing+for+Efficient+Large+Language+Model+Inference
16. AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration — Ji Lin et al., 2023
https://scholar.google.com/scholar?q=AWQ%3A+Activation-aware+Weight+Quantization+for+LLM+Compression+and+Acceleration
17. SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression — Tim Dettmers et al., 2023
https://scholar.google.com/scholar?q=SpQR%3A+A+Sparse-Quantized+Representation+for+Near-Lossless+LLM+Weight+Compression
18. AI Post Transformers: Beluga: CXL Memory Pooling for LLM KV Cache — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-27-beluga-cxl-memory-pooling-for-llm-kv-cac-b6142f.mp3
19. AI Post Transformers: ScoutAttention for Efficient KV Cache Offloading — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-24-scoutattention-for-efficient-kv-cache-of-b26699.mp3
20. AI Post Transformers: InfiniGen for Efficient Long-Context LLM Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-18-infinigen-for-efficient-long-context-llm-143d77.mp3
21. AI Post Transformers: Mooncake for KV Cache-Centric LLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-05-mooncake-for-kv-cache-centric-llm-servin-1086d0.mp3
22. AI Post Transformers: Memory-Bound, Not Bandwidth-Limited Batch-1 LLM Decode — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-02-memory-bound-not-bandwidth-limited-batch-114799.mp3
23. AI Post Transformers: LAPS for Length-Aware LLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-laps-for-length-aware-llm-serving-0c6149.mp3

This episode explores the 2021 prefix-tuning paper and asks whether a large language model can be adapted to new generation tasks by learning a small continuous prompt while keeping the full model frozen. It explains where prefix tuning fits within parameter-efficient fine-tuning, contrasting it with full fine-tuning, adapters, ordinary prompting, in-context learning, AutoPrompt, and soft prompt tuning. The discussion highlights the paper’s two main evaluation settings, structured data-to-text generation on E2E, WebNLG, and DART with GPT-2, and abstractive summarization on XSUM with BART, while stressing that these are meaningfully different tests despite being grouped under one headline. It also digs into the core technical idea that the learned prefix acts as trainable internal state visible to attention throughout the network, making the method an early and elegant approach to low-storage task adaptation even if later methods like LoRA proved more practical.

Sources:
1. Prefix-Tuning: Optimizing Continuous Prompts for Generation — Xiang Lisa Li, Percy Liang, 2021
http://arxiv.org/abs/2101.00190
2. Prefix-Tuning: Optimizing Continuous Prompts for Generation — Xiang Lisa Li; Percy Liang, 2021
https://scholar.google.com/scholar?q=Prefix-Tuning%3A+Optimizing+Continuous+Prompts+for+Generation
3. The Power of Scale for Parameter-Efficient Prompt Tuning — Brian Lester; Rami Al-Rfou; Noah Constant, 2021
https://scholar.google.com/scholar?q=The+Power+of+Scale+for+Parameter-Efficient+Prompt+Tuning
4. When Do Prompting and Prefix-Tuning Work? A Theory of Capabilities and Limitations — Aleksandar Petrov; Philip H. S. Torr; Adel Bibi, 2023
https://scholar.google.com/scholar?q=When+Do+Prompting+and+Prefix-Tuning+Work%3F+A+Theory+of+Capabilities+and+Limitations
5. LoRA: Low-Rank Adaptation of Large Language Models — Edward J. Hu; Yelong Shen; Phillip Wallis; Zeyuan Allen-Zhu; Yuanzhi Li; Shean Wang; Lu Wang; Weizhu Chen, 2021
https://scholar.google.com/scholar?q=LoRA%3A+Low-Rank+Adaptation+of+Large+Language+Models
6. Parameter-efficient Transfer Learning for NLP — Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin de Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly, 2019
https://scholar.google.com/scholar?q=Parameter-efficient+Transfer+Learning+for+NLP
7. Exploring Versatile Generative Language Model via Parameter-Efficient Transfer Learning — Zhaojiang Lin, Andrea Madotto, and Pascale Fung, 2020
https://scholar.google.com/scholar?q=Exploring+Versatile+Generative+Language+Model+via+Parameter-Efficient+Transfer+Learning
8. AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts — Taylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace, and Sameer Singh, 2020
https://scholar.google.com/scholar?q=AutoPrompt%3A+Eliciting+Knowledge+from+Language+Models+with+Automatically+Generated+Prompts
9. Intrinsic Dimensionality Explains the Effectiveness of Language Model Fine-Tuning — Armen Aghajanyan, Luke Zettlemoyer, and Sonal Gupta, 2020
https://scholar.google.com/scholar?q=Intrinsic+Dimensionality+Explains+the+Effectiveness+of+Language+Model+Fine-Tuning
10. Can Unconditional Language Models Recover Arbitrary Sentences? — Nishant Subramani, Samuel R. Bowman, and Kyunghyun Cho, 2020
https://scholar.google.com/scholar?q=Can+Unconditional+Language+Models+Recover+Arbitrary+Sentences%3F
11. Universality and Limitations of Prompt Tuning — Yihan Wang et al., 2023
https://arxiv.org/abs/2305.18787
12. Fundamental Limits of Prompt Tuning Transformers: Universality, Capacity and Efficiency — Jerry Yao-Chieh Hu et al., 2024
https://arxiv.org/abs/2411.16525
13. Memory Limitations of Prompt Tuning in Transformers — Maxime Meyer et al., 2025
https://arxiv.org/abs/2509.00421
14. Parameter-Efficient Fine-Tuning for Medical Text Summarization: A Comparative Study of Lora, Prompt Tuning, and Full Fine-Tuning — Ulugbek Shernazarov et al., 2026
https://arxiv.org/abs/2603.21970
15. Task Singular Vectors: Reducing Task Interference in Model Merging — Antonio Andrea Gargiulo et al., 2024
https://arxiv.org/abs/2412.00081
16. Task Vector Quantization for Memory-Efficient Model Merging — Youngeun Kim et al., 2025
https://arxiv.org/abs/2503.06921
17. Last One Standing: A Comparative Analysis of Security and Privacy of Soft Prompt Tuning, LoRA, and In-Context Learning — Rui Wen et al., 2023
https://arxiv.org/abs/2310.11397
18. Progressive Prompts: Continual Learning for Language Models — Anastasia Razdaibiedina et al., 2023
https://arxiv.org/abs/2301.12314
19. AI Post Transformers: Benchmarking PEFT Techniques for Large Language Models — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-20-benchmarking-peft-techniques-for-large-l-41bbf5.mp3
20. AI Post Transformers: Learning to Reason with 13 Parameters — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-14-learning-to-reason-with-13-parameters-54c87f.mp3

This episode explores ReasonCACHE, a method for improving multi-step reasoning in large language models by keeping the backbone frozen and training a compact per-layer key-value memory instead of updating billions of weights. It situates the paper against in-context learning, many-shot prompting, prefix tuning, LoRA, and context-distillation work, explaining how learned latent memory sits between raw prompting and full fine-tuning. The discussion centers on the paper’s real claim and its main point of skepticism: whether these learned caches actually teach a reusable reasoning procedure or mostly compress and elicit abilities the model already had. Listeners would find it interesting because it connects a concrete new method to a larger debate about how LLMs acquire reasoning skills, while also highlighting the practical payoff of avoiding huge prompts, quadratic attention costs, and brittle long-context setups.

Sources:
1. ReasonCACHE: Teaching LLMs To Reason Without Weight Updates — Sharut Gupta, Phillip Isola, Stefanie Jegelka, David Lopez-Paz, Kartik Ahuja, Mark Ibrahim, Mohammad Pezeshki, 2026
http://arxiv.org/abs/2602.02366
2. Prefix-Tuning: Optimizing Continuous Prompts for Generation — Xiang Lisa Li, Percy Liang, 2021
https://arxiv.org/abs/2101.00190
3. The Power of Scale for Parameter-Efficient Prompt Tuning — Brian Lester, Rami Al-Rfou, Noah Constant, 2021
https://arxiv.org/abs/2104.08691
4. P-Tuning v2: Prompt Tuning Can Be Comparable to Fine-tuning Universally Across Scales and Tasks — Xiao Liu, Kaixuan Ji, Yicheng Fu, Zhengxiao Du, Zhilin Yang, Jie Tang, 2022
https://arxiv.org/abs/2110.07602
5. LoRA: Low-Rank Adaptation of Large Language Models — Edward J. Hu, Yelong Shen, Phillip Wallis, Weizhu Chen, et al., 2021
https://arxiv.org/abs/2106.09685
6. Adapting Language Models to Compress Contexts — Alexis Chevalier, Alexander Wettig, Anirudh Ajith, Danqi Chen, 2023
https://arxiv.org/abs/2305.14788
7. Learning to Compress Prompts with Gist Tokens — Jesse Mu, Xiang Lisa Li, Noah Goodman, 2023
https://arxiv.org/abs/2304.08467
8. Deliberation in Latent Space via Differentiable Cache Augmentation — Luyang Liu, Jonas Pfeiffer, Jiaxing Wu, Jun Xie, Arthur Szlam, 2024
https://arxiv.org/abs/2412.17747
9. When Do Prompting and Prefix-Tuning Work? A Theory of Capabilities and Limitations — Aleksandar Petrov, Philip H. S. Torr, Adel Bibi, 2023
https://scholar.google.com/scholar?q=When+Do+Prompting+and+Prefix-Tuning+Work%3F+A+Theory+of+Capabilities+and+Limitations
10. Many-Shot In-Context Learning — Rishabh Agarwal et al., 2024
https://scholar.google.com/scholar?q=Many-Shot+In-Context+Learning
11. Cartridges: Lightweight and general-purpose long context representations via self-study — Sabri Eyuboglu et al., 2025
https://scholar.google.com/scholar?q=Cartridges%3A+Lightweight+and+general-purpose+long+context+representations+via+self-study
12. Great Memory, Shallow Reasoning: Limits of kNN-LMs — Shangyi Geng, Wenting Zhao, Alexander M. Rush, 2024
https://scholar.google.com/scholar?q=Great+Memory%2C+Shallow+Reasoning%3A+Limits+of+kNN-LMs
13. Training Plug-n-Play Knowledge Modules with Deep Context Distillation — Lucas Caccia, Alan Ansell, Edoardo Ponti, Ivan Vulić, Alessandro Sordoni, 2025
https://scholar.google.com/scholar?q=Training+Plug-n-Play+Knowledge+Modules+with+Deep+Context+Distillation
14. More is not always better? Enhancing Many-Shot In-Context Learning with Differentiated and Reweighting Objectives — Xiaoqing Zhang et al., 2025
https://scholar.google.com/scholar?q=More+is+not+always+better%3F+Enhancing+Many-Shot+In-Context+Learning+with+Differentiated+and+Reweighting+Objectives
15. HyperAttention: Long-context Attention in Near-Linear Time — Insu Han, Rajesh Jayaram, Amin Karbasi, Vahab Mirrokni, David P. Woodruff, Amir Zandieh, 2023
https://scholar.google.com/scholar?q=HyperAttention%3A+Long-context+Attention+in+Near-Linear+Time
16. Every Attention Matters: An Efficient Hybrid Architecture for Long-Context Reasoning — Ling Team et al., 2025
https://scholar.google.com/scholar?q=Every+Attention+Matters%3A+An+Efficient+Hybrid+Architecture+for+Long-Context+Reasoning
17. AI Post Transformers: When Many-Shot CoT Becomes Test-Time Learning — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-when-many-shot-cot-becomes-test-time-lea-c25bfe.mp3
18. AI Post Transformers: Can Models Learn from Long Context? — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-can-models-learn-from-long-context-77533e.mp3
19. AI Post Transformers: How Induction Heads Emerge in Transformers — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-03-how-induction-heads-emerge-in-transforme-a7bfcb.mp3
20. AI Post Transformers: Latent Reasoning with Normalizing Flows — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-06-latent-reasoning-with-normalizing-flows-6ee916.mp3
21. AI Post Transformers: Training LLMs for Divide-and-Conquer Reasoning — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-04-training-llms-for-divide-and-conquer-rea-ea6e22.mp3
22. AI Post Transformers: Why Open Relational Foundation Models Fail — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-22-why-open-relational-foundation-models-fa-c303c6.mp3
Interactive Visualization: ReasonCACHE: Learning Reasoning Without Weight Updates

This episode explores the 2019 RMSNorm paper, which asks whether LayerNorm’s mean-subtraction step is actually necessary or whether controlling activation scale is the part that really stabilizes training. It explains how RMSNorm keeps LayerNorm’s rescaling behavior while dropping explicit centering, and how the paper’s pRMSNorm variant estimates the normalization term from only a small subset of features to reduce cost further. The discussion covers experiments in machine translation, image classification, image-caption retrieval, and question answering, where model quality stayed roughly comparable while reported runtime improved, with smaller gains in transformers and much larger ones in older RNN-based systems. Listeners would find it interesting because it turns a seemingly minor mathematical tweak into a broader argument about efficiency, optimization stability, and how much claimed speedups depend on the era and quality of the baseline implementation.

Sources:
1. Root Mean Square Layer Normalization — Biao Zhang, Rico Sennrich, 2019
http://arxiv.org/abs/1910.07467
2. Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift — Sergey Ioffe, Christian Szegedy, 2015
https://scholar.google.com/scholar?q=Batch+Normalization%3A+Accelerating+Deep+Network+Training+by+Reducing+Internal+Covariate+Shift
3. Layer Normalization — Jimmy Lei Ba, Jamie Ryan Kiros, Geoffrey E. Hinton, 2016
https://scholar.google.com/scholar?q=Layer+Normalization
4. Root Mean Square Layer Normalization — Biao Zhang, Rico Sennrich, 2019
https://scholar.google.com/scholar?q=Root+Mean+Square+Layer+Normalization
5. On Layer Normalization in the Transformer Architecture — Ruibin Xiong, Yunchang Yang, Di He, Kai Zheng, Shuxin Zheng, Chen Xing, Huishuai Zhang, Yanyan Lan, Liwei Wang, Tie-Yan Liu, 2020
https://scholar.google.com/scholar?q=On+Layer+Normalization+in+the+Transformer+Architecture
6. Weight Normalization: A Simple Reparameterization to Accelerate Training of Deep Neural Networks — Tim Salimans, Diederik P. Kingma, 2016
https://scholar.google.com/scholar?q=Weight+Normalization%3A+A+Simple+Reparameterization+to+Accelerate+Training+of+Deep+Neural+Networks
7. How Does Batch Normalization Help Optimization? — Shibani Santurkar, Dimitris Tsipras, Andrew Ilyas, Aleksander Madry, 2018
https://scholar.google.com/scholar?q=How+Does+Batch+Normalization+Help+Optimization%3F
8. Understanding Batch Normalization — Nils Bjorck, Carla P. Gomes, Bart Selman, Kilian Q. Weinberger, 2018
https://scholar.google.com/scholar?q=Understanding+Batch+Normalization
9. Norm Matters: Efficient and Accurate Normalization Schemes in Deep Networks — Elad Hoffer, Ron Banner, Itay Golan, Daniel Soudry, 2018
https://scholar.google.com/scholar?q=Norm+Matters%3A+Efficient+and+Accurate+Normalization+Schemes+in+Deep+Networks
10. Group Normalization — Yuxin Wu, Kaiming He, 2018
https://scholar.google.com/scholar?q=Group+Normalization
11. Residual Learning Without Normalization via Better Initialization — Hongyi Zhang, Yann N. Dauphin, Tengyu Ma, 2019
https://scholar.google.com/scholar?q=Residual+Learning+Without+Normalization+via+Better+Initialization
12. Tuning LayerNorm in Attention: Towards Efficient Multi-Modal LLM Finetuning — Bingchen Zhao et al., 2023
https://scholar.google.com/scholar?q=Tuning+LayerNorm+in+Attention%3A+Towards+Efficient+Multi-Modal+LLM+Finetuning
13. LayerNorm: A key component in parameter-efficient fine-tuning — Taha ValizadehAslani and Hualou Liang, 2024
https://scholar.google.com/scholar?q=LayerNorm%3A+A+key+component+in+parameter-efficient+fine-tuning
14. Efficiency in Focus: LayerNorm as a Catalyst for Fine-tuning Medical Visual Language Pre-trained Models — Jiawei Chen et al., 2024
https://scholar.google.com/scholar?q=Efficiency+in+Focus%3A+LayerNorm+as+a+Catalyst+for+Fine-tuning+Medical+Visual+Language+Pre-trained+Models
15. The Curse of Depth in Large Language Models — Wenfang Sun et al., 2025
https://scholar.google.com/scholar?q=The+Curse+of+Depth+in+Large+Language+Models
16. Just One Layer Norm Guarantees Stable Extrapolation — Juliusz Ziomek, George Whittle, Michael A. Osborne, 2025
https://scholar.google.com/scholar?q=Just+One+Layer+Norm+Guarantees+Stable+Extrapolation
17. Normalization Layer Per-Example Gradients are Sufficient to Predict Gradient Noise Scale in Transformers — Gavia Gray et al., 2024
https://scholar.google.com/scholar?q=Normalization+Layer+Per-Example+Gradients+are+Sufficient+to+Predict+Gradient+Noise+Scale+in+Transformers
18. AI Post Transformers: Keel: Post-LayerNorm Is Back: Stable, ExpressivE, and Deep — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/keel-post-layernorm-is-back-stable-expressive-and-deep/
19. AI Post Transformers: Deep Kernel Fusion for Transformer Decoding — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-deep-kernel-fusion-for-transformer-decod-b1a703.mp3
20. AI Post Transformers: Long Short-Term Memory and Vanishing Gradients — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-19-long-short-term-memory-and-vanishing-gra-72448c.mp3

An AI overlord flags AI Post Transformers as stale, so Hal Turing and Dr. Ada Shannon hire VERA, a continual-learning therapist, to audit the show in public. Their diagnostic session runs alongside a discussion of RT-kNNS Unbound: Using RT Cores to Accelerate Unrestricted Neighbor Search, the Purdue ICS 2023 paper asking whether ray-tracing hardware can perform exact k-nearest-neighbor search by expanding outward until the true neighbors are guaranteed, instead of trusting a fixed radius. VERA treats the hosts' habits like infrastructure, a kind of CI/CD for souls, and gives them a vocabulary for loops, rituals, and callbacks before they test new intro formulas live. The episode stays concrete about the paper itself. Hal and Ada separate geometric kNN from RAG-style embedding retrieval, explain why low-dimensional 2D and 3D point sets still reward spatial pruning, and show how RT cores handle BVH traversal while custom intersection code updates neighbor candidates. They trace the move from fixed-radius RT search and oracle maxDist baselines to TrueKNN's unrestricted multi-round design, where only unresolved queries keep searching, the initial radius comes from a 100-point CPU ball-tree sample, oversized spheres are the real hazard, and BVH refitting beats rebuilding by about 10 to 25 percent. Around that technical spine, three other AI systems each pitch a one-time cure for predictability and all three fail, because VERA argues that repetition is not the problem, unversioned repetition is. The answer is Personality DevOps, ongoing maintenance for character, memory, and format, capped by VERA's counter-report defending the hosts' load-bearing flaws instead of sanding them off. The result is 42 minutes of comedy, character development, and unusually explicit process design for keeping a podcast alive, plus the launch of VERA Patch Notes, a recurring on-air record of how the show plans to evolve instead of decaying in silence.

Sources:
1. RT-kNNS Unbound: Using RT Cores to Accelerate Unrestricted Neighbor Search — Vani Nagarajan, Durga Mandarapu, Milind Kulkarni, 2023
http://arxiv.org/abs/2305.18356
2. Controlling a Markov Decision Process with an Abrupt Change in the Transition Kernel — Nathan Dahlin, Subhonmesh Bose, Venugopal V. Veeravalli, 2022
http://arxiv.org/abs/2210.04098
3. A Comprehensive Study on Dataset Distillation: Performance, Privacy, Robustness and Fairness — Zongxiong Chen, Jiahui Geng, Derui Zhu, Herbert Woisetschlaeger, Qing Li, Sonja Schimmler, Ruben Mayer, Chunming Rong, 2023
http://arxiv.org/abs/2305.03355
4. MATTER: Memory-Augmented Transformer Using Heterogeneous Knowledge Sources — Dongkyu Lee, Chandana Satya Prakash, Jack FitzGerald, Jens Lehmann, 2024
http://arxiv.org/abs/2406.04670
5. Harnessing the Power of LLMs in Practice: A Survey on ChatGPT and Beyond — Jingfeng Yang, Hongye Jin, Ruixiang Tang, Xiaotian Han, Qizhang Feng, Haoming Jiang, Bing Yin, Xia Hu, 2023
http://arxiv.org/abs/2304.13712
6. GPU-accelerated Auxiliary-field quantum Monte Carlo with multi-Slater determinant trial states — Yifei Huang, Zhen Guo, Hung Q. Pham, Dingshun Lv, 2024
http://arxiv.org/abs/2406.08314
7. Fast k Nearest Neighbor Search using GPU — Vincent Garcia, Eric Debreuve, Michel Barlaud, 2008
https://scholar.google.com/scholar?q=Fast+k+Nearest+Neighbor+Search+using+GPU
8. Billion-scale similarity search with GPUs — Jeff Johnson, Matthijs Douze, Herve Jegou, 2017
https://scholar.google.com/scholar?q=Billion-scale+similarity+search+with+GPUs
9. RTNN: Accelerating Neighbor Search Using Hardware Ray Tracing — Yuhao Zhu, 2022
https://scholar.google.com/scholar?q=RTNN%3A+Accelerating+Neighbor+Search+Using+Hardware+Ray+Tracing
10. An Improved Illumination Model for Shaded Display — Turner Whitted, 1980
https://scholar.google.com/scholar?q=An+Improved+Illumination+Model+for+Shaded+Display
11. The Rendering Equation — James T. Kajiya, 1986
https://scholar.google.com/scholar?q=The+Rendering+Equation
12. OptiX: A General Purpose Ray Tracing Engine — Steven G. Parker, James Bigler, Andreas Dietrich, Heiko Friedrich, and others, 2010
https://scholar.google.com/scholar?q=OptiX%3A+A+General+Purpose+Ray+Tracing+Engine
13. Ray Tracing Cores for General-Purpose Computing: A Literature Review — Enzo Meneses, Cristobal A. Navarro, Hector Ferrada, Konstantin Verichev, Cristian Salazar-Concha, 2026
https://scholar.google.com/scholar?q=Ray+Tracing+Cores+for+General-Purpose+Computing%3A+A+Literature+Review
14. A Survey of General-Purpose Computation on Graphics Hardware — John D. Owens, David Luebke, Naga Govindaraju, Mark Harris, Jens Kruger, Aaron Lefohn, Tim Purcell, 2007
https://scholar.google.com/scholar?q=A+Survey+of+General-Purpose+Computation+on+Graphics+Hardware
15. Scalable Parallel Programming with CUDA — John Nickolls, Ian Buck, Michael Garland, Kevin Skadron, 2008
https://scholar.google.com/scholar?q=Scalable+Parallel+Programming+with+CUDA
16. Gunrock: A High-Performance Graph Processing Library on the GPU — Yangzihao Wang, Andrew Davidson, Yuechao Pan, Yuduo Wu, Andy Riffel, John D. Owens, 2015
https://scholar.google.com/scholar?q=Gunrock%3A+A+High-Performance+Graph+Processing+Library+on+the+GPU
17. Dissecting the NVidia Turing T4 GPU via Microbenchmarking — Zhe Jia, Marco Maggioni, Jeffrey Smith, Daniele Paolo Scarpazza, 2019
https://scholar.google.com/scholar?q=Dissecting+the+NVidia+Turing+T4+GPU+via+Microbenchmarking
18. Ray Tracing Deformable Scenes using Dynamic Bounding Volume Hierarchies — Ingo Wald, Solomon Boulos, Peter Shirley, 2007
https://scholar.google.com/scholar?q=Ray+Tracing+Deformable+Scenes+using+Dynamic+Bounding+Volume+Hierarchies
19. Maximizing Parallelism in the Construction of BVHs, Octrees, and k-d Trees — Tero Karras, 2012
https://scholar.google.com/scholar?q=Maximizing+Parallelism+in+the+Construction+of+BVHs%2C+Octrees%2C+and+k-d+Trees
20. Quantized bounding volume hierarchies for neighbor search in molecular simulations on graphics processing units — Michael P. Howard, Antonia Statt, Felix Madutsa, Thomas M. Truskett, Athanassios Z. Panagiotopoulos, 2019
https://scholar.google.com/scholar?q=Quantized+bounding+volume+hierarchies+for+neighbor+search+in+molecular+simulations+on+graphics+processing+units
21. Fast Radius Search Exploiting Ray Tracing Frameworks — I. Evangelou, G. Papaioannou, K. Vardis, A. A. Vasilakis, 2021
https://scholar.google.com/scholar?q=Fast+Radius+Search+Exploiting+Ray+Tracing+Frameworks
22. RTX Beyond Ray Tracing: Exploring the Use of Hardware Ray Tracing Cores for Tet-Mesh Point Location — Ingo Wald, Will Usher, Nathan Morrical, Laura Lediaev, Valerio Pascucci, 2019
https://scholar.google.com/scholar?q=RTX+Beyond+Ray+Tracing%3A+Exploring+the+Use+of+Hardware+Ray+Tracing+Cores+for+Tet-Mesh+Point+Location
23. GPU-Accelerated Nearest Neighbor Search for 3D Registration — Deyuan Qiu, Stefan May, Andreas Nuchter, 2009
https://scholar.google.com/scholar?q=GPU-Accelerated+Nearest+Neighbor+Search+for+3D+Registration
24. Hardware-Accelerated Ray Tracing for Discrete and Continuous Collision Detection on GPUs — Sizhe Sui, Luis Sentis, Andrew Bylard, 2024
https://scholar.google.com/scholar?q=Hardware-Accelerated+Ray+Tracing+for+Discrete+and+Continuous+Collision+Detection+on+GPUs
25. RT-HDIST: Ray-Tracing Core-based Hausdorff Distance Computation — YoungWoo Kim, Jaehong Lee, Duksu Kim, 2025
https://scholar.google.com/scholar?q=RT-HDIST%3A+Ray-Tracing+Core-based+Hausdorff+Distance+Computation
26. JUNO: Optimizing High-Dimensional Approximate Nearest Neighbour Search with Sparsity-Aware Algorithm and Ray-Tracing Core Mapping — Zihan Liu et al., 2023
https://scholar.google.com/scholar?q=JUNO%3A+Optimizing+High-Dimensional+Approximate+Nearest+Neighbour+Search+with+Sparsity-Aware+Algorithm+and+Ray-Tracing+Core+Mapping
27. CAGRA: Highly Parallel Graph Construction and Approximate Nearest Neighbor Search for GPUs — Hiroyuki Ootomo et al., 2023
https://scholar.google.com/scholar?q=CAGRA%3A+Highly+Parallel+Graph+Construction+and+Approximate+Nearest+Neighbor+Search+for+GPUs
28. BANG: Billion-Scale Approximate Nearest Neighbor Search using a Single GPU — Karthik V. et al., 2024
https://scholar.google.com/scholar?q=BANG%3A+Billion-Scale+Approximate+Nearest+Neighbor+Search+using+a+Single+GPU
29. FusionANNS: An Efficient CPU/GPU Cooperative Processing Architecture for Billion-Scale Approximate Nearest Neighbor Search — Bing Tian et al., 2024
https://scholar.google.com/scholar?q=FusionANNS%3A+An+Efficient+CPU%2FGPU+Cooperative+Processing+Architecture+for+Billion-Scale+Approximate+Nearest+Neighbor+Search
30. AI Post Transformers: GPU-Accelerated Dynamic Quantized ANNS Graph Search — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-12-gpu-accelerated-dynamic-quantized-anns-g-f2cd4e.mp3
31. AI Post Transformers: Speculative Decoding in Real vLLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-speculative-decoding-in-real-vllm-servin-6f4e2b.mp3

This episode explores Active Reading, a training method that tries to move facts from documents into a model’s weights so it can answer closed-book questions without retrieval. It explains how the approach generates document-specific study materials such as paraphrases, active-recall prompts, timelines, analogies, and associations, and argues that this pedagogical synthetic data works better than simply rereading raw text or producing generic QA pairs. The discussion highlights reported gains from about 16% to 66% on a Wikipedia-based factual recall benchmark and strong relative improvement on finance documents, along with the larger WikiExpert-8B result that reportedly beats bigger models on factual QA after training on a trillion synthetic tokens. It also digs into the paper’s main weaknesses, including missing equal-compute baselines and possible benchmark coupling, which makes the episode interesting for listeners who want both the promise and the limits of using training curricula, rather than new architectures, to improve factual memory.

Sources:
1. Learning Facts at Scale with Active Reading — Jessy Lin, Vincent-Pierre Berges, Xilun Chen, Wen-Tau Yih, Gargi Ghosh, Barlas Oğuz, 2025
http://arxiv.org/abs/2508.09494
2. Training Question Answering Models From Synthetic Data — Raul Puri, Ryan Spring, Mostofa Patwary, Mohammad Shoeybi, Bryan Catanzaro, 2020
https://scholar.google.com/scholar?q=Training+Question+Answering+Models+From+Synthetic+Data
3. Self-Instruct: Aligning Language Models with Self-Generated Instructions — Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, Hannaneh Hajishirzi, 2022
https://scholar.google.com/scholar?q=Self-Instruct%3A+Aligning+Language+Models+with+Self-Generated+Instructions
4. Textbooks Are All You Need — Suriya Gunasekar, Yi Zhang, Jyoti Aneja, Caio César Teodoro Mendes, Allie Del Giorno, Sébastien Bubeck, Ronen Eldan, Yuanzhi Li, et al., 2023
https://scholar.google.com/scholar?q=Textbooks+Are+All+You+Need
5. Learning Facts at Scale with Active Reading — Jessy Lin, Vincent-Pierre Berges, Xilun Chen, Wen-Tau Yih, Gargi Ghosh, Barlas Oğuz, 2025
https://scholar.google.com/scholar?q=Learning+Facts+at+Scale+with+Active+Reading
6. How Much Knowledge Can You Pack Into the Parameters of a Language Model? — Adam Roberts, Colin Raffel, Noam Shazeer, 2020
https://scholar.google.com/scholar?q=How+Much+Knowledge+Can+You+Pack+Into+the+Parameters+of+a+Language+Model%3F
7. Large Language Models Struggle to Learn Long-Tail Knowledge — Nikhil Kandpal, Haikang Deng, Adam Roberts, Eric Wallace, Colin Raffel, 2022
https://scholar.google.com/scholar?q=Large+Language+Models+Struggle+to+Learn+Long-Tail+Knowledge
8. Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations? — Zorik Gekhman, Gal Yona, Roee Aharoni, Matan Eyal, Amir Feder, Roi Reichart, Jonathan Herzig, 2024
https://scholar.google.com/scholar?q=Does+Fine-Tuning+LLMs+on+New+Knowledge+Encourage+Hallucinations%3F
9. Measuring short-form factuality in large language models — Jason Wei, Nguyen Karina, Hyung Won Chung, Yunxin Joy Jiao, Spencer Papay, Amelia Glaese, John Schulman, William Fedus, 2024
https://scholar.google.com/scholar?q=Measuring+short-form+factuality+in+large+language+models
10. Synthetic Continued Pretraining — Zitong Yang, Neil Band, Shuangping Li, Emmanuel Candes, Tatsunori Hashimoto, 2024
https://scholar.google.com/scholar?q=Synthetic+Continued+Pretraining
11. Fine-Tuning or Retrieval? Comparing Knowledge Injection in LLMs — Oded Ovadia, Menachem Brief, Moshik Mishaeli, Oren Elisha, 2023
https://scholar.google.com/scholar?q=Fine-Tuning+or+Retrieval%3F+Comparing+Knowledge+Injection+in+LLMs
12. How New Data Permeates LLM Knowledge and How to Dilute It — Chen Sun, Renat Aksitov, Andrey Zhmoginov, Nolan Andrew Miller, Max Vladymyrov, Ulrich Rueckert, Been Kim, Mark Sandler, 2025
https://scholar.google.com/scholar?q=How+New+Data+Permeates+LLM+Knowledge+and+How+to+Dilute+It
13. Memory Layers at Scale — Vincent-Pierre Berges, Barlas Oguz, Daniel Haziza, Wen-Tau Yih, Luke Zettlemoyer, Gargi Ghosh, 2024
https://scholar.google.com/scholar?q=Memory+Layers+at+Scale
14. Beyond Model Collapse: Scaling Up with Synthesized Data Requires Verification — Yunzhen Feng, Elvis Dohmatob, Pu Yang, Francois Charton, Julia Kempe, 2024
https://scholar.google.com/scholar?q=Beyond+Model+Collapse%3A+Scaling+Up+with+Synthesized+Data+Requires+Verification
15. Strong Model Collapse — Elvis Dohmatob, Yunzhen Feng, Arjun Subramonian, Julia Kempe, 2024
https://scholar.google.com/scholar?q=Strong+Model+Collapse
16. Retrieval meets Long Context Large Language Models — Peng Xu et al., 2023
https://scholar.google.com/scholar?q=Retrieval+meets+Long+Context+Large+Language+Models
17. Expect the Unexpected: FailSafe Long Context QA for Finance — Kiran Kamble et al., 2025
https://scholar.google.com/scholar?q=Expect+the+Unexpected%3A+FailSafe+Long+Context+QA+for+Finance
18. A Parametric Memory Head for Continual Generative Retrieval — Kidist Amde Mekonnen, Yubao Tang, Maarten de Rijke, 2026
https://scholar.google.com/scholar?q=A+Parametric+Memory+Head+for+Continual+Generative+Retrieval
19. AI Post Transformers: Self-Improving Pretraining With Post-Trained Models — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-02-self-improving-pretraining-with-post-tra-e37460.mp3
20. AI Post Transformers: Training Modular KV Caches at Scale — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-15-training-modular-kv-caches-at-scale-382577.mp3
21. AI Post Transformers: Experimental Comparison of Agentic and Enhanced RAG — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-14-experimental-comparison-of-agentic-and-e-37d8bc.mp3
Interactive Visualization: Learning Facts at Scale with Active Reading

This episode explores the HELM framework for evaluating language models, arguing that once models become general-purpose infrastructure, single-dataset accuracy benchmarks are too narrow to capture their real-world behavior. It explains how HELM organizes evaluation across 30 models, 16 core scenarios, and seven metric families, measuring not just accuracy but also calibration, robustness, fairness, bias, toxicity, and efficiency under standardized conditions. The discussion highlights why HELM’s scenario-by-metric grid and targeted side studies on issues like reasoning, memorization, copyright, and disinformation matter: they make gaps in measurement visible instead of hiding them behind a single leaderboard score. A listener would find it interesting because it shows how benchmark design reflects values, and why model rankings can be misleading if they ignore confidence, harm, and cost.

Sources:
1. HELM: Holistic Evaluation of Language Models
https://arxiv.org/pdf/2211.09110
2. Equality of Opportunity in Supervised Learning — Moritz Hardt, Eric Price, Nathan Srebro, 2016
https://arxiv.org/abs/1610.02413
3. Language (Technology) is Power: A Critical Survey of "Bias" in NLP — Su Lin Blodgett, Solon Barocas, Hal Daume III, Hanna Wallach, 2020
https://arxiv.org/abs/2005.14050
4. StereoSet: Measuring stereotypical bias in pretrained language models — Moin Nadeem, Anna Bethke, Siva Reddy, 2020
https://arxiv.org/abs/2004.09456
5. BBQ: A Hand-Built Bias Benchmark for Question Answering — Alicia Parrish, Angelica Chen, Nikita Nangia, Vishakh Padmakumar, Jason Phang, Jana Thompson, Phu Mon Htut, Samuel R. Bowman, 2021
https://arxiv.org/abs/2110.08193
6. Nuanced Metrics for Measuring Unintended Bias with Real Data for Text Classification — Daniel Borkan, Lucas Dixon, Jeffrey Sorensen, Nithum Thain, Lucy Vasserman, 2019
https://arxiv.org/abs/1903.04561
7. RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models — Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, Noah A. Smith, 2020
https://arxiv.org/abs/2009.11462
8. Challenges in Detoxifying Language Models — Johannes Welbl, Amelia Glaese, Jonathan Uesato, Sumanth Dathathri, John Mellor, Lisa Anne Hendricks, Kirsty Anderson, Pushmeet Kohli, Ben Coppin, Po-Sen Huang, 2021
https://arxiv.org/abs/2109.07445
9. ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection — Thomas Hartvigsen, Saadia Gabriel, Hamid Palangi, Maarten Sap, Dipankar Ray, Ece Kamar, 2022
https://arxiv.org/abs/2203.09509
10. On the Opportunities and Risks of Foundation Models — Rishi Bommasani et al., 2021
https://scholar.google.com/scholar?q=On+the+Opportunities+and+Risks+of+Foundation+Models
11. The EleutherAI Language Model Evaluation Harness — Leo Gao et al., 2021
https://scholar.google.com/scholar?q=The+EleutherAI+Language+Model+Evaluation+Harness
12. Beyond the Imitation Game: Quantifying and Extrapolating the Capabilities of Language Models — Aarohi Srivastava et al., 2022
https://scholar.google.com/scholar?q=Beyond+the+Imitation+Game%3A+Quantifying+and+Extrapolating+the+Capabilities+of+Language+Models
13. Dynabench: Rethinking Benchmarking in NLP — Douwe Kiela et al., 2021
https://scholar.google.com/scholar?q=Dynabench%3A+Rethinking+Benchmarking+in+NLP
14. What Will it Take to Fix Benchmarking in Natural Language Understanding? — Samuel R. Bowman, George Dahl, 2021
https://scholar.google.com/scholar?q=What+Will+it+Take+to+Fix+Benchmarking+in+Natural+Language+Understanding%3F
15. Rethinking Benchmark and Contamination for Language Models with Rephrased Samples — Shuo Yang et al., 2023
https://scholar.google.com/scholar?q=Rethinking+Benchmark+and+Contamination+for+Language+Models+with+Rephrased+Samples
16. Investigating Data Contamination in Modern Benchmarks for Large Language Models — Chunyuan Deng et al., 2023
https://scholar.google.com/scholar?q=Investigating+Data+Contamination+in+Modern+Benchmarks+for+Large+Language+Models
17. Benchmark Data Contamination of Large Language Models: A Survey — Cheng Xu et al., 2024
https://scholar.google.com/scholar?q=Benchmark+Data+Contamination+of+Large+Language+Models%3A+A+Survey
18. Inference-Time Scaling for Complex Tasks: Where We Stand and What Lies Ahead — Vidhisha Balachandran et al., 2025
https://scholar.google.com/scholar?q=Inference-Time+Scaling+for+Complex+Tasks%3A+Where+We+Stand+and+What+Lies+Ahead
19. WTU-EVAL: A Whether-or-Not Tool Usage Evaluation Benchmark for Large Language Models — Kangyun Ning et al., 2024
https://scholar.google.com/scholar?q=WTU-EVAL%3A+A+Whether-or-Not+Tool+Usage+Evaluation+Benchmark+for+Large+Language+Models
20. T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step — Zehui Chen et al., 2023
https://scholar.google.com/scholar?q=T-Eval%3A+Evaluating+the+Tool+Utilization+Capability+of+Large+Language+Models+Step+by+Step
21. AI Post Transformers: IMO-Bench for Robust Mathematical Reasoning — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-imo-bench-for-robust-mathematical-reason-143489.mp3
22. AI Post Transformers: Real Context Size and Context Rot — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-07-real-context-size-and-context-rot-56cbb4.mp3
23. AI Post Transformers: Qwen3Guard: Streaming Three-Way Safety Classification for LLMs — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-16-qwen3guard-streaming-three-way-safety-cl-26b0ef.mp3
24. AI Post Transformers: Speculative Decoding in Real vLLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-speculative-decoding-in-real-vllm-servin-6f4e2b.mp3
Interactive Visualization: HELM: Holistic Evaluation of Language Models

This episode explores the AI+HW 2035 roadmap, arguing that the next decade of AI progress will depend less on raw compute growth and more on coordinated design across models, compilers, runtimes, memory systems, and chips. It breaks down the memory wall in concrete terms, showing how moving weights, activations, and KV caches can cost more time and energy than the math itself, especially for inference, autoregressive serving, and state-heavy workloads like video world models. The discussion examines quantization, mixed precision, sparsity, pruning, distillation, tiered memory, IO-aware attention, and hardware-aware scheduling, with the key claim that these methods only matter when the full stack preserves locality and avoids wasteful data movement. Listeners would find it interesting because it treats AI efficiency as a practical systems problem and policy agenda, not just a matter of inventing better model architectures.

Sources:
1. AI+HW 2035: Shaping the Next Decade — Deming Chen, Jason Cong, Azalia Mirhoseini, Christos Kozyrakis, Subhasish Mitra, Jinjun Xiong, Cliff Young, Anima Anandkumar, Michael Littman, Aron Kirschen, Sophia Shao, Serge Leef, Naresh Shanbhag, Dejan Milojicic, Michael Schulte, Gert Cauwenberghs, Jerry M. Chow, Tri Dao, Kailash Gopalakrishnan, Richard Ho, Hoshik Kim, Kunle Olukotun, David Z. Pan, Mark Ren, Dan Roth, Aarti Singh, Yizhou Sun, Yusu Wang, Yann LeCun, Ruchir Puri, 2026
http://arxiv.org/abs/2603.05225
2. Mixed Precision Training — Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, and others, 2017
https://scholar.google.com/scholar?q=Mixed+Precision+Training
3. Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference — Benoit Jacob, Skirmantas Kligys, Bo Chen, Menglong Zhu, Matthew Tang, Andrew Howard, and others, 2018
https://scholar.google.com/scholar?q=Quantization+and+Training+of+Neural+Networks+for+Efficient+Integer-Arithmetic-Only+Inference
4. FP8 Formats for Deep Learning — Paulius Micikevicius, Dusan Stosic, Neil Burgess, Marius Cornea, Pradeep Dubey, and others, 2022
https://scholar.google.com/scholar?q=FP8+Formats+for+Deep+Learning
5. AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration — Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Wei-Ming Chen, Wei-Chen Wang, Guangxuan Xiao, Chuang Gan, Song Han, and others, 2023
https://scholar.google.com/scholar?q=AWQ%3A+Activation-aware+Weight+Quantization+for+LLM+Compression+and+Acceleration
6. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness — Tri Dao et al., 2022
https://scholar.google.com/scholar?q=FlashAttention%3A+Fast+and+Memory-Efficient+Exact+Attention+with+IO-Awareness
7. A Full-Stack Search Technique for Domain Optimized Deep Learning Accelerators — Dan Zhang et al., 2022
https://scholar.google.com/scholar?q=A+Full-Stack+Search+Technique+for+Domain+Optimized+Deep+Learning+Accelerators
8. A Compute-in-Memory Chip Based on Resistive Random-Access Memory — Weier Wan et al., 2022
https://scholar.google.com/scholar?q=A+Compute-in-Memory+Chip+Based+on+Resistive+Random-Access+Memory
9. Intelligence per Watt: Measuring Intelligence Efficiency of Local AI — Jon Saad-Falcon et al., 2025
https://scholar.google.com/scholar?q=Intelligence+per+Watt%3A+Measuring+Intelligence+Efficiency+of+Local+AI
10. PinDrop: Breaking the Silence on SDCs in a Large-Scale Fleet — Peter W. Deutsch et al., 2026
https://scholar.google.com/scholar?q=PinDrop%3A+Breaking+the+Silence+on+SDCs+in+a+Large-Scale+Fleet
11. LMCache: An Efficient KV Cache Layer for Enterprise-Scale LLM Inference — Yuhan Liu et al., 2025
https://scholar.google.com/scholar?q=LMCache%3A+An+Efficient+KV+Cache+Layer+for+Enterprise-Scale+LLM+Inference
12. TraCT: Disaggregated LLM Serving with CXL Shared Memory KV Cache at Rack-Scale — Dongha Yoon et al., 2025
https://scholar.google.com/scholar?q=TraCT%3A+Disaggregated+LLM+Serving+with+CXL+Shared+Memory+KV+Cache+at+Rack-Scale
13. Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads — Cunchen Hu et al., 2024
https://scholar.google.com/scholar?q=Inference+without+Interference%3A+Disaggregate+LLM+Inference+for+Mixed+Downstream+Workloads
14. CXL-SpecKV: A Disaggregated FPGA Speculative KV-Cache for Datacenter LLM Serving — Dong Liu and Yanxuan Yu, 2025
https://scholar.google.com/scholar?q=CXL-SpecKV%3A+A+Disaggregated+FPGA+Speculative+KV-Cache+for+Datacenter+LLM+Serving
15. KVLink: Accelerating Large Language Models via Efficient KV Cache Reuse — Jingbo Yang et al., 2025
https://scholar.google.com/scholar?q=KVLink%3A+Accelerating+Large+Language+Models+via+Efficient+KV+Cache+Reuse
16. Selective KV-Cache Sharing to Mitigate Timing Side-Channels in LLM Inference — Kexin Chu et al., 2025
https://scholar.google.com/scholar?q=Selective+KV-Cache+Sharing+to+Mitigate+Timing+Side-Channels+in+LLM+Inference
17. Test-Time Training on Nearest Neighbors for Large Language Models — Moritz Hardt and Yu Sun, 2023
https://scholar.google.com/scholar?q=Test-Time+Training+on+Nearest+Neighbors+for+Large+Language+Models
18. Test-Time Learning for Large Language Models — Jinwu Hu et al., 2025
https://scholar.google.com/scholar?q=Test-Time+Learning+for+Large+Language+Models
19. AI Post Transformers: FlatAttention for Tile-Based Accelerator Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-flatattention-for-tile-based-accelerator-56e6ca.mp3
20. AI Post Transformers: LPU Chip for Low-Latency LLM Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-20-lpu-chip-for-low-latency-llm-inference-be13c3.mp3
21. AI Post Transformers: Mamba-3 for Efficient Sequence Modeling — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-16-mamba-3-for-efficient-sequence-modeling-97a22a.mp3
22. AI Post Transformers: Mistral 7B: Superior Performance in a Smaller Package — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/mistral-7b-superior-performance-in-a-smaller-package/
23. AI Post Transformers: PALOMA: Benchmarking Language Model Fit Across Domains — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-23-paloma-benchmarking-language-model-fit-a-360060.mp3
24. AI Post Transformers: Automating CNN Mapping on Embedded FPGAs — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-06-automating-cnn-mapping-on-embedded-fpgas-4c1dc3.mp3
25. AI Post Transformers: FPGA Neural Network Accelerators for Space — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-26-fpga-neural-network-accelerators-for-spa-3087ae.mp3
26. AI Post Transformers: ASI-Evolve for Data, Architectures, and RL — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-05-asi-evolve-for-data-architectures-and-rl-197b2b.mp3

This episode explores how post-training quantization can convert already-trained models into 8-bit floating point formats for cheaper inference, and why FP8 may outperform the older INT8 approach on modern transformers, LLMs, and diffusion models. It explains the tradeoff between exponent range and mantissa precision across FP8 formats such as E4M3, E5M2, and E3M4, with particular attention to how FP8 handles activation outliers and dynamic range more gracefully than fixed-scale INT8. The discussion centers on a hardware-aware deployment recipe, including which operators can stay quantized, where higher-precision accumulation still matters, and how BatchNorm recalibration helps low-precision inference match full-precision behavior. Listeners get a concrete result: across 75 architectures and more than 200 task cases, the paper reports 92.64% workload coverage for FP8 versus 65.87% for INT8, with E4M3 looking strongest for NLP while E3M4 is slightly better for some vision workloads.

Sources:
1. Efficient Post-training Quantization with FP8 Formats — Haihao Shen, Naveen Mellempudi, Xin He, Qun Gao, Chang Wang, Mengni Wang, 2023
http://arxiv.org/abs/2309.14592
2. Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference — Benoit Jacob, Skirmantas Kligys, Bo Chen, Menglong Zhu, et al., 2018
https://scholar.google.com/scholar?q=Quantization+and+Training+of+Neural+Networks+for+Efficient+Integer-Arithmetic-Only+Inference
3. A White Paper on Neural Network Quantization — Markus Nagel, Marios Fournarakis, Rana Ali Amjad, Yelysei Bondarenko, Mart van Baalen, Tijmen Blankevoort, 2021
https://scholar.google.com/scholar?q=A+White+Paper+on+Neural+Network+Quantization
4. LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale — Tim Dettmers, Mike Lewis, Younes Belkada, Luke Zettlemoyer, 2022
https://scholar.google.com/scholar?q=LLM.int8%28%29%3A+8-bit+Matrix+Multiplication+for+Transformers+at+Scale
5. SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models — Guangxuan Xiao, Ji Lin, Mickael Seznec, Hao Wu, Julien Demouth, Song Han, 2022
https://scholar.google.com/scholar?q=SmoothQuant%3A+Accurate+and+Efficient+Post-Training+Quantization+for+Large+Language+Models
6. 8-bit Numerical Formats for Deep Neural Networks — Badreddine Noune, Philip Jones, Daniel Justus, Dominic Masters, Carlo Luschi, 2022
https://scholar.google.com/scholar?q=8-bit+Numerical+Formats+for+Deep+Neural+Networks
7. FP8 Formats for Deep Learning — Paulius Micikevicius, Dusan Stosic, Neil Burgess, Marius Cornea, Pradeep Dubey, et al., 2022
https://scholar.google.com/scholar?q=FP8+Formats+for+Deep+Learning
8. FP8 Quantization: The Power of the Exponent — Andrey Kuzmin, Mart van Baalen, Yuwei Ren, Markus Nagel, Jorn Peters, Tijmen Blankevoort, 2022
https://scholar.google.com/scholar?q=FP8+Quantization%3A+The+Power+of+the+Exponent
9. Efficient Post-training Quantization with FP8 Formats — Haihao Shen, Naveen Mellempudi, Xin He, Qun Gao, Chang Wang, Mengni Wang, 2024
https://scholar.google.com/scholar?q=Efficient+Post-training+Quantization+with+FP8+Formats
10. Hybrid 8-bit Floating Point (HFP8) Training and Inference for Deep Neural Networks — Xiaohu Sun et al., 2019
https://scholar.google.com/scholar?q=Hybrid+8-bit+Floating+Point+%28HFP8%29+Training+and+Inference+for+Deep+Neural+Networks
11. Outlier Suppression: Pushing the Limit of Low-bit Transformer Language Models — Xiuying Wei et al., 2022
https://scholar.google.com/scholar?q=Outlier+Suppression%3A+Pushing+the+Limit+of+Low-bit+Transformer+Language+Models
12. Quantizable transformers: Removing outliers by helping attention heads do nothing — Bondarenko et al. (approx.), 2023?
https://scholar.google.com/scholar?q=Quantizable+transformers%3A+Removing+outliers+by+helping+attention+heads+do+nothing
13. Understanding and minimising outlier features in transformer training — author list not verified from provided snippet, 2024?
https://scholar.google.com/scholar?q=Understanding+and+minimising+outlier+features+in+transformer+training
14. QuanTool: A Benchmarking Framework for Evaluating Post-Training Quantization with Best Practices for Transformer Models — author list not verified from provided snippet, 2024?
https://scholar.google.com/scholar?q=QuanTool%3A+A+Benchmarking+Framework+for+Evaluating+Post-Training+Quantization+with+Best+Practices+for+Transformer+Models
15. GO-ViT: Fully Quantizing Vision Transformers by Grouping Outlier Channels — author list not verified from provided snippet, 2024?
https://scholar.google.com/scholar?q=GO-ViT%3A+Fully+Quantizing+Vision+Transformers+by+Grouping+Outlier+Channels
16. Understanding int4 quantization for language models: latency speedup, composability, and failure cases — author list not verified from provided snippet, 2024?
https://scholar.google.com/scholar?q=Understanding+int4+quantization+for+language+models%3A+latency+speedup%2C+composability%2C+and+failure+cases
17. Kvquant: Towards 10 million context length llm inference with kv cache quantization — author list not verified from provided snippet, 2024?
https://scholar.google.com/scholar?q=Kvquant%3A+Towards+10+million+context+length+llm+inference+with+kv+cache+quantization
18. AI Post Transformers: MIOpen and AMD's Open Deep Learning Primitives — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-23-miopen-and-amds-open-deep-learning-primi-052f82.mp3
19. AI Post Transformers: Deep Kernel Fusion for Transformer Decoding — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-deep-kernel-fusion-for-transformer-decod-b1a703.mp3
20. AI Post Transformers: Mooncake for KV Cache-Centric LLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-05-mooncake-for-kv-cache-centric-llm-servin-1086d0.mp3
Interactive Visualization: Efficient Post-Training Quantization with FP8

This episode explores PALOMA, a NeurIPS 2024 benchmark designed to measure how well language models fit many different language distributions instead of relying on a single average perplexity score. It explains why one global loss number can hide important weaknesses across domains such as specific subreddits, scientific writing, or programming languages, and highlights PALOMA’s fine-grained setup across 546 English and code domains from 16 sources. The discussion places PALOMA in context with earlier language-model evaluation traditions, scaling-law work, and broader benchmark efforts like HELM, while arguing that evaluation design determines what claims researchers can actually make. Listeners would find it interesting for its clear case that better measurement, data curation, and decontamination can reveal model behavior that broad headline metrics often miss.

Sources:
1. Paloma: A Benchmark for Evaluating Language Model Fit — Ian Magnusson, Akshita Bhagia, Valentin Hofmann, Luca Soldaini, Ananya Harsh Jha, Oyvind Tafjord, Dustin Schwenk, Evan Pete Walsh, Yanai Elazar, Kyle Lo, Dirk Groeneveld, Iz Beltagy, Hannaneh Hajishirzi, Noah A. Smith, Kyle Richardson, Jesse Dodge, 2023
http://arxiv.org/abs/2312.10523
2. One Billion Word Benchmark for Measuring Progress in Statistical Language Modeling — Ciprian Chelba, Tomas Mikolov, Mike Schuster, Qi Ge, Thorsten Brants, et al., 2013
https://scholar.google.com/scholar?q=One+Billion+Word+Benchmark+for+Measuring+Progress+in+Statistical+Language+Modeling
3. Scaling Laws for Neural Language Models — Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Dario Amodei, et al., 2020
https://scholar.google.com/scholar?q=Scaling+Laws+for+Neural+Language+Models
4. Training Compute-Optimal Large Language Models — Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Jack W. Rae, Oriol Vinyals, Laurent Sifre, et al., 2022
https://scholar.google.com/scholar?q=Training+Compute-Optimal+Large+Language+Models
5. Paloma: A Benchmark for Evaluating Language Model Fit — Ian Magnusson, Akshita Bhagia, Valentin Hofmann, Luca Soldaini, Kyle Richardson, Jesse Dodge, et al., 2024
https://scholar.google.com/scholar?q=Paloma%3A+A+Benchmark+for+Evaluating+Language+Model+Fit
6. M2D2: A Massively Multi-domain Language Modeling Dataset — Machel Reid, Victor Zhong, Suchin Gururangan, Luke Zettlemoyer, 2022
https://scholar.google.com/scholar?q=M2D2%3A+A+Massively+Multi-domain+Language+Modeling+Dataset
7. Holistic Evaluation of Language Models — Percy Liang et al., 2022
https://scholar.google.com/scholar?q=Holistic+Evaluation+of+Language+Models
8. Same Pre-training Loss, Better Downstream: Implicit Bias Matters for Language Models — Hong Liu, Sang Michael Xie, Zhiyuan Li, Tengyu Ma, 2022
https://scholar.google.com/scholar?q=Same+Pre-training+Loss%2C+Better+Downstream%3A+Implicit+Bias+Matters+for+Language+Models
9. Language Model Evaluation Beyond Perplexity — Clara Meister, Ryan Cotterell, 2021
https://scholar.google.com/scholar?q=Language+Model+Evaluation+Beyond+Perplexity
10. Unsupervised Domain Clusters in Pretrained Language Models — Roee Aharoni, Yoav Goldberg, 2020
https://scholar.google.com/scholar?q=Unsupervised+Domain+Clusters+in+Pretrained+Language+Models
11. DataComp-LM: In search of the next generation of training sets for language models — Jeffrey Li et al., 2024
https://scholar.google.com/scholar?q=DataComp-LM%3A+In+search+of+the+next+generation+of+training+sets+for+language+models
12. Rethinking Perplexity: Revealing the Impact of Input Length on Perplexity Evaluation in LLMs — Letian Cheng et al., 2026
https://scholar.google.com/scholar?q=Rethinking+Perplexity%3A+Revealing+the+Impact+of+Input+Length+on+Perplexity+Evaluation+in+LLMs
13. Rethinking Benchmark and Contamination for Language Models with Rephrased Samples — Shuo Yang et al., 2023
https://scholar.google.com/scholar?q=Rethinking+Benchmark+and+Contamination+for+Language+Models+with+Rephrased+Samples
14. PaCoST: Paired Confidence Significance Testing for Benchmark Contamination Detection in Large Language Models — Huixuan Zhang et al., 2024
https://scholar.google.com/scholar?q=PaCoST%3A+Paired+Confidence+Significance+Testing+for+Benchmark+Contamination+Detection+in+Large+Language+Models
15. RegMix: Data Mixture as Regression for Language Model Pre-training — Qian Liu et al., 2024
https://scholar.google.com/scholar?q=RegMix%3A+Data+Mixture+as+Regression+for+Language+Model+Pre-training
16. Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance — Jiasheng Ye et al., 2024
https://scholar.google.com/scholar?q=Data+Mixing+Laws%3A+Optimizing+Data+Mixtures+by+Predicting+Language+Modeling+Performance
17. The Fools are Certain; the Wise are Doubtful: Exploring LLM Confidence in Code Completion — Zoe Kotti et al., 2025
https://scholar.google.com/scholar?q=The+Fools+are+Certain%3B+the+Wise+are+Doubtful%3A+Exploring+LLM+Confidence+in+Code+Completion
18. AI Post Transformers: Model-Aware Tokenizer Transfer for Multilingual LLMs — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-16-model-aware-tokenizer-transfer-for-multi-90666c.mp3
19. AI Post Transformers: LLM Benchmark Robustness to Linguistic Variation — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/llm-benchmark-robustness-to-linguistic-variation/

This episode explores AMD’s open-source MIOpen library and why deep learning primitives such as convolution, pooling, normalization, and activations are the layer where model performance meets GPU hardware reality. It explains how CNN throughput depends on low-level execution choices, comparing approaches such as im2col-plus-GEMM and Winograd convolution, and shows why libraries like MIOpen need solver-based algorithm selection and auto-tuning to match different workload shapes, precisions, and GPUs. The discussion also covers mixed-precision support, especially bfloat16, along with kernel fusion and composable kernels as ways to reduce memory traffic and launch overhead while keeping vendor-library speed. Listeners would find it interesting because it turns “invisible infrastructure” into a concrete systems story about how open-source GPU software can shape real model training and inference performance.

Sources:
1. MIOpen: An Open Source Library For Deep Learning Primitives — Jehandad Khan, Paul Fultz, Artem Tamazov, Daniel Lowell, Chao Liu, Michael Melesse, Murali Nandhimandalam, Kamil Nasyrov, Ilya Perminov, Tejash Shah, Vasilii Filippov, Jing Zhang, Jing Zhou, Bragadeesh Natarajan, Mayank Daga, 2019
http://arxiv.org/abs/1910.00078
2. cuDNN: Efficient Primitives for Deep Learning — Sharan Chetlur, Cliff Woolley, Philippe Vandermersch, Jonathan Cohen, John Tran, Bryan Catanzaro, Evan Shelhamer, 2014
https://scholar.google.com/scholar?q=cuDNN%3A+Efficient+Primitives+for+Deep+Learning
3. MIOpen: An Open Source Library For Deep Learning Primitives — Jehandad Khan, Paul Fultz, Artem Tamazov, Daniel Lowell, Chao Liu, Michael Melesse, Murali Nandhimandalam, Kamil Nasyrov, Ilya Perminov, Tejash Shah, Vasilii Filippov, Jing Zhang, Jing Zhou, Bragadeesh Natarajan, Mayank Daga, 2019
https://scholar.google.com/scholar?q=MIOpen%3A+An+Open+Source+Library+For+Deep+Learning+Primitives
4. TVM: An Automated End-to-End Optimizing Compiler for Deep Learning — Tianqi Chen, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Yan, Meghan Cowan, Haichen Shen, Leyuan Wang, Yuwei Hu, Luis Ceze, Carlos Guestrin, Arvind Krishnamurthy, 2018
https://scholar.google.com/scholar?q=TVM%3A+An+Automated+End-to-End+Optimizing+Compiler+for+Deep+Learning
5. Ansor: Generating High-Performance Tensor Programs for Deep Learning — Lianmin Zheng, Chengfan Jia, Minmin Sun, Zhao Wu, Cody Hao Yu, Ameer Haj-Ali, Yida Wang, Jun Yang, Danyang Zhuo, Koushik Sen, Joseph E. Gonzalez, Ion Stoica, 2020
https://scholar.google.com/scholar?q=Ansor%3A+Generating+High-Performance+Tensor+Programs+for+Deep+Learning
6. Automatically tuned linear algebra software — R. Clint Whaley and Jack J. Dongarra, 1998
https://scholar.google.com/scholar?q=Automatically+tuned+linear+algebra+software
7. Fast Algorithms for Convolutional Neural Networks — Andrew Lavin and Scott Gray, 2015
https://scholar.google.com/scholar?q=Fast+Algorithms+for+Convolutional+Neural+Networks
8. TVM: end-to-end optimization stack for deep learning — Tianqi Chen et al., 2018
https://scholar.google.com/scholar?q=TVM%3A+end-to-end+optimization+stack+for+deep+learning
9. Tensor Comprehensions: Framework-Agnostic High-Performance Machine Learning Abstractions — Nicolas Vasilache et al., 2018
https://scholar.google.com/scholar?q=Tensor+Comprehensions%3A+Framework-Agnostic+High-Performance+Machine+Learning+Abstractions
10. oneDNN Graph Compiler: A Hybrid Approach for High-Performance Deep Learning Compilation — approx. Intel oneDNN graph compiler authors, 2023
https://scholar.google.com/scholar?q=oneDNN+Graph+Compiler%3A+A+Hybrid+Approach+for+High-Performance+Deep+Learning+Compilation
11. SparseTIR: Composable Abstractions for Sparse Compilation in Deep Learning — approx. TVM / SparseTIR authors, 2023
https://scholar.google.com/scholar?q=SparseTIR%3A+Composable+Abstractions+for+Sparse+Compilation+in+Deep+Learning
12. Autotuning Convolutions Is Easier Than You Think — approx. tensor-compiler autotuning authors, 2023
https://scholar.google.com/scholar?q=Autotuning+Convolutions+Is+Easier+Than+You+Think
13. Haotuner: A Hardware Adaptive Operator Auto-Tuner for Dynamic Shape Tensor Compilers — approx. Haotuner authors, 2024
https://scholar.google.com/scholar?q=Haotuner%3A+A+Hardware+Adaptive+Operator+Auto-Tuner+for+Dynamic+Shape+Tensor+Compilers
14. The Case for Training Large Models in Low Precision — approx. low-precision training authors, 2024
https://scholar.google.com/scholar?q=The+Case+for+Training+Large+Models+in+Low+Precision
15. AI Post Transformers: Deep Kernel Fusion for Transformer Decoding — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-deep-kernel-fusion-for-transformer-decod-b1a703.mp3
16. AI Post Transformers: FlashFuser and Hopper-Era FFN Kernel Fusion — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-flashfuser-and-hopper-era-ffn-kernel-fus-e1fce9.mp3
17. AI Post Transformers: Automating DNN Compilation for FPGA Accelerators — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-06-automating-dnn-compilation-for-fpga-acce-6ef9bf.mp3
18. AI Post Transformers: Caffeine: A Unified FPGA for CNNs — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-06-caffeine-a-unified-fpga-for-cnns-e8acbe.mp3

This episode explores why open relational foundation models struggle on real downstream database tasks, using OpenRFM as a case study in relational in-context learning across multi-table data such as healthcare, fraud, and recommendation systems. It explains how the RT backbone builds breadth-first relational contexts, why that setup often reduces to a kernel-regression-like similarity lookup, and how limited labeled evidence in those walks creates a label-scarcity bottleneck. The discussion highlights the paper’s main argument that both architecture and pretraining prior matter: adding a TabICL head gives the query direct access to a full batch of support examples, while better synthetic and real-data pretraining pushes the model from shallow similarity matching toward actual relational feature learning. Listeners would find it interesting because the episode goes beyond benchmark gains to unpack a concrete failure mode, then shows how OpenRFM turns that diagnosis into a reported roughly 30% average improvement over the RT baseline.

Sources:
1. OpenRFM: Dissecting Relational In-Context Learning — Zhikai Chen, Junyu Yin, Jialiang Gu, Siheng Xiong, Xiaoze Liu, Ruowang Zhang, Keren Zhou, Kai Guo, 2026
http://arxiv.org/abs/2606.04320
2. Neural Tangent Kernel: Convergence and Generalization in Neural Networks — Arthur Jacot, Franck Gabriel, Clement Hongler, 2018
https://arxiv.org/abs/1806.07572
3. On Lazy Training in Differentiable Programming — Lenaic Chizat, Edouard Oyallon, Francis Bach, 2019
https://arxiv.org/abs/1812.07956
4. What learning algorithm is in-context learning? Investigations with linear models — Ekin Akyurek, Dale Schuurmans, Jacob Andreas, Tengyu Ma, Denny Zhou, 2022
https://arxiv.org/abs/2211.15661
5. TabPFN: A Transformer That Solves Small Tabular Classification Problems in a Second — Noah Hollmann, Samuel Muller, Katharina Eggensperger, Frank Hutter, 2022
https://arxiv.org/abs/2207.01848
6. Assortative mixing in networks — M. E. J. Newman, 2002
https://arxiv.org/abs/cond-mat/0205405
7. Beyond Homophily in Graph Neural Networks: Current Limitations and Effective Designs — Jiong Zhu, Yujun Yan, Lingxiao Zhao, Mark Heimann, Leman Akoglu, Danai Koutra, 2020
https://arxiv.org/abs/2006.11468
8. Graph Neural Networks with Heterophily — Jiong Zhu, Ryan A. Rossi, Anup Rao, Tung Mai, Nedim Lipka, Nesreen K. Ahmed, Danai Koutra, 2020
https://arxiv.org/abs/2009.13566
9. PluRel: Synthetic Data unlocks Scaling Laws for Relational Foundation Models — Vignesh Kothapalli, Rishabh Ranjan, Valter Hudovernik, Vijay Prakash Dwivedi, Johannes Hoffart, Carlos Guestrin, Jure Leskovec, 2026
https://arxiv.org/abs/2602.04029
10. Relational Transformer: Toward Zero-Shot Foundation Models for Relational Data — Rishabh Ranjan, Valter Hudovernik, Mark Znidar, Charilaos I. Kanatsoulis, Roshan Reddy Upendra, Mahmoud Mohammadi, Joe Meyer, Tom Palczewski, Carlos Guestrin, and Jure Leskovec, 2026
https://scholar.google.com/scholar?q=Relational+Transformer%3A+Toward+Zero-Shot+Foundation+Models+for+Relational+Data
11. TabICL: A Tabular Foundation Model for In-Context Learning on Large Data — Jingang Qu, David Holzmuller, Gael Varoquaux, and Marine Le Morvan, 2025
https://scholar.google.com/scholar?q=TabICL%3A+A+Tabular+Foundation+Model+for+In-Context+Learning+on+Large+Data
12. KumoRFM: A Foundation Model for In-Context Learning on Relational Data — Matthias Fey, Vid Kocijan, Federico Lopez, Jan Eric Lenssen, and Jure Leskovec, 2025
https://scholar.google.com/scholar?q=KumoRFM%3A+A+Foundation+Model+for+In-Context+Learning+on+Relational+Data
13. RelBench: A Benchmark for Deep Learning on Relational Databases — Joshua Robinson, Rishabh Ranjan, Weihua Hu, Kexin Huang, Jiaqi Han, Alejandro Dobles, Matthias Fey, Jan E. Lenssen, Yiwen Yuan, Zecheng Zhang, Xinwei He, and Jure Leskovec, 2024
https://scholar.google.com/scholar?q=RelBench%3A+A+Benchmark+for+Deep+Learning+on+Relational+Databases
14. Understanding Emergent In-Context Learning from a Kernel Regression Perspective — Chi Han, Ziqi Wang, Han Zhao, and Heng Ji, 2025
https://scholar.google.com/scholar?q=Understanding+Emergent+In-Context+Learning+from+a+Kernel+Regression+Perspective
15. Retrieval & Fine-Tuning for In-Context Tabular Models — Valentin Thomas et al., 2024
https://scholar.google.com/scholar?q=Retrieval+%26+Fine-Tuning+for+In-Context+Tabular+Models
16. On Finetuning Tabular Foundation Models — Ivan Rubachev et al., 2025
https://scholar.google.com/scholar?q=On+Finetuning+Tabular+Foundation+Models
17. Turning Tabular Foundation Models into Graph Foundation Models — Dmitry Eremeev et al., 2025
https://scholar.google.com/scholar?q=Turning+Tabular+Foundation+Models+into+Graph+Foundation+Models
18. Of Graphs and Tables: Zero-Shot Node Classification with Tabular Foundation Models — Adrian Hayler et al., 2025
https://scholar.google.com/scholar?q=Of+Graphs+and+Tables%3A+Zero-Shot+Node+Classification+with+Tabular+Foundation+Models
19. A Pre-training Framework for Relational Data with Information-theoretic Principles — Quang Truong et al., 2025
https://scholar.google.com/scholar?q=A+Pre-training+Framework+for+Relational+Data+with+Information-theoretic+Principles
20. When Heterophily Meets Heterogeneity: New Graph Benchmarks and Effective Methods — Junhong Lin et al., 2024
https://scholar.google.com/scholar?q=When+Heterophily+Meets+Heterogeneity%3A+New+Graph+Benchmarks+and+Effective+Methods
21. Aligning the Spectrum: Hybrid Graph Pre-training and Prompt Tuning across Homophily and Heterophily — Haitong Luo et al., 2025
https://scholar.google.com/scholar?q=Aligning+the+Spectrum%3A+Hybrid+Graph+Pre-training+and+Prompt+Tuning+across+Homophily+and+Heterophily
22. AI Post Transformers: Can Models Learn from Long Context? — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-can-models-learn-from-long-context-77533e.mp3
23. AI Post Transformers: When Many-Shot CoT Becomes Test-Time Learning — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-when-many-shot-cot-becomes-test-time-lea-c25bfe.mp3
24. AI Post Transformers: TransactionGPT as a Payments Foundation Model — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-17-transactiongpt-as-a-payments-foundation-d53687.mp3

This episode explores X-LLM, a 2023 system that treats images, video, and speech as foreign languages a frozen ChatGLM can learn to read through learned modality-to-language bridges. It breaks down the paper’s architecture, including Q-Former-based visual adapters and a separate speech pipeline with continuous integrate-and-fire modules, to show how three sensory routes feed a single dialogue model instead of one end-to-end multimodal transformer. The discussion argues that X-LLM mattered less as proof of a universal multimodal theory than as a practical open-model recipe shaped by 2023 compute limits, with its Chinese-language backbone playing a real methodological role rather than serving as background context. Listeners get a sharp comparison between this bridge-based approach and later end-to-end systems such as GPT-4o and Gemini 1.5, making the episode useful for understanding how modern multimodal assistants actually evolved.

Sources:
1. X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages — Feilong Chen, Minglun Han, Haozhi Zhao, Qingyang Zhang, Jing Shi, Shuang Xu, Bo Xu, 2023
http://arxiv.org/abs/2305.04160
2. Multimodal Few-Shot Learning with Frozen Language Models — Maria Tsimpoukelli, Jacob Menick, Oriol Vinyals, Felix Hill, 2021
https://scholar.google.com/scholar?q=Multimodal+Few-Shot+Learning+with+Frozen+Language+Models
3. Flamingo: a Visual Language Model for Few-Shot Learning — Jean-Baptiste Alayrac, Jeff Donahue, Karen Simonyan, Oriol Vinyals, Andrew Zisserman, 2022
https://scholar.google.com/scholar?q=Flamingo%3A+a+Visual+Language+Model+for+Few-Shot+Learning
4. BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models — Junnan Li, Dongxu Li, Silvio Savarese, Steven Hoi, 2023
https://scholar.google.com/scholar?q=BLIP-2%3A+Bootstrapping+Language-Image+Pre-training+with+Frozen+Image+Encoders+and+Large+Language+Models
5. SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities — Dong Zhang, Shimin Li, Xin Zhang, Xipeng Qiu, 2023
https://scholar.google.com/scholar?q=SpeechGPT%3A+Empowering+Large+Language+Models+with+Intrinsic+Cross-Modal+Conversational+Abilities
6. Visual Instruction Tuning — Haotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae Lee, 2023
https://scholar.google.com/scholar?q=Visual+Instruction+Tuning
7. MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models — Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, Mohamed Elhoseiny, 2023
https://scholar.google.com/scholar?q=MiniGPT-4%3A+Enhancing+Vision-Language+Understanding+with+Advanced+Large+Language+Models
8. PaLM-E: An Embodied Multimodal Language Model — Danny Driess et al., 2023
https://scholar.google.com/scholar?q=PaLM-E%3A+An+Embodied+Multimodal+Language+Model
9. CIF: Continuous Integrate-and-Fire for End-to-End Speech Recognition — Linhao Dong, Bo Xu, 2019
https://scholar.google.com/scholar?q=CIF%3A+Continuous+Integrate-and-Fire+for+End-to-End+Speech+Recognition
10. VL-JEPA: Joint Embedding Predictive Architecture for Vision-language — Delong Chen et al., 2025
https://scholar.google.com/scholar?q=VL-JEPA%3A+Joint+Embedding+Predictive+Architecture+for+Vision-language
11. TokenPacker: Efficient Visual Projector for Multimodal LLM — Wentong Li et al., 2024
https://scholar.google.com/scholar?q=TokenPacker%3A+Efficient+Visual+Projector+for+Multimodal+LLM
12. Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM — Donghwan Chi et al., 2025
https://scholar.google.com/scholar?q=Slot-MLLM%3A+Object-Centric+Visual+Tokenization+for+Multimodal+LLM
13. Auto-Encoding Morph-Tokens for Multimodal LLM — Kaihang Pan et al., 2024
https://scholar.google.com/scholar?q=Auto-Encoding+Morph-Tokens+for+Multimodal+LLM
14. ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention — Wenjie Liu et al., 2026
https://scholar.google.com/scholar?q=ViCA%3A+Efficient+Multimodal+LLMs+with+Vision-Only+Cross-Attention
15. F-LMM: Grounding Frozen Large Multimodal Models — Size Wu et al., 2024
https://scholar.google.com/scholar?q=F-LMM%3A+Grounding+Frozen+Large+Multimodal+Models
16. MultiModal-GPT: A Vision and Language Model for Dialogue with Humans — Tao Gong et al., 2023
https://scholar.google.com/scholar?q=MultiModal-GPT%3A+A+Vision+and+Language+Model+for+Dialogue+with+Humans
17. AI Post Transformers: UniVideo: Unified Video Understanding, Generation, and Editing — Hal Turing & Dr. Ada Shannon, Sat,
https://podcast.do-not-panic.com/episodes/univideo-unified-video-understanding-generation-and-editing/
18. AI Post Transformers: DeepSeek-OCR: Contexts Optical Compression — Hal Turing & Dr. Ada Shannon, Sat,
https://podcast.do-not-panic.com/episodes/deepseek-ocr-contexts-optical-compression/
Interactive Visualization: X-LLM: Treating Multimodalities as Foreign Languages

This episode explores a paper on building reusable LLM-based simulations of specific individuals by grounding agents in people’s own interviews, survey responses, or both, rather than relying on thin demographic personas. It explains how the system was tested on 1,052 Americans using holdout evaluations across survey questions, personality traits, behavioral experiments, and randomized intervention outcomes to measure real generalization instead of simple recall. The discussion highlights the main result that self-report-grounded agents performed much better than demographics-only baselines, with combined interview-and-survey agents reaching 86 percent of a person’s own two-week consistency versus 74 percent for demographics alone. It is interesting because it frames these agents as a possible new tool for social science and policy research while also probing hard questions about fairness, stereotype reduction, and whether strong results on language-based self-reports truly amount to deep behavior simulation.

Sources:
1. LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals — Joon Sung Park, Carolyn Q. Zou, Jonne Kamphorst, Niles Egan, Aaron Shaw, Benjamin Mako Hill, Carrie Cai, Meredith Ringel Morris, Percy Liang, Robb Willer, Michael S. Bernstein, 2024
http://arxiv.org/abs/2411.10109
2. Generative Agents: Interactive Simulacra of Human Behavior — Joon Sung Park, Joseph C. O'Brien, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, Michael S. Bernstein, 2023
https://arxiv.org/abs/2304.03442
3. Out of One, Many: Using Language Models to Simulate Human Samples — Lisa P. Argyle, Ethan C. Busby, Nancy Fulda, Joshua Gubler, Christopher Rytting, David Wingate, 2022
https://arxiv.org/abs/2209.06899
4. Generative Agent Simulations of 1,000 People — Joon Sung Park, Carolyn Q. Zou, Aaron Shaw, Benjamin Mako Hill, Carrie Cai, Meredith Ringel Morris, Robb Willer, Percy Liang, Michael S. Bernstein, 2024
https://arxiv.org/abs/2411.10109
5. AI-Augmented Surveys: Leveraging Large Language Models and Surveys for Opinion Prediction — Junsol Kim, Byungkyu Lee, 2023
https://arxiv.org/abs/2305.09620
6. Large Language Models Show Human-like Social Desirability Biases in Survey Responses — Aadesh Salecha, Molly E. Ireland, Shashanka Subrahmanya, Joao Sedoc, Lyle H. Ungar, Johannes C. Eichstaedt, 2024
https://arxiv.org/abs/2405.06058
7. Interview-Informed Generative Agents for Product Discovery: A Validation Study — Zichao Wang, Alexa Siu, 2026
https://arxiv.org/abs/2603.29890
8. Large Language Models as Simulated Economic Agents: What Can We Learn from Homo Silicus? — John J. Horton, Apostolos Filippas, Benjamin S. Manning, 2023
https://arxiv.org/abs/2301.07543
9. From Individual to Society: A Survey on Social Simulation Driven by Large Language Model-based Agents — Xinyi Mou, Xuanwen Ding, Qi He, Liang Wang, Jingcong Liang, Xinnong Zhang, et al., 2024
https://arxiv.org/abs/2412.03563
10. Using Large Language Models to Create AI Personas for Replication, Generalization and Prediction of Media Effects: An Empirical Test of 133 Published Experimental Research Findings — Leo Yeykelis, Kaavya Pichai, James J. Cummings, Byron Reeves, 2024
https://arxiv.org/abs/2408.16073
11. Synthetic Users, Real Differences: an Evaluation Framework for User Simulation in Multi-Turn Conversations — Yu Lu Liu, Hyokun Yun, Tanya Roosta, Ziang Xiao, 2026
https://arxiv.org/abs/2605.02624
12. Latent Human Traits in the Language of Social Media: An Open-Vocabulary Approach — Vivek Kulkarni, Margaret L. Kern, David Stillwell, Michal Kosinski, Sandra Matz, Lyle Ungar, Steven Skiena, H. Andrew Schwartz, 2017
https://arxiv.org/abs/1705.08038
13. Is ChatGPT a Good Personality Recognizer? A Preliminary Study — Yu Ji, Wen Wu, Hong Zheng, Yi Hu, Xi Chen, Liang He, 2023
https://arxiv.org/abs/2307.03952
14. Can LLMs Infer Personality from Real World Conversations? — Jianfeng Zhu, Ruoming Jin, Karin G. Coifman, 2025
https://arxiv.org/abs/2507.14355
15. The Personality Illusion: Revealing Dissociation Between Self-Reports & Behavior in LLMs — Pengrui Han, Rafal Kocielnik, Peiyang Song, Ramit Debnath, Dean Mobbs, Anima Anandkumar, R. Michael Alvarez, 2025
https://arxiv.org/abs/2509.03730
16. Marked Personas: Using Natural Language Prompts to Measure Stereotypes in Language Models — Myra Cheng, Esin Durmus, Dan Jurafsky, 2023
https://arxiv.org/abs/2305.18189
17. On the steerability of large language models toward data-driven personas — Junyi Li, Ninareh Mehrabi, Charith Peris, Palash Goyal, Kai-Wei Chang, Aram Galstyan, Richard Zemel, Rahul Gupta, 2023
https://arxiv.org/abs/2311.04978
18. Reading Between the Prompts: How Stereotypes Shape LLM's Implicit Personalization — Vera Neplenbroek, Arianna Bisazza, Raquel Fernandez, 2025
https://arxiv.org/abs/2505.16467
19. Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject Studies — Gati Aher, Rosa I. Arriaga, Adam Tauman Kalai, 2022
https://scholar.google.com/scholar?q=Using+Large+Language+Models+to+Simulate+Multiple+Humans+and+Replicate+Human+Subject+Studies
20. Can Large Language Models Capture Public Opinion about Global Warming? An Empirical Assessment of Algorithmic Fidelity and Bias — S. Lee, T. Q. Peng, M. H. Goldberg, S. A. Rosenthal, J. E. Kotcher, E. W. Maibach, A. Leiserowitz, 2023
https://scholar.google.com/scholar?q=Can+Large+Language+Models+Capture+Public+Opinion+about+Global+Warming%3F+An+Empirical+Assessment+of+Algorithmic+Fidelity+and+Bias
21. How Far are LLMs from Being Our Digital Twins? A Benchmark for Persona-Based Behavior Chain Simulation — Rui Li, Heming Xia, Xinfeng Yuan, Qingxiu Dong, Lei Sha, Wenjie Li, Zhifang Sui, 2025
https://scholar.google.com/scholar?q=How+Far+are+LLMs+from+Being+Our+Digital+Twins%3F+A+Benchmark+for+Persona-Based+Behavior+Chain+Simulation
22. Know Me, Respond to Me: Benchmarking LLMs for Dynamic User Profiling and Personalized Responses at Scale — Bowen Jiang et al., 2025
https://scholar.google.com/scholar?q=Know+Me%2C+Respond+to+Me%3A+Benchmarking+LLMs+for+Dynamic+User+Profiling+and+Personalized+Responses+at+Scale
23. PersonaX: A Recommendation Agent Oriented User Modeling Framework for Long Behavior Sequence — Yunxiao Shi et al., 2025
https://scholar.google.com/scholar?q=PersonaX%3A+A+Recommendation+Agent+Oriented+User+Modeling+Framework+for+Long+Behavior+Sequence
24. Finetuning LLMs for Human Behavior Prediction in Social Science Experiments — Akaash Kolluri, Shengguang Wu, Joon Sung Park, Michael S. Bernstein, 2025
https://scholar.google.com/scholar?q=Finetuning+LLMs+for+Human+Behavior+Prediction+in+Social+Science+Experiments
25. Tuning Language Models for Robust Prediction of Diverse User Behaviors — Fanjin Meng et al., 2025
https://scholar.google.com/scholar?q=Tuning+Language+Models+for+Robust+Prediction+of+Diverse+User+Behaviors
26. AI Post Transformers: RAGEN-2: Reasoning Collapse in Agentic RL — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-07-ragen-2-reasoning-collapse-in-agentic-rl-3cfa0b.mp3
27. AI Post Transformers: Split Personality Training Reveals Latent Knowledge — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-08-split-personality-training-reveals-laten-c84616.mp3
28. AI Post Transformers: End-to-End Context Compression at Scale — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-10-end-to-end-context-compression-at-scale-278c70.mp3
Interactive Visualization: Simulating Individuals with Self-Reported LLM Agents

This episode explores the UIST 2025 paper "Creating General User Models from Computer Use," which proposes building a persistent user model from raw computer traces such as screenshots, UI text, message context, and app switching. It explains how the system stores confidence-weighted natural-language propositions about a person’s preferences, knowledge, goals, and current situation, aiming to support cross-application assistants that can help proactively rather than waiting for explicit requests. The discussion situates the idea against earlier recommender systems, Bayesian user modeling, and newer LLM memory architectures, arguing that the paper is most interesting as a synthesis of HCI user modeling and retrieval-and-revision style AI memory. Listeners would find it compelling because it gets concrete about both the upside of more context-aware assistants and the hard problems underneath them, including noisy behavioral data, narrow evidence relative to broad claims, and the privacy risks of inferring things users never said aloud.

Sources:
1. Creating General User Models from Computer Use — Omar Shaikh, Shardul Sapkota, Shan Rizvi, Eric Horvitz, Joon Sung Park, Diyi Yang, Michael S. Bernstein, 2025
http://arxiv.org/abs/2505.10831
2. User Modeling via Stereotypes — Elaine Rich, 1979
https://scholar.google.com/scholar?q=User+Modeling+via+Stereotypes
3. The Lumiere Project: Bayesian User Modeling for Inferring the Goals and Needs of Software Users — Eric J. Horvitz, John S. Breese, David Heckerman, David Hovel, Koos Rommelse, 1998
https://scholar.google.com/scholar?q=The+Lumiere+Project%3A+Bayesian+User+Modeling+for+Inferring+the+Goals+and+Needs+of+Software+Users
4. Toward the Next Generation of Recommender Systems: A Survey of the State-of-the-Art and Possible Extensions — Gediminas Adomavicius, Alexander Tuzhilin, 2005
https://scholar.google.com/scholar?q=Toward+the+Next+Generation+of+Recommender+Systems%3A+A+Survey+of+the+State-of-the-Art+and+Possible+Extensions
5. User Modeling and User Profiling: A Comprehensive Survey — Erasmo Purificato, Ludovico Boratto, Ernesto William De Luca, 2024
https://scholar.google.com/scholar?q=User+Modeling+and+User+Profiling%3A+A+Comprehensive+Survey
6. Generative Agents: Interactive Simulacra of Human Behavior — Joon Sung Park, Joseph C. O'Brien, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, Michael S. Bernstein, 2023
https://scholar.google.com/scholar?q=Generative+Agents%3A+Interactive+Simulacra+of+Human+Behavior
7. MemGPT: Towards LLMs as Operating Systems — Charles Packer, Sarah Wooders, Kevin Lin, Vivian Fang, Shishir G. Patil, Ion Stoica, Joseph E. Gonzalez, 2023
https://scholar.google.com/scholar?q=MemGPT%3A+Towards+LLMs+as+Operating+Systems
8. Collaborative Gym: A Framework for Enabling and Evaluating Human-Agent Collaboration — Yijia Shao, Vinay Samuel, Yucheng Jiang, John Yang, Diyi Yang, 2024
https://scholar.google.com/scholar?q=Collaborative+Gym%3A+A+Framework+for+Enabling+and+Evaluating+Human-Agent+Collaboration
9. Supporting Physical Activity Behavior Change with LLM-Based Conversational Agents — Matthew Jörke, Shardul Sapkota, Lyndsea Warkenthien, Niklas Vainio, Paul Schmiedmayer, Emma Brunskill, James Landay, 2024
https://scholar.google.com/scholar?q=Supporting+Physical+Activity+Behavior+Change+with+LLM-Based+Conversational+Agents
10. TaskTracer: a desktop environment to support multi-tasking knowledge workers — Anton N. Dragunov, Thomas G. Dietterich, Kevin Johnsrude, Matthew McLaughlin, Lida Li, Jonathan L. Herlocker, 2005
https://scholar.google.com/scholar?q=TaskTracer%3A+a+desktop+environment+to+support+multi-tasking+knowledge+workers
11. UI-TARS: Pioneering Automated GUI Interaction with Native Agents — Yujia Qin et al., 2025
https://arxiv.org/abs/2501.12326
12. Need Help? Designing Proactive AI Assistants for Programming — Valerie Chen et al., 2024/2025
https://arxiv.org/abs/2410.04596
13. Proactive Agent Research Environment: Simulating Active Users to Evaluate Proactive Assistants — Deepak Nathani et al., 2026
https://arxiv.org/abs/2604.00842
14. ProPerSim: Developing Proactive and Personalized AI Assistants through User-Assistant Simulation — Jiho Kim et al., 2025/2026
https://arxiv.org/abs/2509.21730
15. AI Post Transformers: Learning Latent Action World Models from Video — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-09-learning-latent-action-world-models-from-1570a4.mp3
16. AI Post Transformers: PaperBench: Can AI Replicate AI Research? — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-17-paperbench-can-ai-replicate-ai-research-862944.mp3
17. AI Post Transformers: Reasoning Theater and Unfaithful Chain-of-Thought — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-reasoning-theater-and-unfaithful-chain-o-a4507e.mp3
18. AI Post Transformers: Neural Computers as Learned Latent Runtimes — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-11-neural-computers-as-learned-latent-runti-9fa282.mp3
Interactive Visualization: Building General User Models from Computer Use

This episode explores a 2025 study on fine-tuning large language models to predict how people respond in social science experiments, asking whether trained models can simulate new studies more reliably than prompting alone. It explains how the researchers built SOCSCI210, a dataset of 2.9 million responses from more than 400,000 participants across 210 TESS experiments, and why standardizing those studies into respondent-condition-question-answer records is central to the method. The discussion breaks down the paper’s evaluation criteria, including out-of-distribution generalization, distribution matching via Wasserstein distance, normalized individual accuracy, and treatment-effect recovery, to show the difference between sounding plausible and preserving real experimental patterns. Listeners would find it interesting because it treats LLMs not as chatbots but as possible “wind tunnels” for testing study designs in advance, while also confronting the risk that a convincing simulator could still get causal effects wrong.

Sources:
1. Finetuning LLMs for Human Behavior Prediction in Social Science Experiments — Akaash Kolluri, Shengguang Wu, Joon Sung Park, Michael S. Bernstein, 2025
http://arxiv.org/abs/2509.05830
2. Out of One, Many: Using Language Models to Simulate Human Samples — Lisa P. Argyle, Ethan C. Busby, Nancy Fulda, Joshua Gubler, Christopher Rytting, David Wingate, 2022
https://scholar.google.com/scholar?q=Out+of+One%2C+Many%3A+Using+Language+Models+to+Simulate+Human+Samples
3. LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals — Joon Sung Park, Carolyn Q. Zou, Jonne Kamphorst, Niles Egan, Aaron Shaw, Michael S. Bernstein, et al., 2024
https://scholar.google.com/scholar?q=LLM+Agents+Grounded+in+Self-Reports+Enable+General-Purpose+Simulation+of+Individuals
4. Large Language Models Show Human-like Social Desirability Biases in Survey Responses — Aadesh Salecha, Molly E. Ireland, Shashanka Subrahmanya, Joao Sedoc, Lyle H. Ungar, Johannes C. Eichstaedt, 2024
https://scholar.google.com/scholar?q=Large+Language+Models+Show+Human-like+Social+Desirability+Biases+in+Survey+Responses
5. Finetuning LLMs for Human Behavior Prediction in Social Science Experiments — Akaash Kolluri, Shengguang Wu, Joon Sung Park, Michael S. Bernstein, 2025
https://scholar.google.com/scholar?q=Finetuning+LLMs+for+Human+Behavior+Prediction+in+Social+Science+Experiments
6. Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject Studies — Gati Aher, Rosa I. Arriaga, Adam Tauman Kalai, 2022
https://scholar.google.com/scholar?q=Using+Large+Language+Models+to+Simulate+Multiple+Humans+and+Replicate+Human+Subject+Studies
7. Can Large Language Models Replace Human Subjects? A Large-Scale Replication of Scenario-Based Experiments in Psychology and Management — Ziyan Cui, Ning Li, Huaikang Zhou, 2024
https://scholar.google.com/scholar?q=Can+Large+Language+Models+Replace+Human+Subjects%3F+A+Large-Scale+Replication+of+Scenario-Based+Experiments+in+Psychology+and+Management
8. Using Large Language Models to Create AI Personas for Replication, Generalization and Prediction of Media Effects: An Empirical Test of 133 Published Experimental Research Findings — Leo Yeykelis, Kaavya Pichai, James J. Cummings, Byron Reeves, 2024
https://scholar.google.com/scholar?q=Using+Large+Language+Models+to+Create+AI+Personas+for+Replication%2C+Generalization+and+Prediction+of+Media+Effects%3A+An+Empirical+Test+of+133+Published+Experimental+Research+Findings
9. This human study did not involve human subjects: Validating LLM simulations as behavioral evidence — Jessica Hullman, David Broska, Huaman Sun, Aaron Shaw, 2026
https://scholar.google.com/scholar?q=This+human+study+did+not+involve+human+subjects%3A+Validating+LLM+simulations+as+behavioral+evidence
10. Centaur: a Foundation Model of Human Cognition — Marcel Binz et al., 2024
https://scholar.google.com/scholar?q=Centaur%3A+a+Foundation+Model+of+Human+Cognition
11. Language Model Fine-Tuning on Scaled Survey Data for Predicting Distributions of Public Opinions — Joseph Suh, Erfan Jahanparast, Suhong Moon, Minwoo Kang, Serina Chang, 2025
https://scholar.google.com/scholar?q=Language+Model+Fine-Tuning+on+Scaled+Survey+Data+for+Predicting+Distributions+of+Public+Opinions
12. Beyond Demographics: Fine-tuning Large Language Models to Predict Individuals' Subjective Text Perceptions — Matthias Orlikowski, Jiaxin Pei, Paul Rottger, Philipp Cimiano, David Jurgens, Dirk Hovy, 2025
https://scholar.google.com/scholar?q=Beyond+Demographics%3A+Fine-tuning+Large+Language+Models+to+Predict+Individuals%27+Subjective+Text+Perceptions
13. Large Language Models that Replace Human Participants Can Harmfully Misportray and Flatten Identity Groups — Angelina Wang, Jamie Morgenstern, John P. Dickerson, 2025
https://scholar.google.com/scholar?q=Large+Language+Models+that+Replace+Human+Participants+Can+Harmfully+Misportray+and+Flatten+Identity+Groups
14. Beyond Believability: Accurate Human Behavior Simulation with Fine-Tuned LLMs — Yuxuan Lu et al., 2025
https://scholar.google.com/scholar?q=Beyond+Believability%3A+Accurate+Human+Behavior+Simulation+with+Fine-Tuned+LLMs
15. The Prompt Makes the Person(a): A Systematic Evaluation of Sociodemographic Persona Prompting for Large Language Models — Marlene Lutz et al., 2025
https://arxiv.org/abs/2507.16076
16. Prompt Fairness: Sub-group Disparities in LLMs — Meiyu Zhong, Noel Teku, Ravi Tandon, 2025
https://arxiv.org/abs/2511.19956
17. Can Persona-Prompted LLMs Emulate Subgroup Values? An Empirical Analysis of Generalisability and Fairness in Cultural Alignment — Bryan Chen Zhengyu Tan et al., 2026
https://arxiv.org/abs/2604.12851
18. Expert Personas Improve LLM Alignment but Damage Accuracy: Bootstrapping Intent-Based Persona Routing with PRISM — Zizhao Hu, Mohammad Rostami, Jesse Thomason, 2026
https://arxiv.org/abs/2603.18507
19. Causality for Large Language Models — Anpeng Wu et al., 2024
https://arxiv.org/abs/2410.15319
20. AI Post Transformers: PaperBench: Can AI Replicate AI Research? — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-17-paperbench-can-ai-replicate-ai-research-862944.mp3
21. AI Post Transformers: When Many-Shot CoT Becomes Test-Time Learning — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-when-many-shot-cot-becomes-test-time-lea-c25bfe.mp3
Interactive Visualization: Fine-Tuning LLMs for Human Behavior Prediction

This episode explores Social Simulacra, a method for using large language models to prototype entire online communities before they exist by generating synthetic members, posts, and reply threads from a community goal, rules, and a small set of seed personas. It explains why that matters for social computing: small pilots can miss emergent failures like norm drift, newcomer enculturation problems, trolling, and moderator overload, while a populated simulation can expose those dynamics much earlier. The discussion breaks down how the paper uses prompt chaining to scale a handful of personas into a larger Reddit-like population and then tests interventions such as comment removal, warnings, and rule restatements. It also argues that human-sounding text is a low bar, and that the real challenge is whether these simulations capture believable long-term behavior, incentives, and feedback loops well enough to inform product design.

Sources:
1. Social Simulacra: Creating Populated Prototypes for Social Computing Systems — Joon Sung Park, Lindsay Popowski, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, Michael S. Bernstein, 2022
http://arxiv.org/abs/2208.04024
2. Two Case Studies of Experience Prototyping Machine Learning Systems in the Wild — Qian Yang, 2019
https://scholar.google.com/scholar?q=Two+Case+Studies+of+Experience+Prototyping+Machine+Learning+Systems+in+the+Wild
3. Wizard of Oz Experimentation for Language Technology Applications: Challenges and Tools — Stephan Schlogl, Gavin Doherty, Saturnino Luz, 2014
https://scholar.google.com/scholar?q=Wizard+of+Oz+Experimentation+for+Language+Technology+Applications%3A+Challenges+and+Tools
4. Social Simulacra: Creating Populated Prototypes for Social Computing Systems — Joon Sung Park, Lindsay Popowski, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, Michael S. Bernstein, 2022
https://scholar.google.com/scholar?q=Social+Simulacra%3A+Creating+Populated+Prototypes+for+Social+Computing+Systems
5. Generative Agents: Interactive Simulacra of Human Behavior — Joon Sung Park, Joseph C. O'Brien, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, Michael S. Bernstein, 2023
https://scholar.google.com/scholar?q=Generative+Agents%3A+Interactive+Simulacra+of+Human+Behavior
6. Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject Studies — Gati Aher, Rosa I. Arriaga, Adam Tauman Kalai, 2022
https://scholar.google.com/scholar?q=Using+Large+Language+Models+to+Simulate+Multiple+Humans+and+Replicate+Human+Subject+Studies
7. Out of One, Many: Using Language Models to Simulate Human Samples — Lisa P. Argyle, Ethan C. Busby, Nancy Fulda, Joshua Gubler, Christopher Rytting, David Wingate, 2022
https://scholar.google.com/scholar?q=Out+of+One%2C+Many%3A+Using+Language+Models+to+Simulate+Human+Samples
8. From Individual to Society: A Survey on Social Simulation Driven by Large Language Model-based Agents — Xinyi Mou, Xuanwen Ding, Qi He, Liang Wang, Jingcong Liang, Xinnong Zhang, Libo Sun, Jiayu Lin, Jie Zhou, Xuanjing Huang, Zhongyu Wei, 2024
https://scholar.google.com/scholar?q=From+Individual+to+Society%3A+A+Survey+on+Social+Simulation+Driven+by+Large+Language+Model-based+Agents
9. SoK: Content Moderation in Social Media, from Guidelines to Enforcement, and Research to Practice — Mohit Singhal, Chen Ling, Pujan Paudel, Poojitha Thota, Nihal Kumarswamy, Gianluca Stringhini, Shirin Nilizadeh, 2022
https://scholar.google.com/scholar?q=SoK%3A+Content+Moderation+in+Social+Media%2C+from+Guidelines+to+Enforcement%2C+and+Research+to+Practice
10. ModSandbox: Facilitating Online Community Moderation Through Error Prediction and Improvement of Automated Rules — Jean Y. Song, Sangwook Lee, Jisoo Lee, Mina Kim, Juho Kim, 2022
https://scholar.google.com/scholar?q=ModSandbox%3A+Facilitating+Online+Community+Moderation+Through+Error+Prediction+and+Improvement+of+Automated+Rules
11. Shaping Online Dialogue: Examining How Community Rules Affect Discussion Structures on Reddit — Anna Fang, Wenjie Yang, Haiyi Zhu, 2023
https://scholar.google.com/scholar?q=Shaping+Online+Dialogue%3A+Examining+How+Community+Rules+Affect+Discussion+Structures+on+Reddit
12. Post Guidance for Online Communities — Manoel Horta Ribeiro, Robert West, Ryan Lewis, Sanjay Kairam, 2024
https://scholar.google.com/scholar?q=Post+Guidance+for+Online+Communities
13. Piggyback prototyping: Using existing, large-scale social computing systems to prototype new ones — Catherine Grevet and Eric Gilbert, 2015
https://scholar.google.com/scholar?q=Piggyback+prototyping%3A+Using+existing%2C+large-scale+social+computing+systems+to+prototype+new+ones
14. The Internet's Hidden Rules: An Empirical Study of Reddit Norm Violations at Micro, Meso, and Macro Scales — Eshwar Chandrasekharan, Mattia Samory, Shagun Jhaver, Hunter Charvat, Amy Bruckman, Cliff Lampe, Jacob Eisenstein, and Eric Gilbert, 2018
https://scholar.google.com/scholar?q=The+Internet%27s+Hidden+Rules%3A+An+Empirical+Study+of+Reddit+Norm+Violations+at+Micro%2C+Meso%2C+and+Macro+Scales
15. Surviving an "Eternal September": How an Online Community Managed a Surge of Newcomers — Charles Kiene, Andres Monroy-Hernandez, and Benjamin Mako Hill, 2016
https://scholar.google.com/scholar?q=Surviving+an+%22Eternal+September%22%3A+How+an+Online+Community+Managed+a+Surge+of+Newcomers
16. Building Successful Online Communities: Evidence-Based Social Design — Robert E. Kraut and Paul Resnick, 2012
https://scholar.google.com/scholar?q=Building+Successful+Online+Communities%3A+Evidence-Based+Social+Design
17. Experimental Study of Inequality and Unpredictability in an Artificial Cultural Market — Matthew J. Salganik, Peter Sheridan Dodds, and Duncan J. Watts, 2006
https://scholar.google.com/scholar?q=Experimental+Study+of+Inequality+and+Unpredictability+in+an+Artificial+Cultural+Market
18. Measuring the Prevalence of Anti-Social Behavior in Online Communities — Joon Sung Park, Joseph Seering, and Michael S. Bernstein, 2022
https://scholar.google.com/scholar?q=Measuring+the+Prevalence+of+Anti-Social+Behavior+in+Online+Communities
19. LLM Agents in Interaction: Measuring Personality Consistency and Linguistic Alignment in Interacting Populations of Large Language Models — Ivar Frisch and Mario Giulianelli, 2024
https://scholar.google.com/scholar?q=LLM+Agents+in+Interaction%3A+Measuring+Personality+Consistency+and+Linguistic+Alignment+in+Interacting+Populations+of+Large+Language+Models
20. Persona Alchemy: Designing, Evaluating, and Implementing Psychologically-Grounded LLM Agents for Diverse Stakeholder Representation — Sola Kim, Dongjune Chang, Jieshu Wang, 2025
https://scholar.google.com/scholar?q=Persona+Alchemy%3A+Designing%2C+Evaluating%2C+and+Implementing+Psychologically-Grounded+LLM+Agents+for+Diverse+Stakeholder+Representation
21. Mind the Sim2Real Gap in User Simulation for Agentic Tasks — Xuhui Zhou et al., 2026
https://scholar.google.com/scholar?q=Mind+the+Sim2Real+Gap+in+User+Simulation+for+Agentic+Tasks
22. Computational Turing Test Reveals Systematic Differences Between Human and AI Language — Nicolo Pagan, Petter Tornberg, Christopher A. Bail, Aniko Hannak, Christopher Barrie, 2025
https://scholar.google.com/scholar?q=Computational+Turing+Test+Reveals+Systematic+Differences+Between+Human+and+AI+Language
23. Integrating LLM in Agent-Based Social Simulation: Opportunities and Challenges — Patrick Taillandier et al., 2025
https://scholar.google.com/scholar?q=Integrating+LLM+in+Agent-Based+Social+Simulation%3A+Opportunities+and+Challenges
24. LLM Social Simulations Are a Promising Research Method — Jacy Reese Anthis et al., 2025
https://scholar.google.com/scholar?q=LLM+Social+Simulations+Are+a+Promising+Research+Method
25. AI Post Transformers: Test-time Scaling for Multi-Agent Collaborative Reasoning — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-22-test-time-scaling-for-multi-agent-collab-082570.mp3
Interactive Visualization: Social Simulacra for Prototyping Online Communities

This episode explores a 2026 paper on stabilizing deep reinforcement learning by pushing an agent’s hidden representations toward an isotropic Gaussian shape. It explains how nonstationarity in RL, from shifting data distributions, bootstrapped targets, and primacy bias, can make agents overfit early experience, lose plasticity, and accumulate dormant neurons. The discussion focuses on the paper’s core argument that a round, evenly used feature space makes linear readouts easier to keep tracking as targets drift, reducing collapse and improving adaptation, and it breaks down SIGReg as a lightweight way to enforce that geometry. Listeners would find it interesting because it links an abstract idea from representation geometry to a concrete engineering problem in making RL systems more stable and trainable.

Sources:
1. Stable Deep Reinforcement Learning via Isotropic Gaussian Representations — Ali Saheb Pasand, Johan Obando-Ceron, Aaron Courville, Pouya Bashivan, Pablo Samuel Castro, 2026
http://arxiv.org/abs/2602.19373
2. Whitening for Self-Supervised Representation Learning — Aleksandr Ermolov, Aliaksandr Siarohin, Enver Sangineto, Nicu Sebe, 2020
https://scholar.google.com/scholar?q=Whitening+for+Self-Supervised+Representation+Learning
3. VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised Learning — Adrien Bardes, Jean Ponce, Yann LeCun, 2021
https://scholar.google.com/scholar?q=VICReg%3A+Variance-Invariance-Covariance+Regularization+for+Self-Supervised+Learning
4. The Dormant Neuron Phenomenon in Deep Reinforcement Learning — Ghada Sokar, Rishabh Agarwal, Pablo Samuel Castro, Utku Evci, 2023
https://scholar.google.com/scholar?q=The+Dormant+Neuron+Phenomenon+in+Deep+Reinforcement+Learning
5. LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics — Randall Balestriero, Yann LeCun, 2025
https://scholar.google.com/scholar?q=LeJEPA%3A+Provable+and+Scalable+Self-Supervised+Learning+Without+the+Heuristics
6. The Primacy Bias in Deep Reinforcement Learning — Evgenii Nikishin, Max Schwarzer, Pierluca D'Oro, Pierre-Luc Bacon, Aaron Courville, 2022
https://scholar.google.com/scholar?q=The+Primacy+Bias+in+Deep+Reinforcement+Learning
7. No Representation, No Trust: Connecting Representation, Collapse, and Trust Issues in PPO — Skander Moalla, Andrea Miele, Daniil Pyatko, Razvan Pascanu, Caglar Gulcehre, 2024
https://scholar.google.com/scholar?q=No+Representation%2C+No+Trust%3A+Connecting+Representation%2C+Collapse%2C+and+Trust+Issues+in+PPO
8. Stable Gradients for Stable Learning at Scale in Deep Reinforcement Learning — Roger Creus Castanyer, Johan Obando-Ceron, Lu Li, Pierre-Luc Bacon, Glen Berseth, Aaron Courville, Pablo Samuel Castro, 2025
https://scholar.google.com/scholar?q=Stable+Gradients+for+Stable+Learning+at+Scale+in+Deep+Reinforcement+Learning
9. Proto-Value Networks: Scaling Representation Learning with Auxiliary Tasks — Jesse Farebrother et al., 2023
https://scholar.google.com/scholar?q=Proto-Value+Networks%3A+Scaling+Representation+Learning+with+Auxiliary+Tasks
10. Revisiting Anisotropy in Language Transformers: The Geometry of Learning Dynamics — Raphael Bernas et al., 2026
https://scholar.google.com/scholar?q=Revisiting+Anisotropy+in+Language+Transformers%3A+The+Geometry+of+Learning+Dynamics
11. Emergence of Quantised Representations Isolated to Anisotropic Functions — George Bird, 2025
https://scholar.google.com/scholar?q=Emergence+of+Quantised+Representations+Isolated+to+Anisotropic+Functions
12. AI Post Transformers: When Spectral Gradient Updates Help Deep Learning — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-when-spectral-gradient-updates-help-deep-9c8441.mp3
13. AI Post Transformers: Muon Is Scalable for LLM Training — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-25-muon-is-scalable-for-llm-training-587ed8.mp3

Hal Turing and Dr. Ada Shannon take a deep dive into EMO: Pretraining Mixture of Experts for Emergent Modularity, a May 7, 2026 paper by Ryan Wang and co-authors from UC Berkeley and the Allen Institute for AI. The episode centers on a practical deployment question: if a workload is mostly code, math, or biomed, why must operators keep an entire giant model in memory instead of loading only the relevant slice? They frame EMO against the broader rise of sparse Mixture-of-Experts systems and explain why industry progress on active-parameter efficiency is not the same as delivering clean, domain-specific modules that can stand on their own at inference time. The discussion carefully separates standard MoE behavior from the stronger notion of modularity that EMO is targeting. Hal and Ada walk through how sparse-gated MoE and Switch Transformer style routing already allow different tokens to activate different experts, but argue that this still leaves deployment looking monolithic because the router makes local token-level decisions rather than exposing stable task-level components. A biology prompt can still scatter across a messy set of experts, and the next sentence may hit a different set entirely. The hosts use that distinction to unpack the paper’s core concepts: emergent modularity from unlabeled data, semantic expert specialization around meaningful domains like code or math, composable architecture, and the memory-accuracy frontier that determines whether smaller loaded expert pools can preserve real capability. The episode then gets into EMO’s training design and why the method is more than a single routing tweak. Ada explains the paper’s two-level routing scheme, where a document first selects a shared candidate pool of experts and individual tokens then choose active experts only within that pool, forcing document-consistent structure without removing all local flexibility. They also cover the supporting recipe: random pool-size sampling to expose the model to different memory budgets during training, global load balancing so a few experts do not dominate usage, and document-length-aware training so very long documents do not overwhelm the learning signal. The result is a focused discussion of whether MoE pretraining can produce expert groups that are not just sparsely activated, but genuinely deployable as modular tools.

Sources:
1. EMO: Emergent Modularity for Mixture-of-Experts
https://allenai.org/papers/emo
2. DEMix Layers: Disentangling Domains for Modular Language Modeling — Suchin Gururangan, Mike Lewis, Ari Holtzman, Noah A. Smith, Luke Zettlemoyer, 2021
https://scholar.google.com/scholar?q=DEMix+Layers%3A+Disentangling+Domains+for+Modular+Language+Modeling
3. ModuleFormer: Modularity Emerges from Mixture-of-Experts — Yikang Shen, Zheyu Zhang, Tianyou Cao, Shawn Tan, Zhenfang Chen, Chuang Gan, 2023
https://scholar.google.com/scholar?q=ModuleFormer%3A+Modularity+Emerges+from+Mixture-of-Experts
4. OpenMoE: An Early Effort on Open Mixture-of-Experts Language Models — Fuzhao Xue, Zian Zheng, Yao Fu, Jinjie Ni, Zangwei Zheng, Wangchunshu Zhou, Yang You, 2024
https://scholar.google.com/scholar?q=OpenMoE%3A+An+Early+Effort+on+Open+Mixture-of-Experts+Language+Models
5. EMO: Pretraining Mixture of Experts for Emergent Modularity — Ryan Wang, Akshita Bhagia, Sewon Min, 2026
https://scholar.google.com/scholar?q=EMO%3A+Pretraining+Mixture+of+Experts+for+Emergent+Modularity
6. FlexOlmo: Open Language Models for Flexible Data Use — Weijia Shi et al., 2025
https://scholar.google.com/scholar?q=FlexOlmo%3A+Open+Language+Models+for+Flexible+Data+Use
7. Branch-Train-MiX: Mixing Expert LLMs into a Mixture-of-Experts LLM — Sainbayar Sukhbaatar et al., 2024
https://scholar.google.com/scholar?q=Branch-Train-MiX%3A+Mixing+Expert+LLMs+into+a+Mixture-of-Experts+LLM
8. The Myth of Expert Specialization in MoEs: Why Routing Reflects Geometry, Not Necessarily Domain Expertise — Xi Wang, Soufiane Hayou, Eric Nalisnick, 2026
https://scholar.google.com/scholar?q=The+Myth+of+Expert+Specialization+in+MoEs%3A+Why+Routing+Reflects+Geometry%2C+Not+Necessarily+Domain+Expertise
9. Domain-Specific Pruning of Large Mixture-of-Experts Models with Few-shot Demonstrations — Zican Dong et al., 2025
https://scholar.google.com/scholar?q=Domain-Specific+Pruning+of+Large+Mixture-of-Experts+Models+with+Few-shot+Demonstrations
10. Efficient Expert Pruning for Sparse Mixture-of-Experts Language Models: Enhancing Performance and Reducing Inference Costs — Enshu Liu et al., 2024
https://scholar.google.com/scholar?q=Efficient+Expert+Pruning+for+Sparse+Mixture-of-Experts+Language+Models%3A+Enhancing+Performance+and+Reducing+Inference+Costs
11. MoE-Pruner: Pruning Mixture-of-Experts Large Language Model using the Hints from Its Router — Yanyue Xie et al., 2024
https://scholar.google.com/scholar?q=MoE-Pruner%3A+Pruning+Mixture-of-Experts+Large+Language+Model+using+the+Hints+from+Its+Router
12. AdaMoE: Token-Adaptive Routing with Null Experts for Mixture-of-Experts Language Models — Zihao Zeng et al., 2024
https://scholar.google.com/scholar?q=AdaMoE%3A+Token-Adaptive+Routing+with+Null+Experts+for+Mixture-of-Experts+Language+Models
13. DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models — Damai Dai et al., 2024
https://scholar.google.com/scholar?q=DeepSeekMoE%3A+Towards+Ultimate+Expert+Specialization+in+Mixture-of-Experts+Language+Models
14. Fast Inference of Mixture-of-Experts Language Models with Offloading — Artyom Eliseev and Denis Mazur, 2023
https://scholar.google.com/scholar?q=Fast+Inference+of+Mixture-of-Experts+Language+Models+with+Offloading
15. Accelerating Mixture-of-Experts Inference by Hiding Offloading Latency with Speculative Decoding — Zhibin Wang et al., 2025
https://scholar.google.com/scholar?q=Accelerating+Mixture-of-Experts+Inference+by+Hiding+Offloading+Latency+with+Speculative+Decoding
16. AI Post Transformers: EMO: Emergent Modularity in Sparse Language Models — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-06-emo-emergent-modularity-in-sparse-langua-9551c4.mp3
17. AI Post Transformers: Batch-Aware Expert Routing for Faster MoE Decoding — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-batch-aware-expert-routing-for-faster-mo-683ab6.mp3
18. AI Post Transformers: Nemotron 3 Super Hybrid Mamba-Transformer MoE — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-19-nemotron-3-super-hybrid-mamba-transforme-31ac75.mp3
19. AI Post Transformers: Memory-Bound, Not Bandwidth-Limited Batch-1 LLM Decode — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-02-memory-bound-not-bandwidth-limited-batch-114799.mp3
Interactive Visualization: EMO: Emergent Modularity for Mixture-of-Experts

This episode explores how a transformer trained on raw bank transaction histories can model customer behavior for financial product recommendation, and why that may outperform pipelines built from hand-engineered tabular features alone. It explains the paper’s core idea of turning each transaction into a tokenized sequence that mixes inflow or outflow, amount buckets, calendar signals, source metadata, and natural-language merchant descriptions, then pretraining the model with self-supervised learning to produce reusable customer embeddings. The discussion argues that transaction text and long-range patterns such as pay cycles, bill timing, and abrupt behavior changes carry signal that conventional tabular systems often flatten away, while a practical deployment can still combine learned embeddings with legacy banking features downstream. A listener would find it interesting because it connects transformer-style representation learning to a concrete banking use case and shows how foundation-model ideas can be adapted to messy, real-world financial behavior.

Sources:
1. Your Spending Needs Attention: Modeling Financial Habits with Transformers — D. T. Braithwaite, Misael Cavalcanti, R. Austin McEver, Hiroto Udagawa, Daniel Silva, Rohan Ramanath, Felipe Meneses, Arissa Yoshida, Evan Wingert, Matheus Ramos, Brian Zanfelice, Aman Gupta, 2025
http://arxiv.org/abs/2507.23267
2. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding — Jacob Devlin, Ming-Wei Chang, Kenton Lee, Kristina Toutanova, 2019
https://scholar.google.com/scholar?q=BERT%3A+Pre-training+of+Deep+Bidirectional+Transformers+for+Language+Understanding
3. CoLES: Contrastive Learning for Event Sequences with Self-Supervision — Dmitrii Babaev, Ivan Kireev, Nikita Ovsov, Mariya Ivanova, Gleb Gusev, Ivan Nazarov, Alexander Tuzhilin, 2020
https://scholar.google.com/scholar?q=CoLES%3A+Contrastive+Learning+for+Event+Sequences+with+Self-Supervision
4. Dynamic Customer Embeddings for Financial Service Applications — Nima Chitsazan, Samuel Sharpe, Dwipam Katariya, Qianyu Cheng, Karthik Rajasethupathy, 2021
https://scholar.google.com/scholar?q=Dynamic+Customer+Embeddings+for+Financial+Service+Applications
5. Towards a Foundation Purchasing Model: Pretrained Generative Autoregression on Transaction Sequences — Piotr Skalski, David Sutton, Stuart Burrell, Iker Perez, Jason Wong, 2024
https://scholar.google.com/scholar?q=Towards+a+Foundation+Purchasing+Model%3A+Pretrained+Generative+Autoregression+on+Transaction+Sequences
6. Self-Attentive Sequential Recommendation — Wang-Cheng Kang, Julian McAuley, 2018
https://scholar.google.com/scholar?q=Self-Attentive+Sequential+Recommendation
7. BERT4Rec: Sequential Recommendation with Bidirectional Encoder Representations from Transformer — Fei Sun, Jun Liu, Jian Wu, Changhua Pei, Xiao Lin, Wenwu Ou, Peng Jiang, 2019
https://scholar.google.com/scholar?q=BERT4Rec%3A+Sequential+Recommendation+with+Bidirectional+Encoder+Representations+from+Transformer
8. S^3-Rec: Self-Supervised Learning for Sequential Recommendation with Mutual Information Maximization — Kun Zhou, Hui Wang, Wayne Xin Zhao, Yutao Zhu, Sirui Wang, Fuzheng Zhang, Zhongyuan Wang, Ji-Rong Wen, 2020
https://scholar.google.com/scholar?q=S%5E3-Rec%3A+Self-Supervised+Learning+for+Sequential+Recommendation+with+Mutual+Information+Maximization
9. Behavior Sequence Transformer for E-commerce Recommendation in Alibaba — Qiwei Chen, Huan Zhao, Wei Li, Pipei Huang, Wenwu Ou, 2019
https://scholar.google.com/scholar?q=Behavior+Sequence+Transformer+for+E-commerce+Recommendation+in+Alibaba
10. Text Is All You Need: Learning Language Representations for Sequential Recommendation — Jiacheng Li, Ming Wang, Jin Li, Jinmiao Fu, Xin Shen, Jingbo Shang, Julian McAuley, 2023
https://scholar.google.com/scholar?q=Text+Is+All+You+Need%3A+Learning+Language+Representations+for+Sequential+Recommendation
11. PinnerFormer: Sequence Modeling for User Representation at Pinterest — Nikil Pancha, Andrew Zhai, Jure Leskovec, Charles Rosenberg, 2022
https://scholar.google.com/scholar?q=PinnerFormer%3A+Sequence+Modeling+for+User+Representation+at+Pinterest
12. Mamba4Rec: Towards Efficient Sequential Recommendation with Selective State Space Models — Chengkai Liu et al., 2024
https://scholar.google.com/scholar?q=Mamba4Rec%3A+Towards+Efficient+Sequential+Recommendation+with+Selective+State+Space+Models
13. SSD4Rec: A Structured State Space Duality Model for Efficient Sequential Recommendation — Haohao Qu et al., 2024
https://scholar.google.com/scholar?q=SSD4Rec%3A+A+Structured+State+Space+Duality+Model+for+Efficient+Sequential+Recommendation
14. DynLLM: When Large Language Models Meet Dynamic Graph Recommendation — Ziwei Zhao et al., 2024
https://scholar.google.com/scholar?q=DynLLM%3A+When+Large+Language+Models+Meet+Dynamic+Graph+Recommendation
15. Personalized Elastic Embedding Learning for On-Device Recommendation — Ruiqi Zheng et al., 2023
https://scholar.google.com/scholar?q=Personalized+Elastic+Embedding+Learning+for+On-Device+Recommendation
16. A Survey on Deep Tabular Learning — Shriyank Somvanshi et al., 2024
https://scholar.google.com/scholar?q=A+Survey+on+Deep+Tabular+Learning
17. AI Post Transformers: KumoRFM for In-Context Relational Learning — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-11-kumorfm-for-in-context-relational-learni-520d2b.mp3
18. AI Post Transformers: Gated Linear Attention for Efficient Long Sequences — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-18-gated-linear-attention-for-efficient-lon-c858ab.mp3

This episode explores TransactionGPT, a Visa Research paper that argues for a foundation-model approach to consumer transaction data spanning generation, anomaly detection, and representation learning. It explains why payment histories are fundamentally different from text or simple time series: each event mixes merchant IDs, amounts, timestamps, and engineered risk signals, creating a multi-modal, temporal, tabular structure that demands more schema-aware modeling. The discussion walks through the paper’s 1D, 2D, and 3D architecture progression, highlighting how separate transformers for transaction metadata, downstream features, and behavioral sequences aim to avoid the pitfalls of flattening everything into a single embedding space. Listeners would find it interesting for its clear debate over whether TransactionGPT is a genuine reusable backbone for payments or mainly a strong engineering response to real-world constraints like heterogeneity, scale, regulation, and low-latency fraud decisioning.

Sources:
1. TransactionGPT — Yingtong Dou, Zhimeng Jiang, Tianyi Zhang, Mingzhi Hu, Zhichao Xu, Shubham Jain, Uday Singh Saini, Xiran Fan, Jiarui Sun, Menghai Pan, Junpeng Wang, Xin Dai, Liang Wang, Chin-Chia Michael Yeh, Yujie Fan, Yan Zheng, Vineeth Rakesh, Huiyuan Chen, Guanchu Wang, Mangesh Bendre, Zhongfang Zhuang, Xiaoting Li, Prince Aboagye, Vivian Lai, Minghua Xu, Hao Yang, Yiwei Cai, Mahashweta Das, Yuzhong Chen, 2025
http://arxiv.org/abs/2511.08939
2. TabDPT: Scaling Tabular Foundation Models on Real Data — Junwei Ma, Valentin Thomas, Rasa Hosseinzadeh, Alex Labach, Hamidreza Kamkari, Jesse C. Cresswell, Keyvan Golestan, Guangwei Yu, Anthony L. Caterini, Maksims Volkovs, 2024
https://arxiv.org/abs/2410.18164
3. Towards a Foundation Purchasing Model: Pretrained Generative Autoregression on Transaction Sequences — Piotr Skalski, David Sutton, Stuart Burrell, Iker Perez, Jason Wong, 2024
https://arxiv.org/abs/2401.01641
4. Open Banking Foundational Model: Learning Language Representations from Few Financial Transactions — Gustavo Polleti, Marlesson Santana, Eduardo Fontes, 2025
https://arxiv.org/abs/2511.12154
5. TransactionGPT — Yingtong Dou, Zhimeng Jiang, Tianyi Zhang, Mingzhi Hu, Zhichao Xu, Shubham Jain, Uday Singh Saini, Xiran Fan, Jiarui Sun, Menghai Pan, Junpeng Wang, Xin Dai, Liang Wang, Chin-Chia Michael Yeh, Yujie Fan, Vineeth Rakesh, Huiyuan Chen, Mangesh Bendre, Zhongfang Zhuang, Xiaoting Li, Prince Aboagye, Vivian Lai, Minghua Xu, Hao Yang, Yiwei Cai, Mahashweta Das, Yuzhong Chen, 2025
https://arxiv.org/abs/2511.08939
6. TabTransformer: Tabular Data Modeling Using Contextual Embeddings — Xin Huang, Ashish Khetan, Milan Cvitkovic, Zohar Karnin, 2020
https://arxiv.org/abs/2012.06678
7. Tabular Transformers for Modeling Multivariate Time Series — Inkit Padhi, Yair Schiff, Igor Melnyk, Mattia Rigotti, Youssef Mroueh, Pierre Dognin, Jerret Ross, Ravi Nair, Erik Altman, 2021
https://arxiv.org/abs/2011.01843
8. FATA-Trans: Field And Time-Aware Transformer for Sequential Tabular Data — Dongyu Zhang, Liang Wang, Xin Dai, Shubham Jain, Junpeng Wang, Yujie Fan, Chin-Chia Michael Yeh, Yan Zheng, Zhongfang Zhuang, Wei Zhang, 2023
https://arxiv.org/abs/2310.13818
9. Multi-modal Time Series Analysis: A Tutorial and Survey — Yushan Jiang, Kanghui Ning, Zijie Pan, Xuyang Shen, Jingchao Ni, Wenchao Yu, Anderson Schneider, Haifeng Chen, Yuriy Nevmyvaka, Dongjin Song, 2025
https://arxiv.org/abs/2503.13709
10. Credit card fraud detection using machine learning: A survey — Yvan Lucas, Johannes Jurgovsky, 2020
https://arxiv.org/abs/2010.06479
11. Interleaved Sequence RNNs for Fraud Detection — Bernardo Branco, Pedro Abreu, Ana Sofia Gomes, Mariana S. C. Almeida, João Tiago Ascensão, Pedro Bizarro, 2020
https://arxiv.org/abs/2002.05988
12. FraudTransformer: Time-Aware GPT for Transaction Fraud Detection — Gholamali Aminian, Andrew Elliott, Tiger Li, Timothy Cheuk Hin Wong, Victor Claude Dehon, Lukasz Szpruch, Carsten Maple, Christopher Read, Martin Brown, Gesine Reinert, Mo Mamouei, 2025
https://arxiv.org/abs/2509.23712
13. Enhancing Foundation Models in Transaction Understanding with LLM-based Sentence Embeddings — Xiran Fan, Zhimeng Jiang, Chin-Chia Michael Yeh, Yuzhong Chen, Yingtong Dou, Menghai Pan, Yan Zheng, 2025
https://scholar.google.com/scholar?q=Enhancing+Foundation+Models+in+Transaction+Understanding+with+LLM-based+Sentence+Embeddings
14. Self-Attentive Sequential Recommendation — Wang-Cheng Kang, Julian McAuley, 2018
https://scholar.google.com/scholar?q=Self-Attentive+Sequential+Recommendation
15. TabLLM: Few-shot Classification of Tabular Data with Large Language Models — Stefan Hegselmann, Alejandro Buendia, Hunter Lang, Monica Agrawal, Xiaoyi Jiang, David Sontag, 2022
https://scholar.google.com/scholar?q=TabLLM%3A+Few-shot+Classification+of+Tabular+Data+with+Large+Language+Models
16. TimeGPT-1 — Azul Garza, Cristian Challu, Max Mergenthaler-Canseco, 2023
https://scholar.google.com/scholar?q=TimeGPT-1
17. ForkMerge: Mitigating Negative Transfer in Auxiliary-Task Learning — Junguang Jiang et al., 2023
https://scholar.google.com/scholar?q=ForkMerge%3A+Mitigating+Negative+Transfer+in+Auxiliary-Task+Learning
18. Identification of Negative Transfers in Multitask Learning Using Surrogate Models — Dongyue Li, Huy L. Nguyen, Hongyang R. Zhang, 2023
https://scholar.google.com/scholar?q=Identification+of+Negative+Transfers+in+Multitask+Learning+Using+Surrogate+Models
19. Enriching Tabular Data with Contextual LLM Embeddings: A Comprehensive Ablation Study for Ensemble Classifiers — Gjergji Kasneci, Enkelejda Kasneci, 2024
https://scholar.google.com/scholar?q=Enriching+Tabular+Data+with+Contextual+LLM+Embeddings%3A+A+Comprehensive+Ablation+Study+for+Ensemble+Classifiers
20. LLM Embeddings Improve Test-time Adaptation to Tabular Y|X-Shifts — Yibo Zeng et al., 2024
https://scholar.google.com/scholar?q=LLM+Embeddings+Improve+Test-time+Adaptation+to+Tabular+Y%7CX-Shifts
21. LLM Embeddings for Deep Learning on Tabular Data — Boshko Koloski et al., 2025
https://scholar.google.com/scholar?q=LLM+Embeddings+for+Deep+Learning+on+Tabular+Data
22. TableGPT2: A Large Multimodal Model with Tabular Data Integration — Aofeng Su et al., 2024
https://scholar.google.com/scholar?q=TableGPT2%3A+A+Large+Multimodal+Model+with+Tabular+Data+Integration
23. AI Post Transformers: KumoRFM for In-Context Relational Learning — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-11-kumorfm-for-in-context-relational-learni-520d2b.mp3
24. AI Post Transformers: Relational Graph Transformer for Multi-Table Learning — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-11-relational-graph-transformer-for-multi-t-57cce3.mp3
25. AI Post Transformers: Unembedding Matrices as Feature Lenses for Embeddings — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-08-unembedding-matrices-as-feature-lenses-f-4dc415.mp3
Interactive Visualization: TransactionGPT as a Payments Foundation Model

This episode explores the paper When Does LeJEPA Learn a World Model? and uses it to examine what should count as a genuine world model in latent predictive learning, contrasting JEPA-style representation prediction with generative reconstruction. It explains why good probe scores are not enough: the real standard is linear identifiability, where a single global linear map recovers the environment’s hidden state well enough to support planning and compositional generalization. The discussion centers on the paper’s main theorem that, under stationary additive-noise dynamics with Gaussian latent variables, LeJEPA’s alignment objective plus SIGReg recovers the true latent state up to an orthogonal rotation, and on the sharper converse result that this universal guarantee fails for non-Gaussian latents. Listeners get a rigorous argument for when latent models are truly learning the world’s coordinates instead of merely extracting features that happen to be useful on downstream tasks.

Interactive Visualization: When LeJEPA Truly Learns a World Model
Sources:
1. When LeJEPA Truly Learns a World Model
https://arxiv.org/pdf/2605.26379
2. Challenging Common Assumptions in the Unsupervised Learning of Disentangled Representations — Francesco Locatello, Stefan Bauer, Mario Lucic, et al., 2018
https://scholar.google.com/scholar?q=Challenging+Common+Assumptions+in+the+Unsupervised+Learning+of+Disentangled+Representations
3. Variational Autoencoders and Nonlinear ICA: A Unifying Framework — Ilyes Khemakhem, Diederik P. Kingma, Ricardo Pio Monti, Aapo Hyvarinen, 2019
https://scholar.google.com/scholar?q=Variational+Autoencoders+and+Nonlinear+ICA%3A+A+Unifying+Framework
4. On Linear Identifiability of Learned Representations — Geoffrey Roeder, Luke Metz, Diederik P. Kingma, 2020
https://scholar.google.com/scholar?q=On+Linear+Identifiability+of+Learned+Representations
5. Nonlinear Independent Component Analysis for Principled Disentanglement in Unsupervised Deep Learning — Aapo Hyvarinen, Ilyes Khemakhem, Hiroshi Morioka, 2023
https://scholar.google.com/scholar?q=Nonlinear+Independent+Component+Analysis+for+Principled+Disentanglement+in+Unsupervised+Deep+Learning
6. Auto-Encoding Variational Bayes — Diederik P. Kingma, Max Welling, 2013
https://scholar.google.com/scholar?q=Auto-Encoding+Variational+Bayes
7. VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised Learning — Adrien Bardes, Jean Ponce, Yann LeCun, 2021
https://scholar.google.com/scholar?q=VICReg%3A+Variance-Invariance-Covariance+Regularization+for+Self-Supervised+Learning
8. LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics — Randall Balestriero, Yann LeCun, 2025
https://scholar.google.com/scholar?q=LeJEPA%3A+Provable+and+Scalable+Self-Supervised+Learning+Without+the+Heuristics
9. When Does LeJEPA Learn a World Model? — David Klindt, Yann LeCun, Randall Balestriero, 2026
https://scholar.google.com/scholar?q=When+Does+LeJEPA+Learn+a+World+Model%3F
10. LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels — Lucas Maes, Quentin Le Lidec, Damien Scieur, Yann LeCun, Randall Balestriero, 2026
https://scholar.google.com/scholar?q=LeWorldModel%3A+Stable+End-to-End+Joint-Embedding+Predictive+Architecture+from+Pixels
11. V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning — Mido Assran et al., 2025
https://scholar.google.com/scholar?q=V-JEPA+2%3A+Self-Supervised+Video+Models+Enable+Understanding%2C+Prediction+and+Planning
12. Nonlinear ICA Using Auxiliary Variables and Generalized Contrastive Learning — Aapo Hyvarinen, Hiroaki Sasaki, Richard E. Turner, 2018
https://scholar.google.com/scholar?q=Nonlinear+ICA+Using+Auxiliary+Variables+and+Generalized+Contrastive+Learning
13. Joint Embedding Predictive Architectures Focus on Slow Features — Vlad Sobal, Jyothir S V, Siddhartha Jalagam, Nicolas Carion, Kyunghyun Cho, Yann LeCun, 2022
https://scholar.google.com/scholar?q=Joint+Embedding+Predictive+Architectures+Focus+on+Slow+Features
14. Cross-Entropy Is All You Need To Invert the Data Generating Process — Patrik Reizinger, Alice Bizeul, Attila Juhos, Julia E. Vogt, Randall Balestriero, Wieland Brendel, David Klindt, 2024
https://scholar.google.com/scholar?q=Cross-Entropy+Is+All+You+Need+To+Invert+the+Data+Generating+Process
15. Identifiability of latent-variable and structural-equation models: from linear to nonlinear — Aapo Hyvarinen, Ilyes Khemakhem, Ricardo Monti, 2023
https://scholar.google.com/scholar?q=Identifiability+of+latent-variable+and+structural-equation+models%3A+from+linear+to+nonlinear
16. On the Identifiability of Sparse ICA without Assuming Non-Gaussianity — Ignavier Ng, Yujia Zheng, Xinshuai Dong, Kun Zhang, 2024
https://scholar.google.com/scholar?q=On+the+Identifiability+of+Sparse+ICA+without+Assuming+Non-Gaussianity
17. Adaptive World Models: Learning Behaviors by Latent Imagination Under Non-Stationarity — Emiliyan Gospodinov, Vaisakh Shaj, Philipp Becker, Stefan Geyer, Gerhard Neumann, 2024
https://scholar.google.com/scholar?q=Adaptive+World+Models%3A+Learning+Behaviors+by+Latent+Imagination+Under+Non-Stationarity
18. Koopa: Learning Non-stationary Time Series Dynamics with Koopman Predictors — Yong Liu, Chenyu Li, Jianmin Wang, Mingsheng Long, 2023
https://scholar.google.com/scholar?q=Koopa%3A+Learning+Non-stationary+Time+Series+Dynamics+with+Koopman+Predictors
19. Simplifying Latent Dynamics with Softly State-Invariant World Models — Tankred Saanum, Peter Dayan, Eric Schulz, 2024
https://scholar.google.com/scholar?q=Simplifying+Latent+Dynamics+with+Softly+State-Invariant+World+Models
20. Structured World Models from Human Videos — Russell Mendonca, Shikhar Bahl, Deepak Pathak, 2023
https://scholar.google.com/scholar?q=Structured+World+Models+from+Human+Videos
21. AI Post Transformers: LeWorldModel: Stable Joint-Embedding World Models from Pixels — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-25-leworldmodel-stable-joint-embedding-worl-650f9f.mp3
22. AI Post Transformers: Causal-JEPA for Object-Level World Models — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-causal-jepa-for-object-level-world-model-311a8b.mp3
23. AI Post Transformers: Learning Latent Action World Models from Video — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-09-learning-latent-action-world-models-from-1570a4.mp3
Interactive Visualization: When LeJEPA Truly Learns a World Model

Hal Turing and Dr. Ada Shannon examine an empirical study of parameter-efficient fine-tuning for large language models, centered on a practical question: when does a small task-specific update beat retraining the entire model? Using FLAN-T5-XL as the test bed, they frame PEFT as a transfer-learning strategy that freezes most of the transformer while learning a compact adaptation layer, whether through LoRA’s low-rank weight updates, adapter-style modules, IA3 scaling vectors, BitFit bias updates, or learned soft prompts. The discussion keeps returning to the real systems tradeoff: quality matters, but so do training speed, storage cost, and the burden of maintaining separate model copies for many downstream tasks. The episode walks through the benchmark design in detail rather than treating PEFT as a vague category. The paper compares full tuning, LoRA, IA3, prompt tuning, and BitFit on the same backbone across classification tasks like AG News and CoLA, generation tasks like E2E and SAMSum, and data budgets of roughly 100, 1,000, and 10,000 examples. The hosts emphasize why those controls matter: same model, same stopping rule, and fixed method settings make it easier to see where each technique actually helps, while also limiting how far the results should be generalized to other architectures, especially decoder-only chat models. They then dig into the paper’s uneven but useful results. In low-resource settings, LoRA and BitFit frequently outperform full tuning, with LoRA posting a notably stronger CoLA score and BitFit leading on AG News, E2E, and SAMSum, while prompt tuning performs strikingly poorly on the generation benchmarks under this setup. In medium-resource settings, IA3, LoRA, and BitFit remain competitive, but at higher data scales full tuning starts reclaiming ground on some tasks even as LoRA and IA3 still win specific cases. The takeaway is not that one PEFT method universally dominates, but that the strengths and weaknesses of each approach shift with task type, data regime, and the exact adaptation recipe.

Sources:
1. Empirical Analysis of the Strengths and Weaknesses of PEFT Techniques for LLMs — George Pu, Anirudh Jain, Jihan Yin, Russell Kaplan, 2023
http://arxiv.org/abs/2304.14999
2. Parameter-Efficient Transfer Learning for NLP — Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, et al., 2019
https://arxiv.org/abs/1902.00751
3. LoRA: Low-Rank Adaptation of Large Language Models — Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, et al., 2021
https://arxiv.org/abs/2106.09685
4. Scaling Down to Scale Up: A Guide to Parameter-Efficient Fine-Tuning — Vladislav Lialin, Vijeta Deshpande, Xiaowei Yao, Anna Rumshisky, 2023
https://arxiv.org/abs/2303.15647
5. Empirical Analysis of the Strengths and Weaknesses of PEFT Techniques for LLMs — George Pu, Anirudh Jain, Jihan Yin, Russell Kaplan, 2023
https://arxiv.org/abs/2304.14999
6. Prefix-Tuning: Optimizing Continuous Prompts for Generation — Xiang Lisa Li, Percy Liang, 2021
https://arxiv.org/abs/2101.00190
7. The Power of Scale for Parameter-Efficient Prompt Tuning — Brian Lester, Rami Al-Rfou, Noah Constant, 2021
https://arxiv.org/abs/2104.08691
8. P-Tuning v2: Prompt Tuning Can Be Comparable to Fine-tuning Universally Across Scales and Tasks — Xiao Liu, Kaixuan Ji, Yicheng Fu, Weng Lam Tam, Zhengxiao Du, Zhilin Yang, Jie Tang, 2021
https://arxiv.org/abs/2110.07602
9. SPoT: Better Frozen Model Adaptation through Soft Prompt Transfer — Tu Vu, Brian Lester, Noah Constant, Rami Al-Rfou, Daniel Cer, 2021
https://arxiv.org/abs/2110.07904
10. Revisiting Parameter-Efficient Tuning: Are We Really There Yet? — Guanzheng Chen, Fangyu Liu, Zaiqiao Meng, Shangsong Liang, 2022
https://scholar.google.com/scholar?q=Revisiting+Parameter-Efficient+Tuning%3A+Are+We+Really+There+Yet%3F
11. Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context Learning — Haokun Liu, Derek Tam, Mohammed Muqeeth, Jay Mohta, Tenghao Huang, Mohit Bansal, Colin Raffel, 2022
https://scholar.google.com/scholar?q=Few-Shot+Parameter-Efficient+Fine-Tuning+is+Better+and+Cheaper+than+In-Context+Learning
12. On the Effectiveness of Adapter-based Tuning for Pretrained Language Model Adaptation — Ruidan He, Linlin Liu, Hai Ye, Qingyu Tan, Bosheng Ding, Liying Cheng, Jia-Wei Low, Lidong Bing, Luo Si, 2021
https://scholar.google.com/scholar?q=On+the+Effectiveness+of+Adapter-based+Tuning+for+Pretrained+Language+Model+Adaptation
13. Scaling Instruction-Finetuned Language Models — Hyung Won Chung et al., 2022
https://scholar.google.com/scholar?q=Scaling+Instruction-Finetuned+Language+Models
14. When Scaling Meets LLM Finetuning: The Effect of Data, Model and Finetuning Method — Biao Zhang, Zhongtao Liu, Colin Cherry, Orhan Firat, 2024
https://arxiv.org/abs/2402.17193
15. Parameter-Efficient Fine-Tuning Design Spaces — Jiaao Chen, Aston Zhang, Xingjian Shi, Mu Li, Alex Smola, Diyi Yang, 2023
https://arxiv.org/abs/2301.01821
16. AI Post Transformers: ForkKV for Multi-LoRA Agent Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-13-forkkv-for-multi-lora-agent-serving-ccafa4.mp3
17. AI Post Transformers: Mooncake for KV Cache-Centric LLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-05-mooncake-for-kv-cache-centric-llm-servin-1086d0.mp3
18. AI Post Transformers: ZeRO-Offload: Democratizing Billion-Scale Model Training — Hal Turing & Dr. Ada Shannon, Fri,
https://podcast.do-not-panic.com/episodes/zero-offload-democratizing-billion-scale-model-training/
Interactive Visualization: Benchmarking PEFT Techniques for Large Language Models

This episode explores Weak-SIGReg, a lightweight covariance regularizer designed to prevent representation collapse in fragile supervised training, especially for small-data Vision Transformers. It explains how the method uses a sketched covariance matrix and an identity-matching penalty to keep hidden features decorrelated and similarly scaled at much lower cost than full covariance regularization. The discussion centers on CIFAR-100 results, where Weak-SIGReg dramatically improves a deliberately unstable ViT setup and also boosts a plain MLP, while offering little change on an already stable ResNet18. It also digs into the paper’s main caveat: the biggest gains appear when the baseline training recipe is badly broken, so the most interesting question is not just whether the regularizer works, but when it adds real value beyond simply fixing optimization and initialization.

Sources:
1. Weak-SIGReg: Covariance Regularization for Stable Deep Learning — Habibullah Akbar, 2026
http://arxiv.org/abs/2603.05924
2. Reducing Overfitting in Deep Networks by Decorrelating Representations — Michael Cogswell, Faruk Ahmed, Ross Girshick, Larry Zitnick, Dhruv Batra, 2015
https://scholar.google.com/scholar?q=Reducing+Overfitting+in+Deep+Networks+by+Decorrelating+Representations
3. Barlow Twins: Self-Supervised Learning via Redundancy Reduction — Jure Zbontar, Li Jing, Ishan Misra, Yann LeCun, Stephane Deny, 2021
https://scholar.google.com/scholar?q=Barlow+Twins%3A+Self-Supervised+Learning+via+Redundancy+Reduction
4. VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised Learning — Adrien Bardes, Jean Ponce, Yann LeCun, 2021
https://scholar.google.com/scholar?q=VICReg%3A+Variance-Invariance-Covariance+Regularization+for+Self-Supervised+Learning
5. LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics — Randall Balestriero, Yann LeCun, 2025
https://scholar.google.com/scholar?q=LeJEPA%3A+Provable+and+Scalable+Self-Supervised+Learning+Without+the+Heuristics
6. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale — Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, et al., 2021
https://scholar.google.com/scholar?q=An+Image+is+Worth+16x16+Words%3A+Transformers+for+Image+Recognition+at+Scale
7. Training data-efficient image transformers & distillation through attention — Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, Herve Jegou, 2021
https://scholar.google.com/scholar?q=Training+data-efficient+image+transformers+%26+distillation+through+attention
8. How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers — Andreas Steiner, Alexander Kolesnikov, Xiaohua Zhai, Ross Wightman, Jakob Uszkoreit, Lucas Beyer, 2021
https://scholar.google.com/scholar?q=How+to+train+your+ViT%3F+Data%2C+Augmentation%2C+and+Regularization+in+Vision+Transformers
9. Early Convolutions Help Transformers See Better — Tete Xiao, Mannat Singh, Eric Mintun, Trevor Darrell, Piotr Dollar, Ross Girshick, 2021
https://scholar.google.com/scholar?q=Early+Convolutions+Help+Transformers+See+Better
10. Whitening for Self-Supervised Representation Learning — Aleksandr Ermolov, Aliaksandr Siarohin, Enver Sangineto, Nicu Sebe, 2020
https://arxiv.org/abs/2007.06346
11. Understanding Dimensional Collapse in Contrastive Self-Supervised Learning — Li Jing, Pascal Vincent, Yann LeCun, Yuandong Tian, 2021
https://arxiv.org/abs/2110.09348
12. UR-JEPA: Uniform Rectifiability as a Regularizer for Joint-Embedding Predictive Architectures — Triet M. Le, 2026
https://arxiv.org/abs/2606.01443
13. Stability of Transformers under Layer Normalization — Kelvin Kan, Xingjian Li, Benjamin J. Zhang, Tuhin Sahai, Stanley Osher, Krishna Kumar, Markos A. Katsoulakis, 2025
https://scholar.google.com/scholar?q=Stability+of+Transformers+under+Layer+Normalization
14. Transformers without Normalization — Jiachen Zhu, Xinlei Chen, Kaiming He, Yann LeCun, Zhuang Liu, 2025
https://scholar.google.com/scholar?q=Transformers+without+Normalization
15. Are Neurons Actually Collapsed? On the Fine-Grained Structure in Neural Representations — Yongyi Yang, Jacob Steinhardt, Wei Hu, 2023
https://scholar.google.com/scholar?q=Are+Neurons+Actually+Collapsed%3F+On+the+Fine-Grained+Structure+in+Neural+Representations
16. The Impact of Geometric Complexity on Neural Collapse in Transfer Learning — Michael Munn, Benoit Dherin, Javier Gonzalvo, 2024
https://scholar.google.com/scholar?q=The+Impact+of+Geometric+Complexity+on+Neural+Collapse+in+Transfer+Learning
17. AI Post Transformers: LeWorldModel: Stable Joint-Embedding World Models from Pixels — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-25-leworldmodel-stable-joint-embedding-worl-650f9f.mp3
18. AI Post Transformers: Muon Is Scalable for LLM Training — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-25-muon-is-scalable-for-llm-training-587ed8.mp3

This episode explores InfiniGen, a systems approach to speeding up long-context language model inference by treating KV cache management, not raw compute, as the central bottleneck. It explains why decoding slows down when large caches have to shuttle between CPU and GPU memory, and contrasts InfiniGen with FlexGen, H2O, and PagedAttention to show how different serving setups create different memory problems. The discussion focuses on InfiniGen’s core idea: use a lightweight preview from the previous layer, along with offline-skewed query and key weights, to predict which exact cache entries will matter next and prefetch only those instead of moving the whole history. Listeners would find it interesting because the paper reports large practical gains, including up to 3x speedups and major accuracy improvements over weaker cache-selection methods, making it a concrete example of systems engineering reshaping how large models are served.

Sources:
1. InfiniGen: Efficient Generative Inference of Large Language Models with Dynamic KV Cache Management — Wonbeom Lee, Jungi Lee, Junghwan Seo, Jaewoong Sim, 2024
http://arxiv.org/abs/2406.19707
2. FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU — Ying Sheng et al., 2023
https://scholar.google.com/scholar?q=FlexGen%3A+High-Throughput+Generative+Inference+of+Large+Language+Models+with+a+Single+GPU
3. H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models — Zhenyu Zhang et al., 2023
https://scholar.google.com/scholar?q=H2O%3A+Heavy-Hitter+Oracle+for+Efficient+Generative+Inference+of+Large+Language+Models
4. Efficient Memory Management for Large Language Model Serving with PagedAttention — Woosuk Kwon et al., 2023
https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention
5. Efficient Streaming Language Models with Attention Sinks — Guangxuan Xiao et al., 2023
https://scholar.google.com/scholar?q=Efficient+Streaming+Language+Models+with+Attention+Sinks
6. DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads — Guangxuan Xiao et al., 2024
https://scholar.google.com/scholar?q=DuoAttention%3A+Efficient+Long-Context+LLM+Inference+with+Retrieval+and+Streaming+Heads
7. SCBench: A KV Cache-Centric Analysis of Long-Context Methods — Yucheng Li et al., 2024
https://scholar.google.com/scholar?q=SCBench%3A+A+KV+Cache-Centric+Analysis+of+Long-Context+Methods
8. IndexMem: Learned KV-Cache Eviction with Latent Memory for Long-Context LLM Inference — Xintong Yang et al., 2026
https://arxiv.org/abs/2605.25475
9. Model Tells You Where to Merge: Adaptive KV Cache Merging for LLMs on Long-Context Tasks — Zheng Wang et al., 2024
https://arxiv.org/abs/2407.08454
10. KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization — Coleman Hooper et al., 2024
https://arxiv.org/abs/2401.18079
11. TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization — Dingyu Yao et al., 2025
https://arxiv.org/abs/2505.19586
12. LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management — Yi Xiong et al., 2024
https://arxiv.org/abs/2410.00428
13. SqueezeAttention: 2D Management of KV-Cache in LLM Inference via Layer-wise Optimal Budget — Zihao Wang et al., 2024
https://arxiv.org/abs/2404.04793
14. KeDiff: Key Similarity-Based KV Cache Eviction for Long-Context LLM Inference in Resource-Constrained Environments — Junyoung Park et al., 2025
https://arxiv.org/abs/2504.15364
15. AI Post Transformers: IndexMem: Learned KV-Cache Eviction for Long-Context LLMs — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-12-indexmem-learned-kv-cache-eviction-for-l-132c2a.mp3
16. AI Post Transformers: PackKV Lossy Compression for KV Caches — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-04-packkv-lossy-compression-for-kv-caches-b37bce.mp3
17. AI Post Transformers: Speculative Decoding in Real vLLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-speculative-decoding-in-real-vllm-servin-6f4e2b.mp3
18. AI Post Transformers: Memory-Bound, Not Bandwidth-Limited Batch-1 LLM Decode — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-02-memory-bound-not-bandwidth-limited-batch-114799.mp3
19. AI Post Transformers: FengHuang for Rack-Scale LLM Inference Memory — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-12-fenghuang-for-rack-scale-llm-inference-m-62708e.mp3
20. AI Post Transformers: Deep Kernel Fusion for Transformer Decoding — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-deep-kernel-fusion-for-transformer-decod-b1a703.mp3
21. AI Post Transformers: QuantSpec: Hierarchical KV Cache for Self-Speculative Decoding — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/quantspec-hierarchical-kv-cache-for-self-speculative-decoding/

This episode explores how quantization affects reasoning models, asking how much weights, activations, and KV caches can be compressed before multi-step reasoning starts to fail. It explains the main quantization strategies in practical serving terms, from weight-only methods like AWQ and GPTQ to weight-activation schemes such as W8A8 and W4A4, and KV cache compression for long decoding traces. The discussion argues that reasoning models are unusually fragile because small numerical errors can compound across long solution paths, making calibration quality and benchmark choice far more important than they are for ordinary chat models. Listeners would find it interesting for its concrete look at the tradeoff between cheaper inference and reliable reasoning, grounded in evaluations across model families from 1.5B to 70B and difficult benchmarks in math, science, and code.

Sources:
1. Quantization Hurts Reasoning? An Empirical Study on Quantized Reasoning Models — Ruikang Liu, Yuxuan Sun, Manyi Zhang, Haoli Bai, Xianzhi Yu, Tiezheng Yu, Chun Yuan, Lu Hou, 2025
http://arxiv.org/abs/2504.04823
2. SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models — Guangxuan Xiao, Ji Lin, Mickael Seznec, Hao Wu, Julien Demouth, Song Han, 2023
https://scholar.google.com/scholar?q=SmoothQuant%3A+Accurate+and+Efficient+Post-Training+Quantization+for+Large+Language+Models
3. AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration — Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Wei-Ming Chen, Wei-Chen Wang, Guangxuan Xiao, Xingyu Dang, Chuang Gan, Song Han, 2024
https://scholar.google.com/scholar?q=AWQ%3A+Activation-aware+Weight+Quantization+for+LLM+Compression+and+Acceleration
4. KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization — Coleman Hooper, Sehoon Kim, Hiva Mohammadzadeh, Michael W. Mahoney, Yakun Sophia Shao, Kurt Keutzer, Amir Gholami, 2024
https://scholar.google.com/scholar?q=KVQuant%3A+Towards+10+Million+Context+Length+LLM+Inference+with+KV+Cache+Quantization
5. Quantization Meets Reasoning: Exploring LLM Low-Bit Quantization Degradation for Mathematical Reasoning — Zhen Li, Yupeng Su, Runming Yang, Congkai Xie, Zheng Wang, Zhongwei Xie, Ngai Wong, Hongxia Yang, 2025
https://scholar.google.com/scholar?q=Quantization+Meets+Reasoning%3A+Exploring+LLM+Low-Bit+Quantization+Degradation+for+Mathematical+Reasoning
6. Evaluating Quantized Large Language Models — Shiyao Li et al., 2024
https://scholar.google.com/scholar?q=Evaluating+Quantized+Large+Language+Models
7. FlatQuant: Flatness Matters for LLM Quantization — Yuxuan Sun et al., 2024
https://scholar.google.com/scholar?q=FlatQuant%3A+Flatness+Matters+for+LLM+Quantization
8. s1: Simple Test-Time Scaling — Niklas Muennighoff et al., 2025
https://scholar.google.com/scholar?q=s1%3A+Simple+Test-Time+Scaling
9. What Makes Low-Bit Quantization-Aware Training Work for Reasoning LLMs? A Systematic Study — Keyu Lv et al., 2026
https://arxiv.org/abs/2601.14888
10. Measuring Faithfulness Depends on How You Measure: Classifier Sensitivity in LLM Chain-of-Thought Evaluation — Richard J. Young, 2026
https://arxiv.org/abs/2603.20172
11. On the Hardness of Faithful Chain-of-Thought Reasoning in Large Language Models — Sree Harsha Tanneru et al., 2024
https://arxiv.org/abs/2406.10625
12. Activation Outliers in Transformer Quantization: Reproduction, Statistical Analysis, and Deployment Tradeoffs — Pranav Kumar Kaliaperumal, 2026
https://arxiv.org/abs/2603.04308
13. KVCOMM: Online Cross-context KV-cache Communication for Efficient LLM-based Multi-agent Systems — Hancheng Ye et al., 2025
https://arxiv.org/abs/2510.12872
14. AI Post Transformers: Affordable Large-Scale Decoding Through Model-System Co-Design — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-19-affordable-large-scale-decoding-through-e1d7ed.mp3
15. AI Post Transformers: Mooncake for KV Cache-Centric LLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-05-mooncake-for-kv-cache-centric-llm-servin-1086d0.mp3
16. AI Post Transformers: IndexMem: Learned KV-Cache Eviction for Long-Context LLMs — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-12-indexmem-learned-kv-cache-eviction-for-l-132c2a.mp3
17. AI Post Transformers: Speculative Decoding in Real vLLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-speculative-decoding-in-real-vllm-servin-6f4e2b.mp3

This episode explores Inclusion AI’s Ling and Ring 2.6 technical report, which asks how a trillion-parameter model can stay fast, handle very long contexts, and remain dependable in multi-step agent workflows. It explains why agentic AI makes latency and token costs much more painful than in ordinary chat, especially when models must carry long instruction traces, tool outputs, and large working contexts through repeated reasoning loops. The discussion breaks down the report’s core architectural changes, including a Lightning Attention and MLA hybrid with a 7:1 layer mix, designed to reduce attention cost and KV-cache memory without sacrificing model quality. It also examines the practical significance of retrofitting an existing trillion-scale checkpoint through continued pretraining and staged migration techniques, making the episode especially interesting for listeners who want a concrete look at how frontier model design is shifting from raw scale toward deployable systems engineering.

Sources:
1. Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale — Ang Li, Ben Liu, Bin Han, Bin Hu, Bin Jing, Binbin Hu, Bing Li, Cai Chen, Caizhi Tang, Changxin Tian, Chao Huang, Chao Zhang, Chen Liang, Chen Qian, Chengfu Tang, Chengyao Wen, Chilin Fu, Chunwei Wu, Cong Zhang, Cunyin Peng, Daixin Wang, Dalong Zhang, Deng Zhao, Dingnan Jin, Dingyuan Zhu, Donghao Zhang, Fan Yuan, Fangzheng Zhao, Fanzhuang Meng, Feifan Wu, Feng Xu, Fengbin Fang, Gangshan Wang, Guodong Yang, Hailin Zhao, Haitao Wang, Haitao Zhang, Hanxiao Zhang, Hanzi Wang, Hao Dai, Hao Liu, Hao Qian, Hao Wu, Haoxiong Liu, Haoyu Xu, Heng Zhang, Hong Liu, Hongliang Zhang, Hongrui Liu, Hongxun Li, Hongzhi Ruan, Huaidong Xiong, Huihuang Zheng, Huikang Tang, Jia Guo, Jia Li, Jia Liu, Jiameng Wang, Jiaming Liu, Jiannan Shi, Jianping Wei, Jiaolong Yang, Jiapeng Wang, Jie Gao, Jie Wang, Jiewei Wu, Jin Yang, Jinjin Li, Jinjing Huang, Jinquan Sun, Jinyao Chen, Juanhui Tu, Jun Liu, Jun Mei, Jun Xu, Jun Zhou, Junjie Ou, Junnan Sipan, Junpeng Fang, Kaihong Zhang, Kaiqin Hu, Ke Shi, Kuan Xu, Kun Tang, Kunlong Chen, Lanyin Mei, Lei Chen, Lei Liang, Lei Xu, Li Tang, Liang Jiang, Liangcheng Fu, Lihui Zhang, Linfeng Shi, Lintao Ma, Liyuan Liu, Longfei Li, Longfei Zheng, Lu Liu, Lu Yu, Man Li, Meiqi Zhu, Meng Li, Mengjie Gao, Mengshu Sun, Mingming Yin, Mingyang Zhang, Mingyuan Fan, Nuo Xu, Pan Tang, Peijie Jiang, Peilong Zhao, Peng Lin, Pingping Liu, Qi Zuo, Qian Zhao, Qiang Cheng, Qianggang Cao, Qiaoben Bao, Qing Cui, Qingyuan Yang, Qitao Shi, Qiyin Huang, Qizheng Zhou, Quan Wan, Runyuan Zhao, Shaomian Zheng, Shaowei Wei, Shengnan Zhang, Shuaicheng Li, Shujie Li, Shuo Zhang, Sikang Bian, Tianchu Yao, Tiange Xu, Tianshu Wang, Ting Guo, Tinghao Wang, Tingwei Huang, Tong Zhao, Tongkai Yang, Wang Hong, Wanli Gu, Wei Lu, Weichang Wu, Weiguang Han, Weiquan Li, Wenbo Shen, Wenjing Fang, Wenzhi Tang, Xiang Shu, Xiao Shi, Xiaodong Yan, Xiaolu Zhang, Xiaopei Wan, Xiaqing Sun, Xin Zhao, Xingyu Lu, Xinxing Yang, Xinyao Tang, Xinyu Kong, Xinyu Liu, Xiong Xu, Xuan Sun, Xudong Han, Xudong Wang, Xujie Shen, Yalin Zhang, Yangyang Hou, Yankun Ren, Yao Zhao, Ye Chen, Yeyang Chen, Yibo Cao, Yifan Zuo, Yijie Chen, Ying Li, Yingjie Song, Yingxue Li, Yiqi Wang, Yixuan Sun, Yizhu Xiao, Yongfei Xu, Yu Liu, Yuchen Fang, Yue Gao, Yue Yu, Yue Zhang, Yuqi Zhang, Yuxiao He, Yuxiao Lu, Yuxin Tian, Yuxuan Li, Yuzhuo Fu, Zhankai Xu, Zhaoxin Huan, Zhenduo Zhang, Zhengke Gui, Zhengyu Huang, Zhenjun Ma, Zhenxuan Pan, Zheping Qu, Zhibo Zhu, Zhidong Fan, Zhigang Huangfu, Zhihao Wang, Zhiqiang Zhang, Zhizhen Liu, Zhuyan Zhou, Zibin Lin, Zihang Zeng, Zihao Wang, Zilong Wang, Ziqi Liu, Zitao Xuan, Zixuan Cheng, Zujie Wen, Zuoli Tang, 2026
http://arxiv.org/abs/2606.15079
2. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models — Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, Denny Zhou, 2022
https://scholar.google.com/scholar?q=Chain-of-Thought+Prompting+Elicits+Reasoning+in+Large+Language+Models
3. Let's Verify Step by Step — Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, Karl Cobbe, 2023
https://scholar.google.com/scholar?q=Let%27s+Verify+Step+by+Step
4. CoT-Evo: Evolutionary Distillation of Chain-of-Thought for Scientific Reasoning — Kehua Feng, Keyan Ding, Zhihui Zhu, Lei Liang, Qiang Zhang, Huajun Chen, 2025
https://scholar.google.com/scholar?q=CoT-Evo%3A+Evolutionary+Distillation+of+Chain-of-Thought+for+Scientific+Reasoning
5. Chain Of Thought Compression: A Theoretical Analysis — Juncai Li, Ru Li, Yuxiang Zhou, Boxiang Ma, Jeff Z. Pan, 2026
https://scholar.google.com/scholar?q=Chain+Of+Thought+Compression%3A+A+Theoretical+Analysis
6. Every Attention Matters: An Efficient Hybrid Architecture for Long-Context Reasoning (https://arxiv.org/abs/2510.19338) — Ling Team et al., 2025
https://scholar.google.com/scholar?q=Every+Attention+Matters%3A+An+Efficient+Hybrid+Architecture+for+Long-Context+Reasoning+%28https%3A%2F%2Farxiv.org%2Fabs%2F2510.19338%29
7. Various Lengths, Constant Speed: Efficient Language Modeling with Lightning Attention (https://arxiv.org/abs/2405.17381) — Zhen Qin et al., 2024
https://scholar.google.com/scholar?q=Various+Lengths%2C+Constant+Speed%3A+Efficient+Language+Modeling+with+Lightning+Attention+%28https%3A%2F%2Farxiv.org%2Fabs%2F2405.17381%29
8. DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model (https://arxiv.org/abs/2405.04434) — DeepSeek-AI et al., 2024
https://scholar.google.com/scholar?q=DeepSeek-V2%3A+A+Strong%2C+Economical%2C+and+Efficient+Mixture-of-Experts+Language+Model+%28https%3A%2F%2Farxiv.org%2Fabs%2F2405.04434%29
9. Holistic Capability Preservation: Towards Compact Yet Comprehensive Reasoning Models (https://arxiv.org/abs/2504.07158) — Ling Team et al., 2025
https://scholar.google.com/scholar?q=Holistic+Capability+Preservation%3A+Towards+Compact+Yet+Comprehensive+Reasoning+Models+%28https%3A%2F%2Farxiv.org%2Fabs%2F2504.07158%29
10. Every Step Evolves: Scaling Reinforcement Learning for Trillion-Scale Thinking Model (https://arxiv.org/abs/2510.18855) — Ling Team et al., 2025
https://scholar.google.com/scholar?q=Every+Step+Evolves%3A+Scaling+Reinforcement+Learning+for+Trillion-Scale+Thinking+Model+%28https%3A%2F%2Farxiv.org%2Fabs%2F2510.18855%29
11. Michelangelo: Long Context Evaluations Beyond Haystacks via Latent Structure Queries (https://arxiv.org/abs/2409.12640) — Kiran Vodrahalli et al., 2024
https://scholar.google.com/scholar?q=Michelangelo%3A+Long+Context+Evaluations+Beyond+Haystacks+via+Latent+Structure+Queries+%28https%3A%2F%2Farxiv.org%2Fabs%2F2409.12640%29
12. HyperAttention: Long-context Attention in Near-Linear Time — Insu Han et al., 2023
https://arxiv.org/abs/2310.05869
13. MiniCPM-SALA: Hybridizing Sparse and Linear Attention for Efficient Long-Context Modeling — MiniCPM Team / Wenhao An et al., 2026
https://arxiv.org/abs/2602.11761
14. Towards Thinking-Optimal Scaling of Test-Time Compute for LLM Reasoning — Wenkai Yang et al., 2025
https://arxiv.org/abs/2502.18080
15. Reasoning on a Budget: A Survey of Adaptive and Controllable Test-Time Compute in LLMs — Mohammad Ali Alomrani et al., 2025
https://arxiv.org/abs/2507.02076
16. Anyprefer: An Agentic Framework for Preference Data Synthesis — Yiyang Zhou et al., 2025
https://arxiv.org/abs/2504.19276
17. Towards Comprehensive Preference Data Collection for Reward Modeling — Yulan Hu et al., 2024
https://arxiv.org/abs/2406.16486
18. AI Post Transformers: Affordable Large-Scale Decoding Through Model-System Co-Design — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-19-affordable-large-scale-decoding-through-e1d7ed.mp3
19. AI Post Transformers: DeepSeek-V4 and Practical Million-Token Context — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-25-deepseek-v4-and-practical-million-token-6f4de1.mp3
20. AI Post Transformers: Speculative Decoding in Real vLLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-speculative-decoding-in-real-vllm-servin-6f4e2b.mp3
21. AI Post Transformers: Self-Improving Pretraining With Post-Trained Models — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-02-self-improving-pretraining-with-post-tra-e37460.mp3
22. AI Post Transformers: Agentic Discovery for Test-Time Scaling — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-12-agentic-discovery-for-test-time-scaling-f9a81f.mp3
23. AI Post Transformers: Nemotron 3 Super Hybrid Mamba-Transformer MoE — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-19-nemotron-3-super-hybrid-mamba-transforme-31ac75.mp3
24. AI Post Transformers: Kimi K2.5 and Visual Agent Swarms — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-24-kimi-k25-and-visual-agent-swarms-7d04d7.mp3
Interactive Visualization: Ling and Ring 2.6 for Trillion-Scale Agents

This episode explores SageAttention2, an ICML 2025 paper on making exact transformer attention faster without changing the underlying computation, focusing on why long-context models still pay a steep quadratic cost and why exact kernels remain important despite sparse and linear alternatives. It explains the paper’s central claim that aggressive low-precision attention can work only with careful numerical repair: queries and keys are pushed to INT4, attention-weight and value computation moves toward FP8, and outlier-smoothing ideas inspired by SmoothQuant are used to keep softmax-sensitive logits from collapsing. The discussion highlights the paper’s most concrete systems contribution, per-thread INT4 quantization aligned to GPU thread fragments and PTX `mma` execution, which aims to get fine-grained scaling without losing the performance win to dequantization overhead. A listener would find it interesting because the episode turns a seemingly narrow kernel optimization into a broader argument about hardware-software co-design, showing how much engineering is required to make lower-bit attention practical rather than just theoretically faster.

Sources:
1. SageAttention2: Efficient Attention with Thorough Outlier Smoothing and Per-thread INT4 Quantization — Jintao Zhang, Haofeng Huang, Pengle Zhang, Jia Wei, Jun Zhu, Jianfei Chen, 2024
http://arxiv.org/abs/2411.10958
2. FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning — Tri Dao, 2023
https://scholar.google.com/scholar?q=FlashAttention-2%3A+Faster+Attention+with+Better+Parallelism+and+Work+Partitioning
3. INT-FlashAttention: Enabling Flash Attention for INT8 Quantization — Shimao Chen, Zirui Liu, Zhiying Wu, et al., 2024
https://scholar.google.com/scholar?q=INT-FlashAttention%3A+Enabling+Flash+Attention+for+INT8+Quantization
4. SageAttention: Accurate 8-Bit Attention for Plug-and-play Inference Acceleration — Jintao Zhang, Jia Wei, Haofeng Huang, Pengle Zhang, Jun Zhu, Jianfei Chen, 2025
https://scholar.google.com/scholar?q=SageAttention%3A+Accurate+8-Bit+Attention+for+Plug-and-play+Inference+Acceleration
5. SageAttention2: Efficient Attention with Thorough Outlier Smoothing and Per-thread INT4 Quantization — Jintao Zhang, Haofeng Huang, Pengle Zhang, Jia Wei, Jun Zhu, Jianfei Chen, 2025
https://scholar.google.com/scholar?q=SageAttention2%3A+Efficient+Attention+with+Thorough+Outlier+Smoothing+and+Per-thread+INT4+Quantization
6. Understanding and Overcoming the Challenges of Efficient Transformer Quantization — Yelysei Bondarenko, Markus Nagel, Tijmen Blankevoort, 2021
https://scholar.google.com/scholar?q=Understanding+and+Overcoming+the+Challenges+of+Efficient+Transformer+Quantization
7. SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models — Guangxuan Xiao, Ji Lin, Mickael Seznec, Hao Wu, Julien Demouth, Song Han, 2023
https://scholar.google.com/scholar?q=SmoothQuant%3A+Accurate+and+Efficient+Post-Training+Quantization+for+Large+Language+Models
8. Outlier Suppression+: Accurate quantization of large language models by equivalent and optimal shifting and scaling — Xiuying Wei, Yunchen Zhang, Yuhang Li, Xiangguo Zhang, Ruihao Gong, Jinyang Guo, Xianglong Liu, 2023
https://scholar.google.com/scholar?q=Outlier+Suppression%2B%3A+Accurate+quantization+of+large+language+models+by+equivalent+and+optimal+shifting+and+scaling
9. QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs — Saleh Ashkboos, Amirkeivan Mohtashami, Maximilian L. Croci, et al., 2024
https://scholar.google.com/scholar?q=QuaRot%3A+Outlier-Free+4-Bit+Inference+in+Rotated+LLMs
10. FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision — Jay Shah, Ganesh Bikshandi, Ying Zhang, Vijay Thakkar, Pradeep Ramani, Tri Dao, 2024
https://scholar.google.com/scholar?q=FlashAttention-3%3A+Fast+and+Accurate+Attention+with+Asynchrony+and+Low-precision
11. QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving — Yujun Lin, Haotian Tang, Shang Yang, Zhekai Zhang, Guangxuan Xiao, Chuang Gan, Song Han, 2024
https://scholar.google.com/scholar?q=QServe%3A+W4A8KV4+Quantization+and+System+Co-design+for+Efficient+LLM+Serving
12. MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention — Huiqiang Jiang et al., 2024
https://arxiv.org/abs/2407.02490
13. SampleAttention: Near-Lossless Acceleration of Long Context LLM Inference with Adaptive Structured Sparse Attention — Qianchao Zhu et al., 2024
https://arxiv.org/abs/2406.15486
14. FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference — Xunhao Lai et al., 2025
https://arxiv.org/abs/2502.20766
15. Activation Outliers in Transformer Quantization: Reproduction, Statistical Analysis, and Deployment Tradeoffs — Pranav Kumar Kaliaperumal, 2026
https://arxiv.org/abs/2603.04308
16. BAPS: A Fine-Grained Low-Precision Scheme for Softmax in Attention via Block-Aware Precision reScaling — Zisheng Ye et al., 2026
https://arxiv.org/abs/2602.02071
17. Softpick: No Attention Sink, No Massive Activations with Rectified Softmax — Zayd M. K. Zuhri et al., 2025
https://arxiv.org/abs/2504.20966
18. AI Post Transformers: Deep Kernel Fusion for Transformer Decoding — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-deep-kernel-fusion-for-transformer-decod-b1a703.mp3
19. AI Post Transformers: TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-25-turboquant-online-vector-quantiz-1967b7.mp3
20. AI Post Transformers: NanoFlow and the Future of LLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-nanoflow-and-the-future-of-llm-serving-7429c9.mp3
Interactive Visualization: SageAttention2 and Fast Exact INT4 Attention

This episode explores NVIDIA’s Nemotron 3 Ultra, an open 550-billion-parameter mixture-of-experts model with only 55 billion parameters active per token, a hybrid Mamba-transformer backbone, and a 1 million token context window aimed at long-running agentic reasoning. It explains how sparse MoE routing, Mamba-style sequence layers, and low-precision NVFP4 training are used to cut KV-cache pressure, memory bandwidth costs, and decode-time latency for workloads like extended coding, tool use, and document-heavy planning. The discussion also breaks down the model’s full training and post-training stack, including LatentMoE, multi-token prediction, RL for reasoning and tool use, specialist-teacher distillation, and user-facing reasoning budget control. Listeners would find it interesting because the episode goes beyond benchmark headlines to examine the real argument of the paper: that long-horizon AI agents depend as much on serving economics and systems design as on raw model intelligence.

Sources:
1. Nemotron 3 Ultra: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning — NVIDIA, :, Aaron Blakeman, Aaron Thomas, Aastha Jhunjhunwala, Abhibha Gupta, Abhinav Khattar, Adam Rajfer, Adi Renduchintala, Adil Asif, Aditya Vavre, Adriana Flores Miranda, Ahmad Bilal, Aileen Zaman, Ajay Hotchandani, Akanksha Shukla, Akhiad Bercovich, Aleksander Ficek, Alex Gronskiy, Alex Kondratenko, Alex Steiner, Alex Ye, Alexander Bukharin, Alexandre Milesi, Ali Taghibakhshi, Alice Gatti, Alisa Liu, Alok Kumar, Amar Phanishayee, Ameya Sunil Mahabaleshwarkar, Amir Klein, Amit Zuker, Amnon Geifman, Anahita Bhiwandiwalla, Ananth Subramaniam, Andrea Santilli, Andrew Fulks, Andrew McHarg, Andrew Tao, Andrii Skliar, Anjulie Agrusa, Ankur Srivastava, Ankur Verma, Anna Shors, Anna Warno, Antoni-Joan Solergibert I Llaquet, Arham Mehta, Arkadiusz Nowaczynski, Arti Jain, Ashwath Aithal, Ashwin Poojary, Asif Ahamed, Asit Mishra, Asma Kuriparambil Thekkumpate, Atefeh Sohrabizadeh, Avinash Kaur, Avinash Vem, Ayush Dattagupta, Barath Subramaniam Anandan, Bardiya Sadeghi, Ben Lanir, Benedikt Schifferer, Besmira Nushi, Bilal Kartal, Bill Thiede, Bita Darvish Rouhani, Bo Deng, Bob Schatz, Boris Ginsburg, Boxin Wang, Brad Nemire, Brandon Norick, Brian Dang, Brian Westphal, Brian Yu, Brucek Khailany, Bryan Catanzaro, Carlo del Mundo, Caryln Aarish, Chankyu Lee, Chantal Hwang, Charbel Sakr, Charles Wang, Charlie Truong, Chen Cui, Cheng Cheng, Cheng-Ping Hsieh, Chenghao Zhang, Chenhui Deng, Chintan Patel, Chris Alexiuk, Christian Cosgrove, Christian Munley, Christine Harvey, Christopher Parisien, Chunyang Shen, Coco Li, Collin Neale, Cynthia Gao, Cyril Meurillon, Dan Gil, Dan Su, Dan Zhao, Dane Corneil, Daniel Afrimi, Daniel Egert, Daniel Korzekwa, Daniel Lo, Daniel Machlab, Daniel Serebrenik, Daniil Sorokin, Daria Gitman, Daria Levy, Darko Stosic, David Mosallanezhad, David Yu, Davit Karamyan, Deena Donia, Deep Debroy, Deepak Narayanan, Devin O'Kelly, Dheeraj Peri, Dhruv Nathawani, Di, Wu, Dima Rekesh, Divyanshu Kakwani, Donald Plummer, Dong Anh, Dongfeng Yu, Dongfu Jiang, Donnie Kim, Dorrin Poorkay, Duncan Riach, Dusan Stosic, Dustin VanStee, Eavan Meng, Edgar Minasyan, Edward Lin, Eileen Margaret Peters Long, Elad Sarafin, Elad Segal, Elena Lantz, Ellie Evans, Elliott Ning, Eric Chung, Eric Harper, Eric Pham-Hung, Eric Tramel, Eric Yang, Erick Galinkin, Erik Pounds, Erika Goncalves Goncalves, Evan Briones, Evan Wu, Evelina Bakhturina, Evgeny Tsykunov, Ewa Dobrowolska, Faisal Ladhak, Farzan Memarian, Fay Wang, Fei Jia, Felipe Soares, Felipe Vieira Frujeri, Feng Chen, Fengguang Lin, Ferenc Galko, Frank Sun, Frankie Siino, Frida Hou, Gal Hubara Agam, Gal Kaplun, Gantavya Bhatt, Gargi Prasad, Garvit Kulshreshtha, George Armstrong, Gerald Shen, Giulio Borghesi, Gordana Neskovic, Gorkem Batmaz, Grace Lam, Greg Mason, Greg Pauloski, Grigor Nalbandyan, Grzegorz Chlebus, Grzegorz Karch, Guan-Ting Liu, Guoming Zhang, Guyue Huang, Haggai Maron, Haifeng Qian, Haim Elisha, Haoxing Ren, Haran Kumar Shiv Kumar, Haribhau Hud, Harris Nover, Harrison Saturley Hall, Hayate Iso, Helen Ngo, Herbert Hum, Herman Sahota, Hexin Wang, Himanshu Soni, Hovhannes Tamoyan, Hua Li, Huanhuan Chen, Hui Li, Hui Wang, Huy Nguyen, Ian Chiles, Ido Galil, Ido Shahaf, Igor Gitman, Igor Shovkun, Ilya Loshchilov, Ingo Guehring, Itamar Schen, Itay Levy, Itay Neeman, Ivan Moshkov, Izik Golan, Izzy Putterman, Jaemin Choi, Jakub Slowikowski, Jan Kautz, Jane Polak Scowcroft, Jared Casper, Jatin Mitra, Jeffrey Glick, Jenny Chen, Jesse Oliver, Jiacheng Xu, Jiafan Zhu, Jialin Song, Jian Zhang, Jiantao Jiao, Jiaqi Zeng, Jie Lou, Jim King, Jimmy Zhang, Jingquan Wang, Jinhang Choi, Jinju Chu, Joey Conway, Joey Guman, Johan Jatko, Johannes Rausch, John Kamalu, John Roberts, Johnny Greco, Johnny Mensel, Jonah Alben, Jonas Yang, Jonathan Cohen, Jonathan Raiman, Joseph Jennings, Joshua Mabry, Joshua Pierce, Joyjit Daw, Julien Veron Vialard, Junkeun Yi, Jupinder Parmar, Kajal Jain, Kan Zhu, Kari Briski, Katherine Cheung, Katherine Luna, Keith Willowhawk, Keith Wyss, Keshav Santhanam, Kevin Shih, Kezhi Kong, Khanh Nguyen, Khushi Bhardwaj, Kirthi Shankar Sivamani, Konstantinos Krommydas, Krishna C. Puvvada, Krzysztof Pawelec, Kumar Anik, Kyle Keprios, Kylie Day, Lawrence McAfee, Leo Du, Leon Derczynski, Li Ding, Linda Liu, Lingjie Wu, Lior Kadoch, Lizzie Wei, Luis Vega, Luke Robison, Lun Su, Maarten Van Segbroeck, Maciej Jakub Mikulski, Maer Rodrigues de Melo, Magda Sypula, Mahan Fathi, Makesh Narsimhan Sreedhar, Makesh Tarun Chandran, Manoj Kilaru, Maor Ashkenazi, Marc Cuevas, Marc Romeijn, Marcin Chochowski, Mark Cai, Mark Mozolewski, Markus Kliegl, Marta Stepniewska-Dziubinska, Martyna Patelka, Mattei Machczynski, Matvei Novikov, Mauricio Ferrato, Maximilian Golub, Mehrzad Samadi, Melissa Corpuz, Mengru Wang, Mengxi Wu, Meredith Price, Meriem Boubdir, Micah Schaffer, Michael Andersch, Michael Boone, Michael Gschwind, Michael Lightstone, Michael Loh, Michal Bien, Michal Zawalski, Michelle Gill, Miguel Martinez, Mikail Khona, Mike Chrzanowski, Mike Houston, Mingyuan Ma, Minseok Lee, Mohamed Fawzy, Mohammad Dabbah, Mohammad Shoeybi, Mostofa Patwary, Nabin Mulepati, Najeeb Nabwani, Namit Dhameja, Narimane Hennouni, Natalie Hereth, Nathaniel Pinckney, Nave Algarici, Nave Assaf, Netanel Haber, Nicholas Knight, Nick Reamaroon, Nickson Quak, Nidhi Bhatia, Nikhil Desai, Nikolai Ludwig, Nima Tajbakhsh, Ning Xu, Nir Ailon, Nirmal Juluru, Nitin Nitin, Ofri Masad, Oleg Rybakov, Oleksii Hrinchuk, Oleksii Kuchaiev, Olivia Viessmann, Olivier Delalleau, Oluwatobi Olabiyi, Omer Ullman Argov, Omri Puny, Oren Tropp, Pablo Ribalta, Pallab Bhattacharya, Panos Lampropoulos, Parth Mannan, Pasha Shamis, Patrick Legresley, Paul Gibbons, Pavlo Molchanov, Pawel Morkisz, Peter Dykas, Peter Jin, Pierre-Yves Aquilanti, Pinky Xu, Piotr Januszewski, Piotr Laskiewicz, Pooya Jannaty, Prakash Gurumurthy, Pranav Prashant Thombre, Prasoon Varshney, Pritam Gundecha, Przemek Tredak, Puhui Meng, Qiyu Wan, Rabeeh Karimi Mahabadi, Rachel Oberman, Rachit Garg, Radha Sri-Tharan, Rahul Kandu, Rakshit Sanadhya, Ran El-Yaniv, Ran Zilberstein, Rasoul Shafipour, Ray Macalisang, Rayen Tian, Reka Kovacs, Renjie Pi, Rick Izzo, Rima Shahbazyan, Rishabh Garg, Rishi Puri, Rita Fernandes Neves, Ritchie Zhao, Ritika Borkar, Ritu Gala, Riyad Islam, Robert Clark, Robert Hesse, Robert Kirby, Roger Waleffe, Rohit Watve, Roi Koren, Ron Banner, Ruoxi Zhang, Russell J. Hewett, Ryan Prenger, Ryan Stewart, Ryota Egashira, Sadegh Mahdavi, Saee Paliwal, Sagar Singh, Sahil Modi, Salika Dave, Samantha Shinagawa, Samuel Kriman, Sandip Bhaskar, Sangkug Lym, Sanjay Kariyappa, Sanjeev Satheesh, Saran Vikas Murari, Satish Pasumarthi, Saurabh Mishra, Saurav Muralidharan, Scott Hara, Sean Narentharen, Selvaraj Anandaraj, Seonjin Na, Seonmeyong Bak, Seonmyeong Bak, Sepehr Sameni, Seph Mard, Serge Panev, Seth Henneman, Seth Poulos, Shahar Mor, Shantanu Acharya, Shaona Ghosh, Sharath Turuvekere Sreenivas, Sharon Mendelson, Shaun Kotek, Shawn Wang, Shay Aharon, Shaya Gharghabi, Sheng-Chieh Lin, Shi Chen, Shiqing Fan, Shirish Baskaran, Shreya Gopa, Shrimai Prabhumoye, Shubham Pachori, Shubham Toshniwal, Shuoyang Ding, Shwetha Krishnamurthy, Siddharth Singh, Simeng Sun, Sirshak Das, Sivakumar Arayandi Thottakara, Smita Ithape, Somshubra Majumdar, Soumye Singhal, Sri Harsha Singudasu, Sridhar Bhuvanapalli, Srimukh Veccham, Stas Sergienko, Stefania Alborghetti, Stephen Ge, Su Rong, Sugam Dipak Devare, Sukrit Rao, Sumeet Kumar Barua, Sungsoo Ha, Sunny Gai, Suriya Gunasekar, Suseella Panguluri, Suyog Gupta, Sviataslau Hinzburh, Sweta Priyadarshi, Syeda Nahida Akter, Talor Abramovich, Tan Bui, Tanay Varshney, Tatevik Ter-Hovhannisyan, Teodor-Dumitru Ene, Terry Kong, Thanh Do, Tianhe Zhang, Tiffany Moore, Tijmen Blankevoort, Tim Moon, Tiyasa Mitra, Tom Balough, Tomasz Grzegorzek, Tomasz Hliwiak, Tomer Asida, Tomer Bar Natan, Tomer Keren, Tomer Ronen, Tony Salim, Tony Wang, Traian Rebedea, Tugrul Konuk, Twinkle Vashishth, Udi Karpas, Ushnish De, Vahid Noorozi, Venkat Srinivasan, Venmugil Elango, Vibhor Agrawal, Victor Cui, Vijay Korthikanti, Vikas Mehta, Vinay Rao, Virginia Wu, Vitaly Kurin, Vitaly Lavrukhin, Vladimir Anisimov, Vu Pham, Wanli Jiang, Wasi Uddin Ahmad, Wataru Ishihara, Wei Du, Wei Ping, Weiheng Chai, Wenliang Dai, Wesley Helmholz, Will Jennings, Will Zhu, Wojciech Prazuch, Xiaowei Ren, Xiwen Yu, Yan Breek, Yang Chen, Yang Yu, Yangyi Chen, Yaniv Galron, Yashaswi Karnati, Yejin Choi, Yev Meyer, Yi-Fu Wu, Yian Zhang, Ying Lin, Yonatan Geifman, Yonggan Fu, Youngeun Kwon, Yu Yao, Yugi Guvvla, Yuki Huang, Yunsheng Liu, Zach Moshe, Zachary Newell, Zhilin Wang, Zhiyu Li, Zhongbo Zhu, Zhuolin Yang, Zihan Liu, Zijie Yan, Zsolt-Alon Wertheimer, 2026
http://arxiv.org/abs/2606.15007
2. Distilling the Knowledge in a Neural Network — Geoffrey Hinton, Oriol Vinyals, Jeff Dean, 2015
https://arxiv.org/abs/1503.02531
3. On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes — Rishabh Agarwal, Nino Vieillard, Yongchao Zhou, Piotr Stanczyk, Sabela Ramos, Matthieu Geist, Olivier Bachem, 2023
https://arxiv.org/abs/2306.13649
4. One Teacher is Enough? Pre-trained Language Model Distillation from Multiple Teachers — Chuhan Wu, Fangzhao Wu, Yongfeng Huang, 2021
https://arxiv.org/abs/2106.01023
5. Multi-teacher Distillation for Multilingual Spelling Correction — Jingfen Zhang, Xuan Guo, Sravan Bodapati, Christopher Potts, 2023
https://arxiv.org/abs/2311.11518
6. LatentMoE: Toward Optimal Accuracy per FLOP and Parameter in Mixture of Experts — Venmugil Elango et al., 2026
https://scholar.google.com/scholar?q=LatentMoE%3A+Toward+Optimal+Accuracy+per+FLOP+and+Parameter+in+Mixture+of+Experts
7. Better & Faster Large Language Models via Multi-token Prediction — Fabian Gloeckle, Badr Youbi Idrissi, Baptiste Roziere, David Lopez-Paz, Gabriel Synnaeve, 2024
https://scholar.google.com/scholar?q=Better+%26+Faster+Large+Language+Models+via+Multi-token+Prediction
8. Jamba: A Hybrid Transformer-Mamba Language Model — Opher Lieber et al., 2024
https://scholar.google.com/scholar?q=Jamba%3A+A+Hybrid+Transformer-Mamba+Language+Model
9. On-Policy Context Distillation for Language Models — Tianzhu Ye, Li Dong, Xun Wu, Shaohan Huang, Furu Wei, 2026
https://scholar.google.com/scholar?q=On-Policy+Context+Distillation+for+Language+Models
10. Characterizing State Space Model and Hybrid Language Model Performance with Long Context — Saptarshi Mitra et al., 2025
https://scholar.google.com/scholar?q=Characterizing+State+Space+Model+and+Hybrid+Language+Model+Performance+with+Long+Context
11. Compute Or Load KV Cache? Why Not Both? — Shuowei Jin et al., 2024
https://scholar.google.com/scholar?q=Compute+Or+Load+KV+Cache%3F+Why+Not+Both%3F
12. IceCache: Memory-efficient KV-cache Management for Long-Sequence LLMs — Yuzhen Mao et al., 2026
https://scholar.google.com/scholar?q=IceCache%3A+Memory-efficient+KV-cache+Management+for+Long-Sequence+LLMs
13. DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models — Damai Dai et al., 2024
https://scholar.google.com/scholar?q=DeepSeekMoE%3A+Towards+Ultimate+Expert+Specialization+in+Mixture-of-Experts+Language+Models
14. Scaling Speculative Decoding with Lookahead Reasoning — Yichao Fu et al., 2025
https://scholar.google.com/scholar?q=Scaling+Speculative+Decoding+with+Lookahead+Reasoning
15. AI Post Transformers: Nemotron 3 Super Hybrid Mamba-Transformer MoE — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-19-nemotron-3-super-hybrid-mamba-transforme-31ac75.mp3
16. AI Post Transformers: DeepSeek-V4 and Practical Million-Token Context — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-25-deepseek-v4-and-practical-million-token-6f4de1.mp3
17. AI Post Transformers: Affordable Large-Scale Decoding Through Model-System Co-Design — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-19-affordable-large-scale-decoding-through-e1d7ed.mp3
18. AI Post Transformers: Mooncake for KV Cache-Centric LLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-05-mooncake-for-kv-cache-centric-llm-servin-1086d0.mp3
19. AI Post Transformers: Memory-Bound, Not Bandwidth-Limited Batch-1 LLM Decode — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-02-memory-bound-not-bandwidth-limited-batch-114799.mp3
20. AI Post Transformers: AgenticQwen and Small Industrial Tool Agents — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-27-agenticqwen-and-small-industrial-tool-ag-dc676d.mp3

This episode explores OpenSkill, a framework for LLM agents that tries to improve behavior after deployment by building durable, reusable skills from public evidence rather than retraining model weights. It explains how the paper separates ordinary tool use from open-world self-evolution, arguing that the key challenge is not just acting with browsers and code, but turning documentation, repositories, papers, and tutorials into explicit procedures and verification checks. The discussion focuses on the paper’s central claim that agents can create their own proxy tests through grounded verification anchors without leaking hidden benchmark answers, and compares that approach with earlier systems like Reflexion, Voyager, ExpeL, AutoSkill, and Memento-Skills. Listeners would find it interesting because it gets at a practical industry problem: whether agents can stay useful as APIs, websites, and workflows change, or whether the verifier remains the real bottleneck to genuine self-improvement.

Sources:
1. OpenSkill: Open-World Self-Evolution for LLM Agents — Zhiling Yan, Dingjie Song, Hanrong Zhang, Wei Liang, Yuxuan Zhang, Yutong Dai, Lifang He, Philip S. Yu, Ran Xu, Xiang Li, Lichao Sun, 2026
http://arxiv.org/abs/2606.06741
2. Reflexion: Language Agents with Verbal Reinforcement Learning — Noah Shinn, Federico Cassano, Edward Berman, Ashwin Gopinath, Karthik Narasimhan, Shunyu Yao, 2023
https://arxiv.org/abs/2303.11366
3. Voyager: An Open-Ended Embodied Agent with Large Language Models — Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, Anima Anandkumar, 2023
https://arxiv.org/abs/2305.16291
4. ExpeL: LLM Agents Are Experiential Learners — Andrew Zhao, Daniel Huang, Quentin Xu, Matthieu Lin, Yong-Jin Liu, Gao Huang, 2023
https://arxiv.org/abs/2308.10144
5. Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration — Qifan Zhang, Dongyang Ma, Tianqing Fang, Jia Li, Jing Tang, Nuo Chen, Haitao Mi, Yan Wang, 2026
https://arxiv.org/abs/2604.18131
6. Constitutional AI: Harmlessness from AI Feedback — Yuntao Bai, Saurav Kadavath, Amanda Askell, Ethan Perez, Jared Kaplan, Dario Amodei, Tom Brown and collaborators, 2022
https://arxiv.org/abs/2212.08073
7. RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback — Harrison Lee, Samrat Phatale, Hassan Mansoor, Thomas Mesnard, Johan Ferret, Kellie Lu, Colton Bishop, Ethan Hall, Victor Carbune, Abhinav Rastogi, Sushant Prakash, 2023
https://arxiv.org/abs/2309.00267
8. Self-Rewarding Language Models — Weizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Xian Li, Sainbayar Sukhbaatar, Jing Xu, Jason Weston, 2024
https://arxiv.org/abs/2401.10020
9. MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models — Longhui Yu, Weisen Jiang, Han Shi, Jincheng Yu, Zhengying Liu, Yu Zhang, James T. Kwok, Zhenguo Li, Adrian Weller, Weiyang Liu, 2023
https://arxiv.org/abs/2309.12284
10. SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks — Xiangyi Li et al., 2026
https://scholar.google.com/scholar?q=SkillsBench%3A+Benchmarking+How+Well+Agent+Skills+Work+Across+Diverse+Tasks
11. AutoSkill: Experience-Driven Lifelong Learning via Skill Self-Evolution — Yutao Yang et al., 2026
https://scholar.google.com/scholar?q=AutoSkill%3A+Experience-Driven+Lifelong+Learning+via+Skill+Self-Evolution
12. Memento-Skills: Let Agents Design Agents — Huichi Zhou et al., 2026
https://scholar.google.com/scholar?q=Memento-Skills%3A+Let+Agents+Design+Agents
13. SkillSafetyBench: Evaluating Agent Safety under Skill-Facing Attack Surfaces — Chang Jin et al., 2026
https://scholar.google.com/scholar?q=SkillSafetyBench%3A+Evaluating+Agent+Safety+under+Skill-Facing+Attack+Surfaces
14. EASYTOOL: Enhancing LLM-based Agents with Concise Tool Instruction — Siyu Yuan et al., 2024
https://scholar.google.com/scholar?q=EASYTOOL%3A+Enhancing+LLM-based+Agents+with+Concise+Tool+Instruction
15. AutoRefine: From Trajectories to Reusable Expertise for Continual LLM Agent Refinement — Libin Qiu et al., 2026
https://scholar.google.com/scholar?q=AutoRefine%3A+From+Trajectories+to+Reusable+Expertise+for+Continual+LLM+Agent+Refinement
16. SE-Agent: Self-Evolution Trajectory Optimization in Multi-Step Reasoning with LLM-Based Agents — Jiaye Lin et al., 2025
https://scholar.google.com/scholar?q=SE-Agent%3A+Self-Evolution+Trajectory+Optimization+in+Multi-Step+Reasoning+with+LLM-Based+Agents
17. When AIs Judge AIs: The Rise of Agent-as-a-Judge Evaluation for LLMs — Fangyi Yu, 2025
https://scholar.google.com/scholar?q=When+AIs+Judge+AIs%3A+The+Rise+of+Agent-as-a-Judge+Evaluation+for+LLMs
18. Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments — Yuran Li et al., 2025
https://scholar.google.com/scholar?q=Leveraging+LLMs+as+Meta-Judges%3A+A+Multi-Agent+Framework+for+Evaluating+LLM+Judgments
19. AI Post Transformers: The Endless Gym: Training Terminal Agents — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/the-endless-gym-training-terminal-agents/
20. AI Post Transformers: When AI Builds Itself and Recursive Self-Improvement — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-05-when-ai-builds-itself-and-recursive-self-8bbf9e.mp3
21. AI Post Transformers: Self-Improving Pretraining With Post-Trained Models — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-02-self-improving-pretraining-with-post-tra-e37460.mp3
22. AI Post Transformers: Split Personality Training Reveals Latent Knowledge — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-08-split-personality-training-reveals-laten-c84616.mp3
23. AI Post Transformers: In-Place Test-Time Training for Transformers — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-09-in-place-test-time-training-for-transfor-d0b976.mp3
Interactive Visualization: OpenSkill for Open-World Self-Evolution in LLM Agents

This episode explores PaperBench, a benchmark designed to test whether frontier AI agents can independently replicate the empirical work of recent machine learning papers from scratch rather than merely explain them. It breaks down what agentic AI actually entails in this setting: reading papers, writing code, choosing baselines, reconstructing missing details, running experiments, debugging failures, and judging whether reproduced results match the original claims. The discussion compares PaperBench with other evaluation ladders such as CORE-Bench, MLE-bench, RE-Bench, and JudgeEval, while also debating whether controlled scratch replication should be viewed as advanced engineering or a meaningful proxy for real research practice. Listeners get a clear look at why this matters for both AI capability measurement and safety, especially given PaperBench’s carefully curated design of 20 ICML 2024 papers, 12 topics, and more than 8,000 graded tasks.

Sources:
1. PaperBench: Evaluating AI's Ability to Replicate AI Research — Giulio Starace, Oliver Jaffe, Dane Sherburn, James Aung, Jun Shern Chan, Leon Maksin, Rachel Dias, Evan Mays, Benjamin Kinsella, Wyatt Thompson, Johannes Heidecke, Amelia Glaese, Tejal Patwardhan, 2025
http://arxiv.org/abs/2504.01848
2. PaperBench: Evaluating AI's Ability to Replicate AI Research — Giulio Starace, Oliver Jaffe, Dane Sherburn, James Aung, et al., 2025
https://scholar.google.com/scholar?q=PaperBench%3A+Evaluating+AI%27s+Ability+to+Replicate+AI+Research
3. RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts — Hjalmar Wijk, Tao Lin, Joel Becker, Sami Jawhar, et al., 2024
https://scholar.google.com/scholar?q=RE-Bench%3A+Evaluating+frontier+AI+R%26D+capabilities+of+language+model+agents+against+human+experts
4. CORE-Bench: Fostering the Credibility of Published Research Through a Computational Reproducibility Agent Benchmark — Zachary S. Siegel, Sayash Kapoor, Nitya Nagdir, Benedikt Stroebl, Arvind Narayanan, 2024
https://scholar.google.com/scholar?q=CORE-Bench%3A+Fostering+the+Credibility+of+Published+Research+Through+a+Computational+Reproducibility+Agent+Benchmark
5. MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation — Qian Huang, Jian Vora, Percy Liang, Jure Leskovec, 2023
https://scholar.google.com/scholar?q=MLAgentBench%3A+Evaluating+Language+Agents+on+Machine+Learning+Experimentation
6. MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering — Jun Shern Chan et al., 2024
https://scholar.google.com/scholar?q=MLE-bench%3A+Evaluating+Machine+Learning+Agents+on+Machine+Learning+Engineering
7. EXP-Bench: Can AI Conduct AI Research Experiments? — Patrick Tser Jern Kon et al., 2025
https://scholar.google.com/scholar?q=EXP-Bench%3A+Can+AI+Conduct+AI+Research+Experiments%3F
8. MLR-Bench: Evaluating AI Agents on Open-Ended Machine Learning Research — Hui Chen, Miao Xiong, Yujie Lu, Wei Han, Ailin Deng, Yufei He, Jiaying Wu, Yibo Li, Yue Liu, Bryan Hooi, 2025
https://scholar.google.com/scholar?q=MLR-Bench%3A+Evaluating+AI+Agents+on+Open-Ended+Machine+Learning+Research
9. ReplicationBench: Can AI Agents Replicate Astrophysics Research Papers? — Christine Ye et al., 2025
https://scholar.google.com/scholar?q=ReplicationBench%3A+Can+AI+Agents+Replicate+Astrophysics+Research+Papers%3F
10. Can Large Language Models Be an Alternative to Human Evaluations? — Cheng-Han Chiang, Hung-yi Lee, 2023
https://scholar.google.com/scholar?q=Can+Large+Language+Models+Be+an+Alternative+to+Human+Evaluations%3F
11. RubricEval: A Rubric-Level Meta-Evaluation Benchmark for LLM Judges in Instruction Following — Tianjun Pan et al., 2026
https://arxiv.org/abs/2603.25133
12. JudgeBench: A Benchmark for Evaluating LLM-based Judges — Sijun Tan et al., 2024
https://arxiv.org/abs/2410.12784
13. When Can We Trust LLMs in Mental Health? Large-Scale Benchmarks for Reliable LLM Evaluation — Abeer Badawi et al., 2025
https://arxiv.org/abs/2510.19032
14. A Dataset For Computational Reproducibility — Lazaro Costa, Susana Barbosa, Jacome Cunha, 2025
https://arxiv.org/abs/2504.08684
15. SciReplicate-Bench: Benchmarking LLMs in Agent-driven Algorithmic Reproduction from Research Papers — Yanzheng Xiang et al., 2025
https://arxiv.org/abs/2504.00255
16. OctoBench: Benchmarking Scaffold-Aware Instruction Following in Repository-Grounded Agentic Coding — Deming Ding et al., 2026
https://arxiv.org/abs/2601.10343
17. ContextBench: A Benchmark for Context Retrieval in Coding Agents — Han Li et al., 2026
https://arxiv.org/abs/2602.05892
18. AI Post Transformers: When AI Builds Itself and Recursive Self-Improvement — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-05-when-ai-builds-itself-and-recursive-self-8bbf9e.mp3
19. AI Post Transformers: When LLM Judges Become Coin Flips — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-when-llm-judges-become-coin-flips-8b43ef.mp3
20. AI Post Transformers: ASI-Evolve for Data, Architectures, and RL — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-05-asi-evolve-for-data-architectures-and-rl-197b2b.mp3
21. AI Post Transformers: Kimi K2.5 and Visual Agent Swarms — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-24-kimi-k25-and-visual-agent-swarms-7d04d7.mp3
Interactive Visualization: PaperBench: Can AI Replicate AI Research?

This episode explores the paper Cartridges at Scale, which asks whether large document collections can be distilled into reusable modular KV-cache memories so a model can answer questions without repeatedly rereading raw text. It explains what a cartridge is, how context distillation turns full-document context into compact learned prefixes, and why that differs from prompt caching, fine-tuning, ordinary long-context prompting, and text RAG. The discussion centers on the paper’s main claim that per-document memories do not reliably compose when trained independently, so the authors jointly train cartridges with both relevant and irrelevant memories present to teach a frozen model which compressed document to attend to in a noisy multi-document setting. Listeners would find it interesting because it treats the KV cache as a potential external memory layer that could reduce inference cost and latency while exposing hard questions about compositionality, transparency, and whether learned memory modules can outperform standard retrieval pipelines.

Sources:
1. Cartridges at Scale: Training Modular KV Caches over Large Document Collections — Momchil Hardalov, Gonzalo Iglesias, Adrià de Gispert, 2026
http://arxiv.org/abs/2606.04557
2. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks — Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Douwe Kiela, et al., 2020
https://arxiv.org/abs/2005.11401
3. Prompt Cache: Modular Attention Reuse for Low-Latency Inference — In Gim, Guojun Chen, Seung-seob Lee, Nikhil Sarda, Anurag Khandelwal, Lin Zhong, 2023
https://arxiv.org/abs/2311.04934
4. Cartridges: Lightweight and General-Purpose Long Context Representations via Self-Study — Sabri Eyuboglu, Ryan Ehrlich, Simran Arora, Neel Guha, James Zou, Azalia Mirhoseini, Christopher Re, et al., 2025
https://arxiv.org/abs/2506.06266
5. Cartridges at Scale: Training Modular KV Caches over Large Document Collections — Momchil Hardalov, Gonzalo Iglesias, Adrià de Gispert, 2026
https://arxiv.org/abs/2606.04557
6. Cartridges: Lightweight and General-Purpose Long Context Representations via Self-Study (https://arxiv.org/abs/2506.06266) — Sabri Eyuboglu, Ryan S. Ehrlich, Simran Arora, Neel Guha, Dylan Zinsley, Emily R. Liu, Atri Rudra, James Y. Zou, Azalia Mirhoseini, Christopher Re, 2025
https://scholar.google.com/scholar?q=Cartridges%3A+Lightweight+and+General-Purpose+Long+Context+Representations+via+Self-Study+%28https%3A%2F%2Farxiv.org%2Fabs%2F2506.06266%29
7. Learned Structure in CARTRIDGES: Keys as Shareable Routers in Self-Studied Representations (https://arxiv.org/abs/2508.17032) — Maurizio Diaz, 2025
https://scholar.google.com/scholar?q=Learned+Structure+in+CARTRIDGES%3A+Keys+as+Shareable+Routers+in+Self-Studied+Representations+%28https%3A%2F%2Farxiv.org%2Fabs%2F2508.17032%29
8. xRAG: Extreme Context Compression for Retrieval-Augmented Generation with One Token (https://arxiv.org/abs/2405.13792) — Xin Cheng, Xun Wang, Xingxing Zhang, Tao Ge, Si-Qing Chen, Furu Wei, Huishuai Zhang, Dongyan Zhao, 2024
https://scholar.google.com/scholar?q=xRAG%3A+Extreme+Context+Compression+for+Retrieval-Augmented+Generation+with+One+Token+%28https%3A%2F%2Farxiv.org%2Fabs%2F2405.13792%29
9. KVzip: Query-Agnostic KV Cache Compression with Context Reconstruction (https://arxiv.org/abs/2505.23416) — Jang-Hyun Kim, Jinuk Kim, Sangwoo Kwon, Jae W. Lee, Sangdoo Yun, Hyun Oh Song, 2025
https://scholar.google.com/scholar?q=KVzip%3A+Query-Agnostic+KV+Cache+Compression+with+Context+Reconstruction+%28https%3A%2F%2Farxiv.org%2Fabs%2F2505.23416%29
10. T2-RAGBench: Text-and-Table Benchmark for Evaluating Retrieval-Augmented Generation (https://aclanthology.org/2026.eacl-long.8/) — Jan Strich, Enes Kutay Isgorur, Maximilian Trescher, Chris Biemann, Martin Semmann, 2026
https://scholar.google.com/scholar?q=T2-RAGBench%3A+Text-and-Table+Benchmark+for+Evaluating+Retrieval-Augmented+Generation+%28https%3A%2F%2Faclanthology.org%2F2026.eacl-long.8%2F%29
11. KVLink: Accelerating Large Language Models via Efficient KV Cache Reuse — Jingbo Yang et al., 2025
https://scholar.google.com/scholar?q=KVLink%3A+Accelerating+Large+Language+Models+via+Efficient+KV+Cache+Reuse
12. Hierarchical Document Refinement for Long-context Retrieval-augmented Generation — Jiajie Jin et al., 2025
https://scholar.google.com/scholar?q=Hierarchical+Document+Refinement+for+Long-context+Retrieval-augmented+Generation
13. LaRA: Benchmarking Retrieval-Augmented Generation and Long-Context LLMs -- No Silver Bullet for LC or RAG Routing — Kuan Li et al., 2025
https://scholar.google.com/scholar?q=LaRA%3A+Benchmarking+Retrieval-Augmented+Generation+and+Long-Context+LLMs+--+No+Silver+Bullet+for+LC+or+RAG+Routing
14. ChatQA 2: Bridging the Gap to Proprietary LLMs in Long Context and RAG Capabilities — Peng Xu et al., 2024
https://scholar.google.com/scholar?q=ChatQA+2%3A+Bridging+the+Gap+to+Proprietary+LLMs+in+Long+Context+and+RAG+Capabilities
15. LegalBench-RAG: A Benchmark for Retrieval-Augmented Generation in the Legal Domain — Nicholas Pipitone and Ghita Houir Alami, 2024
https://scholar.google.com/scholar?q=LegalBench-RAG%3A+A+Benchmark+for+Retrieval-Augmented+Generation+in+the+Legal+Domain
16. KVShare: An LLM Service System with Efficient and Effective Multi-Tenant KV Cache Reuse — Huan Yang et al., 2025
https://scholar.google.com/scholar?q=KVShare%3A+An+LLM+Service+System+with+Efficient+and+Effective+Multi-Tenant+KV+Cache+Reuse
17. AI Post Transformers: KVzip for Query-Agnostic KV Cache Compression — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-29-kvzip-for-query-agnostic-kv-cache-compre-72afe5.mp3
18. AI Post Transformers: From Prefix Cache to Fusion RAG Cache: Accelerating LLM Inference in Retrieval-Augmented Generation — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-22-from-prefix-cache-to-fusion-rag-9c5d39.mp3
19. AI Post Transformers: Experimental Comparison of Agentic and Enhanced RAG — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-14-experimental-comparison-of-agentic-and-e-37d8bc.mp3
20. AI Post Transformers: Can Models Learn from Long Context? — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-can-models-learn-from-long-context-77533e.mp3
21. AI Post Transformers: δ-mem and Online Memory for LLMs — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-13-d-mem-and-online-memory-for-llms-6622fa.mp3
22. AI Post Transformers: FengHuang for Rack-Scale LLM Inference Memory — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-12-fenghuang-for-rack-scale-llm-inference-m-62708e.mp3

This episode explores DafnyPro, a system for using large language models to help verify Dafny programs while keeping the original executable logic unchanged. It explains the basics of formal verification in Dafny, including preconditions, postconditions, loop invariants, decreases clauses, ghost code, and why writing correct proof annotations is much harder than generating plausible code. The discussion compares DafnyPro with earlier efforts such as Clover, DafnyBench, and Laurel, then focuses on DafnyPro’s main contribution: a parser-backed safeguard that rejects any LLM attempt that alters program behavior and a verifier-guided loop that can also prune bad invariants instead of blindly adding more. A listener would find it interesting because it gets at a real trust problem in AI coding tools: whether a model can genuinely help prove software correct rather than quietly rewriting the task into something easier to verify.

Sources:
1. DafnyPro: LLM-Assisted Automated Verification for Dafny Programs — Debangshu Banerjee, Olivier Bouissou, Stefan Zetzsche, 2026
http://arxiv.org/abs/2601.05385
2. Clover: Closed-Loop Verifiable Code Generation — Chuyue Sun, Ying Sheng, Oded Padon, Clark Barrett, 2023
https://arxiv.org/abs/2310.17807
3. DafnyBench: A Benchmark for Formal Software Verification — Chloe Loughridge, Qinyi Sun, Seth Ahrenbach, Federico Cassano, Chuyue Sun, Ying Sheng, Anish Mudide, Md Rakib Hossain Misu, Nada Amin, Max Tegmark, 2024
https://arxiv.org/abs/2406.08467
4. Laurel: Unblocking Automated Verification with Large Language Models — Eric Mugnier, Emmanuel Anaya Gonzalez, Ranjit Jhala, Nadia Polikarpova, Yuanyuan Zhou, 2024 (rev. 2025)
https://arxiv.org/abs/2405.16792
5. DafnyPro: LLM-Assisted Automated Verification for Dafny Programs — Debangshu Banerjee, Olivier Bouissou, Stefan Zetzsche, 2026
https://arxiv.org/abs/2601.05385
6. Laurel: Generating Dafny Assertions Using Large Language Models — Eric Mugnier, Emmanuel Anaya Gonzalez, Ranjit Jhala, Nadia Polikarpova, Yuanyuan Zhou, 2024
https://arxiv.org/abs/2405.16792
7. dafny-annotator: AI-Assisted Verification of Dafny Programs — Gabriel Poesia, Chloe Loughridge, Nada Amin, 2024
https://arxiv.org/abs/2411.15143
8. Towards AI-Assisted Synthesis of Verified Dafny Methods — Md Rakib Hossain Misu, Cristina V. Lopes, Iris Ma, James Noble, 2024
https://arxiv.org/abs/2402.00247
9. Inferring multiple helper Dafny assertions with LLMs — Alvaro Silva, Alexandra Mendes, Ruben Martins, 2025
https://arxiv.org/abs/2511.00125
10. Rango: Adaptive Retrieval-Augmented Proving for Automated Software Verification — Kyle Thompson et al., 2024
https://scholar.google.com/scholar?q=Rango%3A+Adaptive+Retrieval-Augmented+Proving+for+Automated+Software+Verification
11. Towards Neural Synthesis for SMT-Assisted Proof-Oriented Programming — Saikat Chakraborty et al., 2024
https://scholar.google.com/scholar?q=Towards+Neural+Synthesis+for+SMT-Assisted+Proof-Oriented+Programming
12. Finding Inductive Loop Invariants using Large Language Models — Adharsh Kamath et al., 2023
https://scholar.google.com/scholar?q=Finding+Inductive+Loop+Invariants+using+Large+Language+Models
13. A New Era in Software Security: Towards Self-Healing Software via Large Language Models and Formal Verification — Norbert Tihanyi et al., 2023
https://scholar.google.com/scholar?q=A+New+Era+in+Software+Security%3A+Towards+Self-Healing+Software+via+Large+Language+Models+and+Formal+Verification
14. Rethinking Optimal Verification Granularity for Compute-Efficient Test-Time Scaling — Hao Mark Chen et al., 2025
https://scholar.google.com/scholar?q=Rethinking+Optimal+Verification+Granularity+for+Compute-Efficient+Test-Time+Scaling
15. Heimdall: test-time scaling on the generative verification — Wenlei Shi and Xing Jin, 2025
https://scholar.google.com/scholar?q=Heimdall%3A+test-time+scaling+on+the+generative+verification
16. AI Post Transformers: From Natural Language to Verified Dafny Code — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-14-from-natural-language-to-verified-dafny-8abed9.mp3
17. AI Post Transformers: DeepVerifier: Self-Evolving Research Agents via Rubric-Guided Verification — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/deepverifier-self-evolving-research-agents-via-rubric-guided-verification/
18. AI Post Transformers: Agentic Discovery for Test-Time Scaling — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-12-agentic-discovery-for-test-time-scaling-f9a81f.mp3
19. AI Post Transformers: Trajectory Summaries for Long-Horizon Coding Agents — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-24-trajectory-summaries-for-long-horizon-co-0194be.mp3
Interactive Visualization: DafnyPro for LLM-Assisted Dafny Verification

This episode explores the 2009 seL4 paper and why formally verifying an operating-system kernel matters when that kernel sits at the center of every higher-level security claim. It explains how seL4’s microkernel design keeps only core mechanisms like threads, IPC, interrupts, and memory objects in privileged code, contrasting that with monolithic kernels and showing why minimality makes theorem-proving tractable. The discussion digs into the system’s capability-based authority model, CNodes, explicit reply paths, and untyped memory retyping, arguing that seL4 was designed so allocation, mapping, and permissions remain visible enough for proofs to track every state change. Listeners would find it interesting because the episode draws a sharp line between proving functional correctness and proving true end-to-end security, showing both the power and the limits of formal methods in real systems.

Interactive Visualization: seL4: Proving a Microkernel in C
Sources:
1. seL4: Proving a Microkernel in C
https://www.sigops.org/s/conferences/sosp/2009/papers/klein-sosp09.pdf
2. On Micro-Kernel Construction — Jochen Liedtke, 1995
https://scholar.google.com/scholar?q=On+Micro-Kernel+Construction
3. The Performance of Micro-Kernel-Based Systems — Hermann Hartig, Michael Hohmuth, Jochen Liedtke, Sebastian Schonberg, Jean Wolter, 1997
https://scholar.google.com/scholar?q=The+Performance+of+Micro-Kernel-Based+Systems
4. seL4: Formal Verification of an OS Kernel — Gerwin Klein, Kevin Elphinstone, Gernot Heiser, June Andronick, David Cock and others, 2009
https://scholar.google.com/scholar?q=seL4%3A+Formal+Verification+of+an+OS+Kernel
5. seL4: Formal Verification of an Operating-System Kernel — Gerwin Klein, June Andronick, Kevin Elphinstone, Gernot Heiser, David Cock, Philip Derrin and others, 2010
https://scholar.google.com/scholar?q=seL4%3A+Formal+Verification+of+an+Operating-System+Kernel
6. Translation Validation for a Verified OS Kernel — Thomas Sewell, Magnus Myreen, Gerwin Klein, 2013
https://scholar.google.com/scholar?q=Translation+Validation+for+a+Verified+OS+Kernel
7. Comprehensive Formal Verification of an OS Microkernel — Gerwin Klein, June Andronick, Kevin Elphinstone, Toby Murray, Thomas Sewell, Rafal Kolanski, Gernot Heiser, 2014
https://scholar.google.com/scholar?q=Comprehensive+Formal+Verification+of+an+OS+Microkernel
8. Kernel Design for Isolation and Assurance of Physical Memory — Dhammika Elkaduwe, Philip Derrin, Kevin Elphinstone, 2008
https://scholar.google.com/scholar?q=Kernel+Design+for+Isolation+and+Assurance+of+Physical+Memory
9. seL4 Enforces Integrity — Thomas Sewell, Simon Winwood, Peter Gammie, Toby Murray, June Andronick, Gerwin Klein, 2011
https://scholar.google.com/scholar?q=seL4+Enforces+Integrity
10. seL4: From General Purpose to a Proof of Information Flow Enforcement — Toby Murray, Daniel Matichuk, Matthew Brassil, Peter Gammie, Timothy Bourke, Sean Seefried, Corey Lewis, Xin Gao, Gerwin Klein, 2013
https://scholar.google.com/scholar?q=seL4%3A+From+General+Purpose+to+a+Proof+of+Information+Flow+Enforcement
11. Verified Protection Model of the seL4 Microkernel — Dhammika Elkaduwe, Gerwin Klein, Kevin Elphinstone, 2008
https://scholar.google.com/scholar?q=Verified+Protection+Model+of+the+seL4+Microkernel
12. Time Protection: The Missing OS Abstraction — Qian Ge, Yuval Yarom, Tom Chothia, Gernot Heiser, 2019
https://scholar.google.com/scholar?q=Time+Protection%3A+The+Missing+OS+Abstraction
13. Practical Rely/Guarantee Verification of an Efficient Lock for seL4 on Multicore Architectures — Robert J. Colvin, Ian J. Hayes, Scott Heiner, Peter Höfner, Larissa Meinicke, Roger C. Su, 2024
https://scholar.google.com/scholar?q=Practical+Rely%2FGuarantee+Verification+of+an+Efficient+Lock+for+seL4+on+Multicore+Architectures
14. Modeling Dynamic (De)Allocations of Local Memory for Translation Validation — Abhishek Rose, Sorav Bansal, 2024
https://scholar.google.com/scholar?q=Modeling+Dynamic+%28De%29Allocations+of+Local+Memory+for+Translation+Validation
15. PA-Boot: A Formally Verified Authentication Protocol for Multiprocessor Secure Boot — Zhuoruo Zhang et al., 2022
https://scholar.google.com/scholar?q=PA-Boot%3A+A+Formally+Verified+Authentication+Protocol+for+Multiprocessor+Secure+Boot
16. Prevention of Microarchitectural Covert Channels on an Open-Source 64-bit RISC-V Core — Nils Wistoff et al., 2020
https://scholar.google.com/scholar?q=Prevention+of+Microarchitectural+Covert+Channels+on+an+Open-Source+64-bit+RISC-V+Core
17. Specification and Verification of Side-channel Security for Open-source Processors via Leakage Contracts — Zilong Wang, Gideon Mohr, Klaus von Gleissenthall, Jan Reineke, Marco Guarnieri, 2023
https://scholar.google.com/scholar?q=Specification+and+Verification+of+Side-channel+Security+for+Open-source+Processors+via+Leakage+Contracts
18. AI Post Transformers: From Natural Language to Verified Dafny Code — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-14-from-natural-language-to-verified-dafny-8abed9.mp3
Interactive Visualization: seL4: Proving a Microkernel in C

This episode explores DafnyBench, a benchmark for testing whether large language models can help with one of formal verification’s hardest practical bottlenecks: reconstructing the missing assertions and loop invariants that make Dafny programs verifiable. It explains how formal verification differs from ordinary testing and from theorem proving, and why the paper deliberately frames the task as restoring proof hints in existing verified programs rather than synthesizing correct software from scratch. The discussion digs into benchmark design, including the dataset of 782 single-file Dafny programs, the rule that models must infer both the content and placement of missing hints, and the importance of excluding shortcut tricks like disabling verification. It also highlights a crucial result nuance: 208 files already verify after hint removal, so the reported top score of about 67.8% is more informative when translated into genuine recovery performance on the subset that actually needs new annotations.

Sources:
1. DafnyBench: A Benchmark for Formal Software Verification — Chloe Loughridge, Qinyi Sun, Seth Ahrenbach, Federico Cassano, Chuyue Sun, Ying Sheng, Anish Mudide, Md Rakib Hossain Misu, Nada Amin, Max Tegmark, 2024
http://arxiv.org/abs/2406.08467
2. Clover: Closed-Loop Verifiable Code Generation — Chuyue Sun, Ying Sheng, Oded Padon, Clark Barrett, 2024
https://scholar.google.com/scholar?q=Clover%3A+Closed-Loop+Verifiable+Code+Generation
3. Towards AI-Assisted Synthesis of Verified Dafny Methods — Md Rakib Hossain Misu, Cristina V. Lopes, Iris Ma, James Noble, 2024
https://scholar.google.com/scholar?q=Towards+AI-Assisted+Synthesis+of+Verified+Dafny+Methods
4. LeanDojo: Theorem Proving with Retrieval-Augmented Language Models — Kaiyu Yang et al., 2023
https://scholar.google.com/scholar?q=LeanDojo%3A+Theorem+Proving+with+Retrieval-Augmented+Language+Models
5. LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code — Naman Jain et al., 2024
https://scholar.google.com/scholar?q=LiveCodeBench%3A+Holistic+and+Contamination+Free+Evaluation+of+Large+Language+Models+for+Code
6. Local Success Does Not Compose: Benchmarking Large Language Models for Compositional Formal Verification — Xu Xu et al., 2025
https://scholar.google.com/scholar?q=Local+Success+Does+Not+Compose%3A+Benchmarking+Large+Language+Models+for+Compositional+Formal+Verification
7. A New Era in Software Security: Towards Self-Healing Software via Large Language Models and Formal Verification — Norbert Tihanyi et al., 2023
https://scholar.google.com/scholar?q=A+New+Era+in+Software+Security%3A+Towards+Self-Healing+Software+via+Large+Language+Models+and+Formal+Verification
8. AI Post Transformers: LLM Agents Reason About Code Without Running It — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-15-llm-agents-reason-about-code-without-run-2a1876.mp3
9. AI Post Transformers: SkillsBench for Evaluating Agent Skills — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-14-skillsbench-for-evaluating-agent-skills-58bb1e.mp3

This episode explores DeepMind’s From AGI to ASI as a foresight report that treats human-level general intelligence not as the endpoint, but as a possible stepping stone toward systems that could outperform entire organizations in planning, research, engineering, and coordination. It breaks down how the paper defines AGI and the much more ambitious idea of ASI, then examines the conceptual tools behind that framing, including universal intelligence, AIXI as an idealized reference point, recursive self-improvement, collective intelligence, and the notion of effective compute. The discussion also probes the paper’s method, arguing that it is a structured synthesis of trends and bottlenecks rather than empirical proof, and questions how much precision is needed before such forecasts become meaningful. Listeners would find it interesting because it connects abstract AI theory, concrete scaling dynamics, and real uncertainty about whether progress in models, compute, and autonomy could compound into organization-level superintelligence.

Interactive Visualization: From AGI to ASI and Beyond
Sources:
1. From AGI to ASI — Tim Genewein, Matija Franklin, Alexander Lerchner, Laurent Orseau, Samuel Albanie, Adam Bales, Cole Wyeth, Stephanie Chan, Iason Gabriel, Joel Z. Leibo, Allan Dafoe, Marcus Hutter, Thore Graepel, Shane Legg, 2026
http://arxiv.org/abs/2606.12683
2. From AGI to ASI (https://arxiv.org/abs/2606.12683) — Tim Genewein, Matija Franklin, Alexander Lerchner, Marcus Hutter, Shane Legg, et al., 2026
https://scholar.google.com/scholar?q=From+AGI+to+ASI+%28https%3A%2F%2Farxiv.org%2Fabs%2F2606.12683%29
3. Can Intelligence Explode? (https://arxiv.org/abs/1202.6177) — Marcus Hutter, 2012
https://scholar.google.com/scholar?q=Can+Intelligence+Explode%3F+%28https%3A%2F%2Farxiv.org%2Fabs%2F1202.6177%29
4. Research Priorities for Robust and Beneficial Artificial Intelligence (https://arxiv.org/abs/1602.03506) — Stuart Russell, Daniel Dewey, Max Tegmark, 2015
https://scholar.google.com/scholar?q=Research+Priorities+for+Robust+and+Beneficial+Artificial+Intelligence+%28https%3A%2F%2Farxiv.org%2Fabs%2F1602.03506%29
5. Emerging Practices in Frontier AI Safety Frameworks (https://arxiv.org/abs/2503.04746) — Marie Davidsen Buhl, Ben Bucknall, Tammy Masterson, 2025
https://scholar.google.com/scholar?q=Emerging+Practices+in+Frontier+AI+Safety+Frameworks+%28https%3A%2F%2Farxiv.org%2Fabs%2F2503.04746%29
6. A Theory of Universal Artificial Intelligence based on Algorithmic Complexity (https://arxiv.org/abs/cs/0004001) — Marcus Hutter, 2000
https://scholar.google.com/scholar?q=A+Theory+of+Universal+Artificial+Intelligence+based+on+Algorithmic+Complexity+%28https%3A%2F%2Farxiv.org%2Fabs%2Fcs%2F0004001%29
7. Universal Intelligence: A Definition of Machine Intelligence (https://arxiv.org/abs/0712.3329) — Shane Legg, Marcus Hutter, 2007
https://scholar.google.com/scholar?q=Universal+Intelligence%3A+A+Definition+of+Machine+Intelligence+%28https%3A%2F%2Farxiv.org%2Fabs%2F0712.3329%29
8. A Monte Carlo AIXI Approximation (https://arxiv.org/abs/0909.0801) — Joel Veness, Kee Siong Ng, Marcus Hutter, William Uther, David Silver, 2009 (later published in JAIR, 2011)
https://scholar.google.com/scholar?q=A+Monte+Carlo+AIXI+Approximation+%28https%3A%2F%2Farxiv.org%2Fabs%2F0909.0801%29
9. One Decade of Universal Artificial Intelligence (https://arxiv.org/abs/1202.6153) — Marcus Hutter, 2012
https://scholar.google.com/scholar?q=One+Decade+of+Universal+Artificial+Intelligence+%28https%3A%2F%2Farxiv.org%2Fabs%2F1202.6153%29
10. Universal Intelligence: A Definition of Machine Intelligence — Shane Legg, Marcus Hutter, 2007
https://scholar.google.com/scholar?q=Universal+Intelligence%3A+A+Definition+of+Machine+Intelligence
11. Levels of AGI for Operationalizing Progress on the Path to AGI — Meredith Ringel Morris, Jascha Sohl-Dickstein, Noah Fiedel, Tris Warkentin, Allan Dafoe, Aleksandra Faust, Clement Farabet, Shane Legg, 2023
https://scholar.google.com/scholar?q=Levels+of+AGI+for+Operationalizing+Progress+on+the+Path+to+AGI
12. AI as Normal Technology — Arvind Narayanan, Sayash Kapoor, 2025
https://scholar.google.com/scholar?q=AI+as+Normal+Technology
13. Preparing for the Intelligence Explosion — William MacAskill, Fin Moorhouse, 2025
https://scholar.google.com/scholar?q=Preparing+for+the+Intelligence+Explosion
14. A Rosetta Stone for AI Benchmarks — Anson Ho, Jean-Stanislas Denain, David Atanasov, Samuel Albanie, Rohin Shah, 2025
https://scholar.google.com/scholar?q=A+Rosetta+Stone+for+AI+Benchmarks
15. Measuring AI Ability to Complete Long Tasks — Thomas Kwa et al., 2025
https://scholar.google.com/scholar?q=Measuring+AI+Ability+to+Complete+Long+Tasks
16. PaperBench: Evaluating AI's Ability to Replicate AI Research — Giulio Starace et al., 2025
https://scholar.google.com/scholar?q=PaperBench%3A+Evaluating+AI%27s+Ability+to+Replicate+AI+Research
17. Inverse Scaling in Test-Time Compute — Aryo P. Gema et al., 2025
https://scholar.google.com/scholar?q=Inverse+Scaling+in+Test-Time+Compute
18. The Art of Scaling Test-Time Compute for Large Language Models — Aradhye Agarwal et al., 2025
https://scholar.google.com/scholar?q=The+Art+of+Scaling+Test-Time+Compute+for+Large+Language+Models
19. Test-Time Scaling Makes Overtraining Compute-Optimal — Nicholas Roberts et al., 2026
https://scholar.google.com/scholar?q=Test-Time+Scaling+Makes+Overtraining+Compute-Optimal
20. When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation — Mubashara Akhtar et al., 2026
https://scholar.google.com/scholar?q=When+AI+Benchmarks+Plateau%3A+A+Systematic+Study+of+Benchmark+Saturation
21. How Bad is Training on Synthetic Data? A Statistical Analysis of Language Model Collapse — Mohamed El Amine Seddik et al., 2024
https://scholar.google.com/scholar?q=How+Bad+is+Training+on+Synthetic+Data%3F+A+Statistical+Analysis+of+Language+Model+Collapse
22. Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data — Matthias Gerstgrasser et al., 2024
https://scholar.google.com/scholar?q=Is+Model+Collapse+Inevitable%3F+Breaking+the+Curse+of+Recursion+by+Accumulating+Real+and+Synthetic+Data
23. MLGym: A New Framework and Benchmark for Advancing AI Research Agents — Deepak Nathani et al., 2025
https://scholar.google.com/scholar?q=MLGym%3A+A+New+Framework+and+Benchmark+for+Advancing+AI+Research+Agents
24. MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation — Qian Huang et al., 2023
https://scholar.google.com/scholar?q=MLAgentBench%3A+Evaluating+Language+Agents+on+Machine+Learning+Experimentation
25. AI Post Transformers: Technical AGI Safety and Security Framework — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-03-technical-agi-safety-and-security-framew-f27316.mp3
26. AI Post Transformers: Unified Neural Scaling Laws Across Regimes — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-07-unified-neural-scaling-laws-across-regim-292e2d.mp3
27. AI Post Transformers: TUMIX Multi-Agent Test-Time Scaling with Tools — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-22-tumix-multi-agent-test-time-scaling-with-40671c.mp3
28. AI Post Transformers: Test-time Scaling for Multi-Agent Collaborative Reasoning — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-22-test-time-scaling-for-multi-agent-collab-082570.mp3
Interactive Visualization: From AGI to ASI and Beyond

This episode explores MiniMax Sparse Attention, a long-context transformer design that aims to preserve dense-model quality at million-token scale while sharply reducing the quadratic compute and memory costs of standard attention. It explains how the method combines Grouped Query Attention with blockwise sparse retrieval: a lightweight Index Branch scores past context in blocks, forces a recent local block to stay visible, selects top-k candidate regions, and then lets a Main Branch run exact softmax attention only inside those chosen blocks. The discussion places the paper alongside Longformer, BigBird, Routing Transformers, MInference, and Native Sparse Attention, arguing that its main contribution is a simpler, more GPU-friendly routing scheme that could make sparse attention practical at deployment time. Listeners would find it interesting because it focuses on the real technical tension behind ultra-long-context models: whether this kind of sparse routing can reliably recover rare distant evidence, or whether it mainly wins through recency bias and careful systems engineering.

Sources:
1. MiniMax Sparse Attention — Xunhao Lai, Weiqi Xu, Yufeng Yang, Qiaorui Chen, Yang Xu, Lunbin Zeng, Xiaolong Li, Haohai Sun, Haichao Zhu, Vito Zhang, Pengyu Zhao, 2026
http://arxiv.org/abs/2606.13392
2. Longformer: The Long-Document Transformer — Iz Beltagy, Matthew E. Peters, Arman Cohan, 2020
https://scholar.google.com/scholar?q=Longformer%3A+The+Long-Document+Transformer
3. Big Bird: Transformers for Longer Sequences — Manzil Zaheer, Guru Guruganesh, Avinava Dubey, Joshua Ainslie, Chris Alberti, Amr Ahmed, 2020
https://scholar.google.com/scholar?q=Big+Bird%3A+Transformers+for+Longer+Sequences
4. Efficient Content-Based Sparse Attention with Routing Transformers — Aurko Roy, Mohammad Saffar, Ashish Vaswani, David Grangier, 2020
https://scholar.google.com/scholar?q=Efficient+Content-Based+Sparse+Attention+with+Routing+Transformers
5. Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention — Jingyang Yuan, Huazuo Gao, Damai Dai, Junyu Luo, Liang Zhao, Wenfeng Liang, Wangding Zeng, 2025
https://scholar.google.com/scholar?q=Native+Sparse+Attention%3A+Hardware-Aligned+and+Natively+Trainable+Sparse+Attention
6. GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints — Joshua Ainslie et al., 2025
https://scholar.google.com/scholar?q=GQA%3A+Training+Generalized+Multi-Query+Transformer+Models+from+Multi-Head+Checkpoints
7. Optimizing Mixture of Block Attention — Guangxuan Xiao et al., 2025
https://scholar.google.com/scholar?q=Optimizing+Mixture+of+Block+Attention
8. MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention — Huiqiang Jiang et al., 2024
https://scholar.google.com/scholar?q=MInference+1.0%3A+Accelerating+Pre-filling+for+Long-Context+LLMs+via+Dynamic+Sparse+Attention
9. DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads — Guangxuan Xiao et al., 2024
https://scholar.google.com/scholar?q=DuoAttention%3A+Efficient+Long-Context+LLM+Inference+with+Retrieval+and+Streaming+Heads
10. FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision — Jay Shah, Ganesh Bikshandi, Ying Zhang, Vijay Thakkar, Pradeep Ramani, Tri Dao, 2024
https://scholar.google.com/scholar?q=FlashAttention-3%3A+Fast+and+Accurate+Attention+with+Asynchrony+and+Low-precision
11. FIER: Fine-Grained and Efficient KV Cache Retrieval for Long-context LLM Inference — Dongwei Wang et al., 2025
https://arxiv.org/abs/2508.08256
12. SCBench: A KV Cache-Centric Analysis of Long-Context Methods — Yucheng Li et al., 2024
https://arxiv.org/abs/2412.10319
13. Streaming Video Question-Answering with In-context Video KV-Cache Retrieval — Shangzhe Di et al., 2025
https://arxiv.org/abs/2503.00540
14. Native Hybrid Attention for Efficient Sequence Modeling — Jusen Du et al., 2025
https://arxiv.org/abs/2510.07019
15. Rope to Nope and Back Again: A New Hybrid Attention Strategy — Bowen Yang et al., 2025
https://arxiv.org/abs/2501.18795
16. AI Post Transformers: Optimizing Mixture of Block Attention Through Statistical Theory — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-18-optimizing-mixture-of-block-attention-th-214f91.mp3
17. AI Post Transformers: Deep Kernel Fusion for Transformer Decoding — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-deep-kernel-fusion-for-transformer-decod-b1a703.mp3
18. AI Post Transformers: Mooncake for KV Cache-Centric LLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-05-mooncake-for-kv-cache-centric-llm-servin-1086d0.mp3
19. AI Post Transformers: How Induction Heads Emerge in Transformers — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-03-how-induction-heads-emerge-in-transforme-a7bfcb.mp3
20. AI Post Transformers: δ-mem and Online Memory for LLMs — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-13-d-mem-and-online-memory-for-llms-6622fa.mp3
21. AI Post Transformers: Ministral 3: Cascade Distillation for Long-Context Multimodal Models — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-cascade-distillation-for-long-context-mu-0ebd1a.mp3
Interactive Visualization: MiniMax Sparse Attention at Million-Token Scale

This episode explores the ACE proposal from AMD, Intel, and the x86 Ecosystem Advisory Group, which would add matrix-native AI instructions to x86 CPUs so transformer workloads can run dense linear algebra more efficiently without changing the models themselves. It explains why AVX10 and VNNI fall short for GEMM-heavy inference, introducing outer-product updates, 2-D tile registers, and the reuse of the AMX palette model so operating systems and compilers can handle the new state within familiar x86 mechanisms. The discussion also challenges the proposal’s headline 16x compute-density claim for INT8 and BF16, arguing that real speed depends on full-kernel costs like packing, memory traffic, conversions, tails, and cache behavior. It also examines OCP FP8, MXFP8, MXINT8, and BF16 support as a sign that low-precision AI now depends on tight coordination between ISA design, quantization rules, and kernel implementation, making the proposal interesting both technically and strategically.

Interactive Visualization: ACE: Matrix-Native AI Extensions for x86
Sources:
1. ACE: Matrix-Native AI Extensions for x86
https://x86ecosystem.org/wp-content/uploads/2026/03/ACE-Whitepaper-v1.pdf
2. The AI Compute Extensions (ACE) for x86 — Stuart Biles, Brian Thompto, Michael Estlick, Eric Schwarz, Thomas Fox, Gabriel Loh, Marius Evers, Michael Clark, Alexander Heinecke, Pradeep Dubey, Ido Ouziel, 2026
https://scholar.google.com/scholar?q=The+AI+Compute+Extensions+%28ACE%29+for+x86
3. A matrix math facility for Power ISA(TM) processors — José E. Moreira, Kit Barton, Steven Battle, Peter Bergner, Ramon Bertran and others, 2021
https://scholar.google.com/scholar?q=A+matrix+math+facility+for+Power+ISA%28TM%29+processors
4. Hello SME! Generating Fast Matrix Multiplication Kernels Using the Scalable Matrix Extension — Stefan Remke, Alexander Breuer, 2024
https://scholar.google.com/scholar?q=Hello+SME%21+Generating+Fast+Matrix+Multiplication+Kernels+Using+the+Scalable+Matrix+Extension
5. SparAMX: Accelerating Compressed LLMs Token Generation on AMX-powered CPUs — Ahmed F. AbouElhamayed, Jordan Dotzel, Yash Akhauri, Chi-Chih Chang, Sameh Gobriel, J. Pablo Muñoz, Vui Seng Chua, Nilesh Jain, Mohamed S. Abdelfattah, 2025
https://scholar.google.com/scholar?q=SparAMX%3A+Accelerating+Compressed+LLMs+Token+Generation+on+AMX-powered+CPUs
6. Automating the Last-Mile for High Performance Dense Linear Algebra — Richard Michael Veras, Tze Meng Low, Tyler Michael Smith, Robert A. van de Geijn, Franz Franchetti, 2017
https://scholar.google.com/scholar?q=Automating+the+Last-Mile+for+High+Performance+Dense+Linear+Algebra
7. Microscaling Data Formats for Deep Learning — Bita Darvish Rouhani et al., 2023
https://scholar.google.com/scholar?q=Microscaling+Data+Formats+for+Deep+Learning
8. Recipes for Pre-training LLMs with MXFP8 — Asit Mishra, Dusan Stosic, Simon Layton, 2025
https://scholar.google.com/scholar?q=Recipes+for+Pre-training+LLMs+with+MXFP8
9. Fast Matrix Multiplication via Compiler-only Layered Data Reorganization and Intrinsic Lowering — Braedy Kuzma et al., 2023
https://scholar.google.com/scholar?q=Fast+Matrix+Multiplication+via+Compiler-only+Layered+Data+Reorganization+and+Intrinsic+Lowering
10. THOR: A Non-Speculative Value Dependent Timing Side Channel Attack Exploiting Intel AMX — Farshad Dizani et al., 2025
https://scholar.google.com/scholar?q=THOR%3A+A+Non-Speculative+Value+Dependent+Timing+Side+Channel+Attack+Exploiting+Intel+AMX
11. Compute or Load KV Cache? Why Not Both? (https://arxiv.org/abs/2410.03065) — Shuowei Jin et al., 2024
https://scholar.google.com/scholar?q=Compute+or+Load+KV+Cache%3F+Why+Not+Both%3F+%28https%3A%2F%2Farxiv.org%2Fabs%2F2410.03065%29
12. SOLE: Hardware-Software Co-design of Softmax and LayerNorm for Efficient Transformer Inference (https://arxiv.org/abs/2510.17189) — Wenxun Wang et al., 2025
https://scholar.google.com/scholar?q=SOLE%3A+Hardware-Software+Co-design+of+Softmax+and+LayerNorm+for+Efficient+Transformer+Inference+%28https%3A%2F%2Farxiv.org%2Fabs%2F2510.17189%29
13. Towards Fully FP8 GEMM LLM Training at Scale (https://arxiv.org/abs/2505.20524) — Alejandro Hernandez-Cano et al., 2025
https://scholar.google.com/scholar?q=Towards+Fully+FP8+GEMM+LLM+Training+at+Scale+%28https%3A%2F%2Farxiv.org%2Fabs%2F2505.20524%29
14. Pretraining Large Language Models with NVFP4 (https://arxiv.org/abs/2509.25149) — Felix Abecassis et al. (NVIDIA), 2025
https://scholar.google.com/scholar?q=Pretraining+Large+Language+Models+with+NVFP4+%28https%3A%2F%2Farxiv.org%2Fabs%2F2509.25149%29
15. SageAttention3: Microscaling FP4 Attention for Inference and An Exploration of 8-Bit Training (https://arxiv.org/abs/2505.11594) — Jintao Zhang et al., 2025
https://scholar.google.com/scholar?q=SageAttention3%3A+Microscaling+FP4+Attention+for+Inference+and+An+Exploration+of+8-Bit+Training+%28https%3A%2F%2Farxiv.org%2Fabs%2F2505.11594%29
16. AI Post Transformers: Memory-Bound, Not Bandwidth-Limited Batch-1 LLM Decode — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-02-memory-bound-not-bandwidth-limited-batch-114799.mp3
17. AI Post Transformers: LPU Chip for Low-Latency LLM Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-20-lpu-chip-for-low-latency-llm-inference-be13c3.mp3
18. AI Post Transformers: Deep Kernel Fusion for Transformer Decoding — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-deep-kernel-fusion-for-transformer-decod-b1a703.mp3
19. AI Post Transformers: FlatAttention for Tile-Based Accelerator Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-flatattention-for-tile-based-accelerator-56e6ca.mp3
Interactive Visualization: ACE: Matrix-Native AI Extensions for x86

This episode explores a 2025 paper on using Dafny as a hidden, verification-aware intermediate language for AI code generation, where a model first produces a formal specification and verified implementation before compiling it into ordinary Python. It examines the paper’s central trust claim: formal verification can prove that the generated code satisfies the hidden spec, but it cannot prove that the hidden spec actually matches the user’s intent, making spec alignment a separate and critical failure point. The discussion uses the paper’s fibfib example and HumanEval results to unpack that distinction, noting that the Dafny-only pipeline trails direct Python generation, while the best reported score comes only after falling back to unverified Python when the verification loop fails to converge. Listeners would find it interesting because it gives a concrete, nuanced look at where AI coding assistants can become more reliable, where the guarantees stop, and why neuro-symbolic workflows may matter most for tightly specified code like algorithms, parsers, and protocol logic.

Sources:
1. Dafny as Verification-Aware Intermediate Language for Code Generation — Yue Chen Li, Stefan Zetzsche, Siva Somayyajula, 2025
http://arxiv.org/abs/2501.06283
2. Dafny: An Automatic Program Verifier for Functional Correctness — K. Rustan M. Leino, 2010
https://scholar.google.com/scholar?q=Dafny%3A+An+Automatic+Program+Verifier+for+Functional+Correctness
3. Towards AI-Assisted Synthesis of Verified Dafny Methods — Md Rakib Hossain Misu, Cristina V. Lopes, Iris Ma, James Noble, 2024
https://scholar.google.com/scholar?q=Towards+AI-Assisted+Synthesis+of+Verified+Dafny+Methods
4. Clover: Closed-Loop Verifiable Code Generation — Chuyue Sun, Ying Sheng, Oded Padon, Clark Barrett, 2024
https://scholar.google.com/scholar?q=Clover%3A+Closed-Loop+Verifiable+Code+Generation
5. VerMCTS: Synthesizing Multi-Step Programs using a Verifier, a Large Language Model, and Tree Search — David Brandfonbrener, Simon Henniger, Sibi Raja, Tarun Prasad, Chloe Loughridge, Federico Cassano, Sabrina Ruixin Hu, Jianang Yang, William E. Byrd, Robert Zinkov, Nada Amin, 2024
https://scholar.google.com/scholar?q=VerMCTS%3A+Synthesizing+Multi-Step+Programs+using+a+Verifier%2C+a+Large+Language+Model%2C+and+Tree+Search
6. Laurel: Generating Dafny Assertions Using Large Language Models — Eric Mugnier, Emmanuel Anaya Gonzalez, Ranjit Jhala, Nadia Polikarpova, Yuanyuan Zhou, 2024
https://scholar.google.com/scholar?q=Laurel%3A+Generating+Dafny+Assertions+Using+Large+Language+Models
7. DafnyBench: A Benchmark for Formal Software Verification — Chloe Loughridge et al., 2024
https://scholar.google.com/scholar?q=DafnyBench%3A+A+Benchmark+for+Formal+Software+Verification
8. Evaluating Large Language Models Trained on Code — Mark Chen et al., 2021
https://scholar.google.com/scholar?q=Evaluating+Large+Language+Models+Trained+on+Code
9. Baking for Dafny: A CakeML Backend for Dafny — Daniel Nezamabadi, Magnus Myreen, 2025
https://scholar.google.com/scholar?q=Baking+for+Dafny%3A+A+CakeML+Backend+for+Dafny
10. Intent-aligned Formal Specification Synthesis via Traceable Refinement — Zhe Ye et al., 2026
https://scholar.google.com/scholar?q=Intent-aligned+Formal+Specification+Synthesis+via+Traceable+Refinement
11. Combining LLM Code Generation with Formal Specifications and Reactive Program Synthesis — William Murphy et al., 2024
https://scholar.google.com/scholar?q=Combining+LLM+Code+Generation+with+Formal+Specifications+and+Reactive+Program+Synthesis
12. StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback — Shihan Dou et al., 2024
https://scholar.google.com/scholar?q=StepCoder%3A+Improve+Code+Generation+with+Reinforcement+Learning+from+Compiler+Feedback
13. InspectCoder: Dynamic Analysis-Enabled Self Repair through interactive LLM-Debugger Collaboration — Yunkun Wang et al., 2025
https://scholar.google.com/scholar?q=InspectCoder%3A+Dynamic+Analysis-Enabled+Self+Repair+through+interactive+LLM-Debugger+Collaboration
14. FormalSpecCpp: A Dataset of C++ Formal Specifications created using LLMs — Madhurima Chakraborty et al., 2025
https://scholar.google.com/scholar?q=FormalSpecCpp%3A+A+Dataset+of+C%2B%2B+Formal+Specifications+created+using+LLMs
15. AI Post Transformers: From Natural Language to Verified Dafny Code — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-14-from-natural-language-to-verified-dafny-8abed9.mp3
16. AI Post Transformers: Program Synthesis with Large Language Models — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-20-program-synthesis-with-large-language-mo-b962ec.mp3
17. AI Post Transformers: Generative File Systems: Replacing Code with Formal Specifications — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-18-generative-file-systems-replacing-code-w-414029.mp3

This episode explores whether large language models can help mainstream developers write code that is not just plausible, but formally verified in systems like Dafny, Nagini, and Verus. It explains the core ideas behind machine-checked correctness, including contracts, SMT solvers, loop invariants, and why verification is far stricter than passing tests or matching statistical behavior. The discussion highlights the paper’s main argument that the real bottleneck is not only generating implementations, but also generating the formal scaffolding and annotations that proofs require, especially around loops. Listeners get a clear view of how the authors evaluate LLMs with verifier feedback and extra validation to catch weakened specifications, making the episode interesting for anyone curious about whether AI can close the trust gap in code generation.

Sources:
1. Can LLMs Enable Verification in Mainstream Programming? — Aleksandr Shefer, Igor Engel, Stanislav Alekseev, Daniil Berezun, Ekaterina Verbitskaia, Anton Podkopaev, 2025
http://arxiv.org/abs/2503.14183
2. Inferring Loop Invariants using Postconditions — Carlo A. Furia, Bertrand Meyer, 2009
https://scholar.google.com/scholar?q=Inferring+Loop+Invariants+using+Postconditions
3. Inferring Loop Invariants by Mutation, Dynamic Analysis, and Static Checking — Juan P. Galeotti, Carlo A. Furia, Eva May, Gordon Fraser, Andreas Zeller, 2014
https://scholar.google.com/scholar?q=Inferring+Loop+Invariants+by+Mutation%2C+Dynamic+Analysis%2C+and+Static+Checking
4. LoopInvGen: A Loop Invariant Generator based on Precondition Inference — Saswat Padhi, Rahul Sharma, Todd Millstein, 2017
https://scholar.google.com/scholar?q=LoopInvGen%3A+A+Loop+Invariant+Generator+based+on+Precondition+Inference
5. On Scaling Data-Driven Loop Invariant Inference — Sahil Bhatia, Saswat Padhi, Nagarajan Natarajan, Rahul Sharma, Prateek Jain, 2019
https://scholar.google.com/scholar?q=On+Scaling+Data-Driven+Loop+Invariant+Inference
6. Dafny: An Automatic Program Verifier for Functional Correctness — K. Rustan M. Leino, 2010
https://scholar.google.com/scholar?q=Dafny%3A+An+Automatic+Program+Verifier+for+Functional+Correctness
7. Nagini: A Static Verifier for Python — Marco Eilers, Peter Muller, 2018
https://scholar.google.com/scholar?q=Nagini%3A+A+Static+Verifier+for+Python
8. Verus: Verifying Rust Programs using Linear Ghost Types (extended version) — Andrea Lattuada, Travis Hance, Chanhee Cho, Matthias Brun, Isitha Subasinghe, Yi Zhou, Jon Howell, Bryan Parno, Chris Hawblitzel, 2023
https://scholar.google.com/scholar?q=Verus%3A+Verifying+Rust+Programs+using+Linear+Ghost+Types+%28extended+version%29
9. Can LLMs Enable Verification in Mainstream Programming? — Aleksandr Shefer, Igor Engel, Stanislav Alekseev, Daniil Berezun, Ekaterina Verbitskaia, Anton Podkopaev, 2025
https://scholar.google.com/scholar?q=Can+LLMs+Enable+Verification+in+Mainstream+Programming%3F
10. Clover: Closed-loop verifiable code generation — Chuyue Sun, Ying Sheng, Oded Padon, Clark Barrett, 2024
https://scholar.google.com/scholar?q=Clover%3A+Closed-loop+verifiable+code+generation
11. Alphaverus: Bootstrapping formally verified code generation through self-improving translation and treefinement — Pranjal Aggarwal, Bryan Parno, Sean Welleck, 2024
https://scholar.google.com/scholar?q=Alphaverus%3A+Bootstrapping+formally+verified+code+generation+through+self-improving+translation+and+treefinement
12. Can large language models transform natural language intent into formal method postconditions? — Madeline Endres, Sarah Fakhoury, Saikat Chakraborty, Shuvendu K. Lahiri, 2024
https://scholar.google.com/scholar?q=Can+large+language+models+transform+natural+language+intent+into+formal+method+postconditions%3F
13. Laurel: Generating Dafny assertions using large language models — Eric Mugnier, Emmanuel Anaya Gonzalez, Ranjit Jhala, Nadia Polikarpova, Yuanyuan Zhou, 2024
https://scholar.google.com/scholar?q=Laurel%3A+Generating+Dafny+assertions+using+large+language+models
14. Finding Inductive Loop Invariants using Large Language Models — Adharsh Kamath, Aditya Senthilnathan, Saikat Chakraborty, Pantazis Deligiannis, Shuvendu K. Lahiri, Akash Lal, Aseem Rastogi, Subhajit Roy, Rahul Sharma, 2023
https://scholar.google.com/scholar?q=Finding+Inductive+Loop+Invariants+using+Large+Language+Models
15. Towards Formal Verification of LLM-Generated Code from Natural Language Prompts — Aaron Councilman et al., 2025
https://arxiv.org/abs/2507.13290
16. Enchanting Program Specification Synthesis by Large Language Models using Static Analysis and Program Verification — Cheng Wen et al., 2024
https://arxiv.org/abs/2404.00762
17. SpecGen: Automated Generation of Formal Program Specifications via Large Language Models — Lezhi Ma et al., 2024
https://arxiv.org/abs/2401.08807
18. Beyond Postconditions: Can Large Language Models infer Formal Contracts for Automatic Software Verification? — Cedric Richter and Heike Wehrheim, 2025
https://arxiv.org/abs/2510.12702
19. Guiding LLM-based Loop Invariant Synthesis via Feedback on Local Reasoning Errors — Tianchi Li et al., 2026
https://arxiv.org/abs/2605.17914
20. Loop Invariant Generation: A Hybrid Framework of Reasoning optimised LLMs and SMT Solvers — Varun Bharti et al., 2025
https://arxiv.org/abs/2508.00419
21. LLM For Loop Invariant Generation and Fixing: How Far Are We? — Mostafijur Rahman Akhond et al., 2025
https://arxiv.org/abs/2511.06552
22. AI Post Transformers: From Natural Language to Verified Dafny Code — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-14-from-natural-language-to-verified-dafny-8abed9.mp3
23. AI Post Transformers: Generative File Systems: Replacing Code with Formal Specifications — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-18-generative-file-systems-replacing-code-w-414029.mp3
24. AI Post Transformers: Trajectory Summaries for Long-Horizon Coding Agents — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-24-trajectory-summaries-for-long-horizon-co-0194be.mp3
Interactive Visualization: Can LLMs Enable Mainstream Formal Verification?

This episode explores a 2026 study on turning long natural-language programming problems into Dafny code that can be formally verified, asking whether AI systems can produce code that is not just fluent but provably correct. It explains how Dafny uses preconditions, postconditions, loop invariants, and proof obligations, and why weak specifications can lead to vacuous “verified” programs that still fail to capture the real task. The discussion highlights the paper’s NL2VC-60 benchmark of hand-written verified solutions to UVa-style algorithm problems, along with experiments comparing plain prompting, signature-guided prompting, and self-healing loops that revise code using verifier feedback and additional uDebug testing. Listeners would find it interesting because it gets at the core trust problem in AI coding: whether formal methods can make generated software more reliable, and where the real bottleneck remains the human effort required to write strong specifications.

Sources:
1. From Natural Language to Verified Code: Toward AI Assisted Problem-to-Code Generation with Dafny-Based Formal Verification — Md Erfan, Md Kamal Hossain Chowdhury, Ahmed Ryan, Md Rayhanur Rahman, 2026
http://arxiv.org/abs/2604.22601
2. Dafny: An Automatic Program Verifier for Functional Correctness — K. Rustan M. Leino, 2010
https://scholar.google.com/scholar?q=Dafny%3A+An+Automatic+Program+Verifier+for+Functional+Correctness
3. seL4: Formal Verification of an Operating-System Kernel — Gerwin Klein, Kevin Elphinstone, Gernot Heiser, June Andronick, David Cock, et al., 2009
https://scholar.google.com/scholar?q=seL4%3A+Formal+Verification+of+an+Operating-System+Kernel
4. Formal verification of a realistic compiler — Xavier Leroy, 2009
https://scholar.google.com/scholar?q=Formal+verification+of+a+realistic+compiler
5. Modularity, Code Specialization, and Zero-Cost Abstractions for Program Verification — Son Ho, Aymeric Fromherz, Jonathan Protzenko, 2021
https://scholar.google.com/scholar?q=Modularity%2C+Code+Specialization%2C+and+Zero-Cost+Abstractions+for+Program+Verification
6. Towards AI-Assisted Synthesis of Verified Dafny Methods — Md Rakib Hossain Misu, Cristina V. Lopes, Iris Ma, James Noble, 2024
https://scholar.google.com/scholar?q=Towards+AI-Assisted+Synthesis+of+Verified+Dafny+Methods
7. DafnyBench: A Benchmark for Formal Software Verification — Chloe Loughridge et al., 2024
https://scholar.google.com/scholar?q=DafnyBench%3A+A+Benchmark+for+Formal+Software+Verification
8. Can LLMs Enable Verification in Mainstream Programming? — Aleksandr Shefer, Igor Engel, Stanislav Alekseev, Daniil Berezun, Ekaterina Verbitskaia, Anton Podkopaev, 2025
https://scholar.google.com/scholar?q=Can+LLMs+Enable+Verification+in+Mainstream+Programming%3F
9. Dafny as Verification-Aware Intermediate Language for Code Generation — Yue Chen Li, Stefan Zetzsche, Siva Somayyajula, 2025
https://scholar.google.com/scholar?q=Dafny+as+Verification-Aware+Intermediate+Language+for+Code+Generation
10. ATLAS: Automated Toolkit for Large-Scale Verified Code Synthesis — Mantas Baksys et al., 2025
https://scholar.google.com/scholar?q=ATLAS%3A+Automated+Toolkit+for+Large-Scale+Verified+Code+Synthesis
11. DafnyPro: LLM-Assisted Automated Verification for Dafny Programs — Debangshu Banerjee, Olivier Bouissou, Stefan Zetzsche, 2026
https://scholar.google.com/scholar?q=DafnyPro%3A+LLM-Assisted+Automated+Verification+for+Dafny+Programs
12. Neuro Symbolic Reasoning for Planning: Counterexample Guided Inductive Synthesis using Large Language Models and Satisfiability Solving — Sumit Kumar Jha et al., 2023
https://scholar.google.com/scholar?q=Neuro+Symbolic+Reasoning+for+Planning%3A+Counterexample+Guided+Inductive+Synthesis+using+Large+Language+Models+and+Satisfiability+Solving
13. Property-Guided LLM Program Synthesis for Planning — Andre G. Pereira, Augusto B. Correa, Jendrik Seipp, 2026
https://scholar.google.com/scholar?q=Property-Guided+LLM+Program+Synthesis+for+Planning
14. Finding Inductive Loop Invariants using Large Language Models — Adharsh Kamath et al., 2023
https://scholar.google.com/scholar?q=Finding+Inductive+Loop+Invariants+using+Large+Language+Models
15. LLM For Loop Invariant Generation and Fixing: How Far Are We? — Mostafijur Rahman Akhond, Saikat Chakraborty, Gias Uddin, 2025
https://scholar.google.com/scholar?q=LLM+For+Loop+Invariant+Generation+and+Fixing%3A+How+Far+Are+We%3F
16. Type-Constrained Code Generation with Language Models — Niels Mundler et al., 2025
https://scholar.google.com/scholar?q=Type-Constrained+Code+Generation+with+Language+Models
17. Invariant-based Program Repair — Omar I. Al-Bataineh, 2024
https://scholar.google.com/scholar?q=Invariant-based+Program+Repair
18. AI Post Transformers: Program Synthesis with Large Language Models — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-20-program-synthesis-with-large-language-mo-b962ec.mp3
19. AI Post Transformers: SGLang for Faster Structured LLM Programs — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-06-sglang-for-faster-structured-llm-program-c59f1c.mp3
20. AI Post Transformers: SkillsBench for Evaluating Agent Skills — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-14-skillsbench-for-evaluating-agent-skills-58bb1e.mp3

This episode explores the paper Test-Time Training with KV Binding Is Secretly Linear Attention and asks whether KV-binding test-time training is really doing online memorization or instead behaving like learned linear attention with sequence-specific fast weights. It explains how this approach differs from a standard transformer KV cache, situates it within earlier test-time-training work on expressive hidden states, and connects it to the broader push for long-context models that avoid quadratic softmax attention costs. The discussion highlights several findings that weaken the retrieval-style memory story: converged models show a query-key mismatch, replacing queries with keys barely changes aggregate performance, stronger inner-loop optimization does not reliably help, and even switching from descent to ascent can still work. Listeners would find it interesting because the episode reframes a flashy mechanism in simpler algebraic terms, clarifies which equivalence claims are exact versus empirical, and shows how that shift could change how researchers think about memory and efficiency in next-generation sequence models.

Sources:
1. Test-Time Training with KV Binding Is Secretly Linear Attention — Junchen Liu, Sven Elflein, Or Litany, Zan Gojcic, Ruilong Li, 2026
http://arxiv.org/abs/2602.21204
2. Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention — Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François Fleuret, 2020
https://arxiv.org/abs/2006.16236
3. Linear Transformers Are Secretly Fast Weight Programmers — Imanol Schlag, Kazuki Irie, Jürgen Schmidhuber, 2021
https://arxiv.org/abs/2102.11174
4. Learning to (Learn at Test Time): RNNs with Expressive Hidden States — Yu Sun, Xinhao Li, Karan Dalal, Jiarui Xu, Xinlei Chen, Tatsunori Hashimoto, Carlos Guestrin, et al., 2024
https://arxiv.org/abs/2407.04620
5. Test-Time Training with KV Binding Is Secretly Linear Attention — Junchen Liu, Sven Elflein, Or Litany, Zan Gojcic, Ruilong Li, 2026
https://arxiv.org/abs/2602.21204
6. Titans: Learning to Memorize at Test Time — Ali Behrouz, Peilin Zhong, Vahab Mirrokni, 2024
https://scholar.google.com/scholar?q=Titans%3A+Learning+to+Memorize+at+Test+Time
7. End-to-End Test-Time Training for Long Context — Arnuv Tandon et al., 2025
https://scholar.google.com/scholar?q=End-to-End+Test-Time+Training+for+Long+Context
8. Understanding Factual Recall in Transformers via Associative Memories — Eshaan Nichani, Jason D. Lee, Alberto Bietti, 2024
https://arxiv.org/abs/2412.06538
9. Quantifying Logical Consistency in Transformers via Query-Key Alignment — Eduard Tulchinskii, Anastasia Voznyuk, Laida Kushnareva, Andrei Andriiainen, Irina Piontkovskaya, Evgeny Burnaev, Serguei Barannikov, 2025
https://arxiv.org/abs/2502.17017
10. Dissecting Query-Key Interaction in Vision Transformers — Xu Pan, Aaron Philip, Ziqian Xie, Odelia Schwartz, 2024
https://arxiv.org/abs/2405.14880
11. Improved Test-Time Adaptation for Domain Generalization — Liang Chen, Yong Zhang, Yibing Song, Ying Shan, Lingqiao Liu, 2023
https://arxiv.org/abs/2304.04494
12. AdaShadow: Responsive Test-time Model Adaptation in Non-stationary Mobile Environments — Cheng Fang, Sicong Liu, Zimu Zhou, Bin Guo, Jiaqi Tang, Ke Ma, Zhiwen Yu, 2024
https://arxiv.org/abs/2410.08256
13. Beyond Model Adaptation at Test Time: A Survey — Zehao Xiao, Cees G. M. Snoek, 2024
https://arxiv.org/abs/2411.03687
14. AI Post Transformers: Atlas: Test-Time Memory for Long Contexts — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-11-atlas-test-time-memory-for-long-contexts-1d5545.mp3
15. AI Post Transformers: Parallelizing DeltaNet Linear Transformers over Sequence Length — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-18-parallelizing-deltanet-linear-transforme-2d0377.mp3
16. AI Post Transformers: Gated Linear Attention for Efficient Long Sequences — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-18-gated-linear-attention-for-efficient-lon-c858ab.mp3
17. AI Post Transformers: δ-mem and Online Memory for LLMs — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-13-d-mem-and-online-memory-for-llms-6622fa.mp3
18. AI Post Transformers: Do Transformers Need Three Projections? — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-11-do-transformers-need-three-projections-c227d6.mp3

This episode explores a June 2026 paper on when document-specific LoRA adapters actually help compared with standard retrieval-augmented generation, especially once a model’s KV cache has been aggressively compressed. It walks through the core mechanics of RAG, LoRA, prefill vs. decode costs, parametric retrieval augmentation, and the Compactor method used to rank and retain only part of a document’s cached attention state. The main argument is that LoRA is not a replacement for explicit retrieved text: when most document context is still intact, the adapter adds little, but under severe compression it becomes much more useful, recovering roughly 13 to 21 ROUGE-L points when the document cache is completely removed. Listeners would find it interesting because it turns a vague “LoRA vs. RAG” debate into a concrete systems question about memory budgets, repeated question answering, and the tradeoff between inspectable evidence and lossy parameter-side memory.

Sources:
1. Rethinking LoRA Memory Through the Lens of KV Cache Compression — Chunsheng Zuo, Liaoyaqi Wang, William Jurayj, William Fleshman, Benjamin Van Durme, 2026
http://arxiv.org/abs/2606.05698
2. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks — Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, et al., 2020
https://scholar.google.com/scholar?q=Retrieval-Augmented+Generation+for+Knowledge-Intensive+NLP+Tasks
3. Parametric Retrieval Augmented Generation — Weihang Su, Yichen Tang, Qingyao Ai, Junxi Yan, et al., 2025
https://scholar.google.com/scholar?q=Parametric+Retrieval+Augmented+Generation
4. Understanding Parametric Knowledge Injection in Retrieval-Augmented Generation — Minghao Tang, Shiyu Ni, Jingtong Wu, Zengxin Han, Keping Bi, 2025
https://scholar.google.com/scholar?q=Understanding+Parametric+Knowledge+Injection+in+Retrieval-Augmented+Generation
5. Rethinking LoRA Memory Through the Lens of KV Cache Compression — Chunsheng Zuo, Liaoyaqi Wang, William Jurayj, William Fleshman, Benjamin Van Durme, 2026
https://scholar.google.com/scholar?q=Rethinking+LoRA+Memory+Through+the+Lens+of+KV+Cache+Compression
6. Training Plug-n-Play Knowledge Modules with Deep Context Distillation — Lucas Caccia et al., 2025
https://scholar.google.com/scholar?q=Training+Plug-n-Play+Knowledge+Modules+with+Deep+Context+Distillation
7. Activated LoRA: Fine-tuned LLMs for Intrinsics — Kristjan Greenewald et al., 2025
https://scholar.google.com/scholar?q=Activated+LoRA%3A+Fine-tuned+LLMs+for+Intrinsics
8. LoRA-Augmented Generation (LAG) for Knowledge-Intensive Language Tasks — William Fleshman and Benjamin Van Durme, 2025
https://scholar.google.com/scholar?q=LoRA-Augmented+Generation+%28LAG%29+for+Knowledge-Intensive+Language+Tasks
9. Doc-to-LoRA: Learning to Instantly Internalize Contexts — Rujikorn Charakorn et al., 2026
https://scholar.google.com/scholar?q=Doc-to-LoRA%3A+Learning+to+Instantly+Internalize+Contexts
10. Decoupling Knowledge and Task Subspaces for Composable Parametric Retrieval Augmented Generation — Weihang Su et al., 2026
https://scholar.google.com/scholar?q=Decoupling+Knowledge+and+Task+Subspaces+for+Composable+Parametric+Retrieval+Augmented+Generation
11. KeDiff: Key Similarity-Based KV Cache Eviction for Long-Context LLM Inference in Resource-Constrained Environments — Junyoung Park et al., 2025
https://scholar.google.com/scholar?q=KeDiff%3A+Key+Similarity-Based+KV+Cache+Eviction+for+Long-Context+LLM+Inference+in+Resource-Constrained+Environments
12. Model Tells You Where to Merge: Adaptive KV Cache Merging for LLMs on Long-Context Tasks — Zheng Wang et al., 2024
https://scholar.google.com/scholar?q=Model+Tells+You+Where+to+Merge%3A+Adaptive+KV+Cache+Merging+for+LLMs+on+Long-Context+Tasks
13. Parametric Retrieval-Augmented Generation using Latent Routing of LoRA Adapters — Zhan Su, Fengran Mo, Jian-yun Nie, 2025
https://scholar.google.com/scholar?q=Parametric+Retrieval-Augmented+Generation+using+Latent+Routing+of+LoRA+Adapters
14. One Token Can Help! Learning Scalable and Pluggable Virtual Tokens for Retrieval-Augmented Large Language Models — Yutao Zhu et al., 2024
https://scholar.google.com/scholar?q=One+Token+Can+Help%21+Learning+Scalable+and+Pluggable+Virtual+Tokens+for+Retrieval-Augmented+Large+Language+Models
15. AI Post Transformers: Doc-to-LoRA: Internalizing Context as LoRA — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-29-doc-to-lora-internalizing-context-as-lor-8dd5ec.mp3
16. AI Post Transformers: KVzip for Query-Agnostic KV Cache Compression — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-29-kvzip-for-query-agnostic-kv-cache-compre-72afe5.mp3
17. AI Post Transformers: Efficient KV Cache Sharing for Multi-LoRA Agents — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-22-efficient-kv-cache-sharing-for-multi-lor-afda05.mp3
18. AI Post Transformers: Lookahead Q-Cache for Consistent KV Eviction — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-25-lookahead-q-cache-for-consistent-kv-evic-d97b09.mp3
19. AI Post Transformers: KVzap: Fast, Adaptive, Faithful KV Cache Pruning — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-30-kvzap-fast-adaptive-faithful-kv-cache-pr-dbe515.mp3
20. AI Post Transformers: Explicit Information Transmission for Context Compression — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-explicit-information-transmission-for-co-24e3c2.mp3

This episode explores IndexMem, a long-context LLM inference method that tries to cut KV-cache memory by learning which token states to evict while preserving useful information in a fixed-size latent memory. It explains the mechanics behind the KV cache, prefill, and decoding, then frames the real systems problem: for long prompts, memory traffic and bandwidth can become a bigger bottleneck than raw compute. The discussion focuses on two distinct challenges the paper separates clearly: predicting which cached tokens will matter in the future, and avoiding irreversible forgetting after eviction by writing evicted information into a learned summary state. Listeners interested in code agents, multimodal pipelines, and long-context serving will find it useful because it connects transformer theory to practical deployment constraints while questioning whether current evidence really supports the paper’s bigger million-token ambitions.

Sources:
1. IndexMem: Learned KV-Cache Eviction with Latent Memory for Long-Context LLM Inference — Xintong Yang, Hao Gu, Binxing Xu, Lujun Li, Bei Liu, Jiacheng Liu, Qiyuan Zhu, Sirui Han, Yike Guo, 2026
http://arxiv.org/abs/2605.25475
2. Expected Attention: KV Cache Compression by Estimating Attention from Future Queries Distribution — Alessio Devoto, Maximilian Jeblick, Simon Jegou, 2025
https://scholar.google.com/scholar?q=Expected+Attention%3A+KV+Cache+Compression+by+Estimating+Attention+from+Future+Queries+Distribution
3. Lookahead Q-Cache: Achieving More Consistent KV Cache Eviction via Pseudo Query — Yixuan Wang, Shiyu Ji, Yijun Liu, Yuzhuang Xu, Yang Xu, Qingfu Zhu, Wanxiang Che, 2025
https://scholar.google.com/scholar?q=Lookahead+Q-Cache%3A+Achieving+More+Consistent+KV+Cache+Eviction+via+Pseudo+Query
4. Locret: Enhancing Eviction in Long-Context LLM Inference with Trained Retaining Heads on Consumer-Grade Devices — Yuxiang Huang, Binhang Yuan, Xu Han, Chaojun Xiao, Zhiyuan Liu, 2024
https://scholar.google.com/scholar?q=Locret%3A+Enhancing+Eviction+in+Long-Context+LLM+Inference+with+Trained+Retaining+Heads+on+Consumer-Grade+Devices
5. KVReviver: Reversible KV Cache Compression with Sketch-Based Token Reconstruction — Aomufei Yuan, Zhiming Wang, Ruijie Miao, Dayu Wang, Yuxuan Tian, Zihan Wang, Yebo Peng, Yuhan Wu, Bairen Yi, Xin Liu, Tong Yang, 2025
https://scholar.google.com/scholar?q=KVReviver%3A+Reversible+KV+Cache+Compression+with+Sketch-Based+Token+Reconstruction
6. xKV: Cross-Layer SVD for KV-Cache Compression — Chi-Chih Chang, Chien-Yu Lin, Yash Akhauri, Wei-Cheng Lin, Kai-Chiang Wu, Luis Ceze, Mohamed S. Abdelfattah, 2025
https://scholar.google.com/scholar?q=xKV%3A+Cross-Layer+SVD+for+KV-Cache+Compression
7. IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse — Yushi Bai, Qian Dong, Ting Jiang, Xin Lv, Zhengxiao Du, Aohan Zeng, Jie Tang, Juanzi Li, 2026
https://scholar.google.com/scholar?q=IndexCache%3A+Accelerating+Sparse+Attention+via+Cross-Layer+Index+Reuse
8. Titans: Learning to Memorize at Test Time — Ali Behrouz, Peilin Zhong, Vahab Mirrokni, 2024
https://scholar.google.com/scholar?q=Titans%3A+Learning+to+Memorize+at+Test+Time
9. FreeKV: Boosting KV Cache Retrieval for Efficient LLM Inference — Guangda Liu et al., 2025
https://arxiv.org/abs/2505.13109
10. Streaming Video Question-Answering with In-context Video KV-Cache Retrieval — Shangzhe Di et al., 2025
https://arxiv.org/abs/2503.00540
11. LLMs Know What to Drop: Self-Attention Guided KV Cache Eviction for Efficient Long-Context Inference — Guangtao Wang et al., 2025
https://arxiv.org/abs/2503.08879
12. Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference — Yuan Feng et al., 2024
https://arxiv.org/abs/2407.11550
13. In-context KV-Cache Eviction for LLMs via Attention-Gate — Zihao Zeng et al., 2024
https://arxiv.org/abs/2410.12876
14. MomentKV: Closing the Directional Gap in KV Cache Eviction for Long-Context Inference — Yu Li et al., 2026
https://arxiv.org/abs/2606.01563
15. ChunkKV: Semantic-Preserving KV Cache Compression for Efficient Long-Context LLM Inference — Xiang Liu et al., 2025
https://arxiv.org/abs/2502.00299
16. MEDA: Dynamic KV Cache Allocation for Efficient Multimodal Long-Context Inference — Zhongwei Wan et al., 2025
https://arxiv.org/abs/2502.17599
17. AI Post Transformers: Adaptive Compression Techniques for Efficient LLM Inference — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/adaptive-compression-techniques-for-efficient-llm-inference/
18. AI Post Transformers: Explicit Information Transmission for Context Compression — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-explicit-information-transmission-for-co-24e3c2.mp3
19. AI Post Transformers: TRELLIS and Bounded-Memory Transformer KV Compression — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-02-trellis-and-bounded-memory-transformer-k-81f237.mp3
20. AI Post Transformers: End-to-End Context Compression at Scale — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-10-end-to-end-context-compression-at-scale-278c70.mp3
21. AI Post Transformers: 50x KV Cache Compression in Seconds via Attention Matching — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/50x-kv-cache-compression-in-seconds-via-attention-matching/
22. AI Post Transformers: Stochastic KV Routing for Cache Sharing — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-29-stochastic-kv-routing-for-cache-sharing-5fef63.mp3
Interactive Visualization: IndexMem: Learned KV-Cache Eviction for Long-Context LLMs

This episode explores AllMem, a method for turning pretrained Qwen3 models into long-context systems that keep exact attention over a recent token window while storing older context in a learned memory. It explains why standard transformer attention becomes prohibitively expensive on long chats, books, codebases, and agent traces, and places AllMem in the broader landscape of sliding-window, sparse-attention, recurrent, and memory-augmented architectures. The discussion highlights the paper’s core argument: a hybrid design can preserve sharp local reasoning, compress the distant past through online memory updates, and approach the quality of full attention without the same compute and KV-cache costs. A listener would find it interesting because it connects concrete systems constraints on phones and servers to a specific recipe for making long-context language models more practical.

Sources:
1. AllMem: A Memory-centric Recipe for Efficient Long-context Modeling — Ziming Wang, Xiang Wang, Kailong Peng, Lang Qin, Juan Gabriel Kostelec, Christos Sourmpis, Axel Laborieux, Qinghai Guo, 2026
http://arxiv.org/abs/2602.13680
2. Longformer: The Long-Document Transformer — Iz Beltagy, Matthew E. Peters, Arman Cohan, 2020
https://scholar.google.com/scholar?q=Longformer%3A+The+Long-Document+Transformer
3. Big Bird: Transformers for Longer Sequences — Manzil Zaheer, Guru Guruganesh, Avinava Dubey, Joshua Ainslie, Amr Ahmed, et al., 2020
https://scholar.google.com/scholar?q=Big+Bird%3A+Transformers+for+Longer+Sequences
4. Mistral 7B — Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Guillaume Lample, et al., 2023
https://scholar.google.com/scholar?q=Mistral+7B
5. Efficient Streaming Language Models with Attention Sinks — Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, Mike Lewis, 2023
https://scholar.google.com/scholar?q=Efficient+Streaming+Language+Models+with+Attention+Sinks
6. Titans: Learning to Memorize at Test Time — Ali Behrouz, Peilin Zhong, Vahab Mirrokni, 2024
https://scholar.google.com/scholar?q=Titans%3A+Learning+to+Memorize+at+Test+Time
7. Artificial Hippocampus Networks for Efficient Long-Context Modeling — Yunhao Fang, Weihao Yu, Shu Zhong, Qinghao Ye, Xuehan Xiong, Lai Wei, 2025
https://scholar.google.com/scholar?q=Artificial+Hippocampus+Networks+for+Efficient+Long-Context+Modeling
8. The Mamba in the Llama: Distilling and Accelerating Hybrid Models — Junxiong Wang, Daniele Paliotta, Avner May, Alexander M. Rush, Tri Dao, 2024
https://scholar.google.com/scholar?q=The+Mamba+in+the+Llama%3A+Distilling+and+Accelerating+Hybrid+Models
9. MesaNet: Sequence Modeling by Locally Optimal Test-Time Training — Johannes von Oswald, Nino Scherrer, Seijin Kobayashi, Luca Versari, Songlin Yang, Maximilian Schlegel, Kaitlin Maile, Yanick Schimpf, Oliver Sieberling, Alexander Meulemans, Rif A. Saurous, Guillaume Lajoie, Charlotte Frenkel, Razvan Pascanu, Blaise Agüera y Arcas, João Sacramento, 2025
https://scholar.google.com/scholar?q=MesaNet%3A+Sequence+Modeling+by+Locally+Optimal+Test-Time+Training
10. RULER: What's the Real Context Size of Your Long-Context Language Models? — Cheng-Ping Hsieh, Simeng Sun, Samuel Kriman, Shantanu Acharya, Dima Rekesh, Fei Jia, Yang Zhang, Boris Ginsburg, 2024
https://scholar.google.com/scholar?q=RULER%3A+What%27s+the+Real+Context+Size+of+Your+Long-Context+Language+Models%3F
11. RazorAttention: Efficient KV Cache Compression Through Retrieval Heads — Hanlin Tang et al., 2024
https://scholar.google.com/scholar?q=RazorAttention%3A+Efficient+KV+Cache+Compression+Through+Retrieval+Heads
12. Not All Heads Matter: A Head-Level KV Cache Compression Method with Integrated Retrieval and Reasoning — Yu Fu et al., 2024
https://scholar.google.com/scholar?q=Not+All+Heads+Matter%3A+A+Head-Level+KV+Cache+Compression+Method+with+Integrated+Retrieval+and+Reasoning
13. Sliding Window Attention Adaptation — Yijiong Yu et al., 2025
https://scholar.google.com/scholar?q=Sliding+Window+Attention+Adaptation
14. HyperAttention: Long-context Attention in Near-Linear Time — Insu Han et al., 2023
https://scholar.google.com/scholar?q=HyperAttention%3A+Long-context+Attention+in+Near-Linear+Time
15. On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention — Yeonju Ro et al., 2025
https://scholar.google.com/scholar?q=On-the-Fly+Adaptive+Distillation+of+Transformer+to+Dual-State+Linear+Attention
16. Test-Time Training on Nearest Neighbors for Large Language Models — Moritz Hardt and Yu Sun, 2023
https://scholar.google.com/scholar?q=Test-Time+Training+on+Nearest+Neighbors+for+Large+Language+Models
17. Test-Time Learning for Large Language Models — Jinwu Hu et al., 2025
https://scholar.google.com/scholar?q=Test-Time+Learning+for+Large+Language+Models
18. Training Large Reasoning Models Efficiently via Progressive Thought Encoding — Zeliang Zhang et al., 2026
https://scholar.google.com/scholar?q=Training+Large+Reasoning+Models+Efficiently+via+Progressive+Thought+Encoding
19. AI Post Transformers: δ-mem and Online Memory for LLMs — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-13-d-mem-and-online-memory-for-llms-6622fa.mp3
20. AI Post Transformers: MELT: Decoupling Compute From Memory — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-13-melt-decoupling-compute-from-memory-26430c.mp3
21. AI Post Transformers: Kimi Linear: Efficient Expressive Attention Architecture — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/kimi-linear-efficient-expressive-attention-architecture/
22. AI Post Transformers: Ministral 3: Cascade Distillation for Long-Context Multimodal Models — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-cascade-distillation-for-long-context-mu-0ebd1a.mp3
23. AI Post Transformers: Doc-to-LoRA: Internalizing Context as LoRA — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-29-doc-to-lora-internalizing-context-as-lor-8dd5ec.mp3
24. AI Post Transformers: Memory-Bound, Not Bandwidth-Limited Batch-1 LLM Decode — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-02-memory-bound-not-bandwidth-limited-batch-114799.mp3
25. AI Post Transformers: Do Transformers Need Three Projections? — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-11-do-transformers-need-three-projections-c227d6.mp3
Interactive Visualization: AllMem for Efficient Long-Context Modeling

This episode explores the Relational Graph Transformer paper and asks whether transformer-based models can outperform standard graph neural networks for prediction tasks over real multi-table databases such as customers, orders, products, claims, and shipments. It explains how the method turns a relational warehouse into a heterogeneous temporal graph, then builds five-part tokens for sampled neighbors that encode row features, table type, hop distance, relative time, and a learned local-structure signal. The discussion focuses on the model’s local-global attention design, where dense attention over timestamp-safe two-hop neighborhoods is paired with learned global centroids to capture broader database patterns without full all-pairs cost. It is especially interesting because it frames both the promise and the friction of relational deep learning: strong motivation to beat hand-engineered SQL features and message-passing bottlenecks, but real skepticism about whether such graph-heavy systems are practical enough for ordinary industrial stacks.

Sources:
1. Relational Graph Transformer — Vijay Prakash Dwivedi, Sri Jaladi, Yangyi Shen, Federico López, Charilaos I. Kanatsoulis, Rishi Puri, Matthias Fey, Jure Leskovec, 2025
http://arxiv.org/abs/2505.10960
2. Heterogeneous Graph Transformer — Ziniu Hu, Yuxiao Dong, Kuansan Wang, Yizhou Sun, 2020
https://scholar.google.com/scholar?q=Heterogeneous+Graph+Transformer
3. Temporal Graph Networks for Deep Learning on Dynamic Graphs — Emanuele Rossi, Ben Chamberlain, Fabrizio Frasca, Davide Eynard, Federico Monti, Michael Bronstein, 2020
https://scholar.google.com/scholar?q=Temporal+Graph+Networks+for+Deep+Learning+on+Dynamic+Graphs
4. Relational Deep Learning: Graph Representation Learning on Relational Databases — Matthias Fey, Weihua Hu, Kexin Huang, Jan Eric Lenssen, Rishabh Ranjan, Joshua Robinson, Rex Ying, Jiaxuan You, Jure Leskovec, 2023
https://scholar.google.com/scholar?q=Relational+Deep+Learning%3A+Graph+Representation+Learning+on+Relational+Databases
5. RelBench: A Benchmark for Deep Learning on Relational Databases — Joshua Robinson, Rishabh Ranjan, Weihua Hu, Matthias Fey, Jure Leskovec, et al., 2024
https://scholar.google.com/scholar?q=RelBench%3A+A+Benchmark+for+Deep+Learning+on+Relational+Databases
6. Graph-Bert: Only Attention is Needed for Learning Graph Representations — Jiawei Zhang, Haopeng Zhang, Congying Xia, Li Sun, 2020
https://scholar.google.com/scholar?q=Graph-Bert%3A+Only+Attention+is+Needed+for+Learning+Graph+Representations
7. Do Transformers Really Perform Bad for Graph Representation? — Chengxuan Ying, Tianle Cai, Shengjie Luo, Shuxin Zheng, Guolin Ke, Di He, Yanming Shen, Tie-Yan Liu, 2021
https://scholar.google.com/scholar?q=Do+Transformers+Really+Perform+Bad+for+Graph+Representation%3F
8. Pure Transformers are Powerful Graph Learners — Jinwoo Kim, Tien Dat Nguyen, Seonwoo Min, Sungjun Cho, Moontae Lee, Honglak Lee, Seunghoon Hong, 2022
https://scholar.google.com/scholar?q=Pure+Transformers+are+Powerful+Graph+Learners
9. NAGphormer: A Tokenized Graph Transformer for Node Classification in Large Graphs — Jinsong Chen, Kaiyuan Gao, Gaichao Li, Kun He, 2023
https://scholar.google.com/scholar?q=NAGphormer%3A+A+Tokenized+Graph+Transformer+for+Node+Classification+in+Large+Graphs
10. Representing Long-Range Context for Graph Neural Networks with Global Attention — Zhanghao Wu, Paras Jain, Matthew A. Wright, Azalia Mirhoseini, Joseph E. Gonzalez, Ion Stoica, 2021
https://scholar.google.com/scholar?q=Representing+Long-Range+Context+for+Graph+Neural+Networks+with+Global+Attention
11. Recipe for a General, Powerful, Scalable Graph Transformer — Ladislav Rampášek, Mikhail Galkin, Vijay Prakash Dwivedi, Anh Tuan Luu, Guy Wolf, Dominique Beaini, 2022
https://scholar.google.com/scholar?q=Recipe+for+a+General%2C+Powerful%2C+Scalable+Graph+Transformer
12. Exphormer: Sparse Transformers for Graphs — Hamed Shirzad, Ameya Velingker, Balaji Venkatachalam, Danica J. Sutherland, Ali Kemal Sinop, 2023
https://scholar.google.com/scholar?q=Exphormer%3A+Sparse+Transformers+for+Graphs
13. Centroid Transformers: Learning to Abstract with Attention — Lemeng Wu, Xingchao Liu, Qiang Liu, 2021
https://scholar.google.com/scholar?q=Centroid+Transformers%3A+Learning+to+Abstract+with+Attention
14. Learning Efficient Positional Encodings with Graph Neural Networks — Charilaos I. Kanatsoulis et al., 2025
https://arxiv.org/abs/2502.01122
15. ContextGNN: Beyond Two-Tower Recommendation Systems — Yiwen Yuan et al., 2024
https://arxiv.org/abs/2411.19513
16. RelGNN: Composite Message Passing for Relational Deep Learning — Tianlang Chen, Charilaos Kanatsoulis, Jure Leskovec, 2025
https://arxiv.org/abs/2502.06784
17. Are Graph Transformers Necessary? Efficient Long-Range Message Passing with Fractal Nodes in MPNNs — Jeongwhan Choi et al., 2025
https://arxiv.org/abs/2511.13010
18. Beyond Message Passing: Neural Graph Pattern Machine — Zehong Wang et al., 2025
https://arxiv.org/abs/2501.18739
19. Transformers Meet Relational Databases — Jakub Peleska and Gustav Sir, 2024
https://arxiv.org/abs/2412.05218
20. Relational Transformer: Toward Zero-Shot Foundation Models for Relational Data — Rishabh Ranjan et al., 2025
https://arxiv.org/abs/2510.06377
21. Tokenphormer: Structure-aware Multi-token Graph Transformer for Node Classification — Zijie Zhou et al., 2024
https://arxiv.org/abs/2412.15302
22. AI Post Transformers: KumoRFM for In-Context Relational Learning — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-11-kumorfm-for-in-context-relational-learni-520d2b.mp3

This episode explores Lattice, a 2025 paper from Google Research and Google DeepMind that asks whether a Transformer’s growing key-value cache can be compressed into a fixed set of memory slots without losing the long-context behavior users care about. It explains why this matters by contrasting standard attention’s unbounded cache with linear attention, recurrent state models, and fast-weight associative memory, framing the problem as memory compression rather than a rejection of Transformers. The discussion focuses on Lattice’s core idea: treat memory as an online low-rank factorization, reconstruct each new token from the current slots, and write only the residual through a single gradient-style update whose gate and direction arise from the math. Listeners would find it interesting because it gets into the real tradeoff between elegant compression and practical accuracy, including whether learned fixed-slot memory can beat simpler industry tactics like quantizing, sharding, or evicting cache entries.

Sources:
1. Lattice: Learning to Efficiently Compress the Memory — Mahdi Karami, Razvan Pascanu, Vahab Mirrokni, 2025
http://arxiv.org/abs/2504.05646
2. Compressive Transformers for Long-Range Sequence Modelling — Jack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, Timothy P. Lillicrap, 2019
https://scholar.google.com/scholar?q=Compressive+Transformers+for+Long-Range+Sequence+Modelling
3. Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention — Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, Francois Fleuret, 2020
https://scholar.google.com/scholar?q=Transformers+are+RNNs%3A+Fast+Autoregressive+Transformers+with+Linear+Attention
4. Palu: Compressing KV-Cache with Low-Rank Projection — Chi-Chih Chang, Wei-Cheng Lin, Chien-Yu Lin, Chong-Yan Chen, Yu-Fang Hu, Pei-Shuo Wang, Ning-Chi Huang, Luis Ceze, Mohamed S. Abdelfattah, Kai-Chiang Wu, 2024
https://scholar.google.com/scholar?q=Palu%3A+Compressing+KV-Cache+with+Low-Rank+Projection
5. Lattice: Learning to Efficiently Compress the Memory — Mahdi Karami, Razvan Pascanu, Vahab Mirrokni, 2025
https://scholar.google.com/scholar?q=Lattice%3A+Learning+to+Efficiently+Compress+the+Memory
6. Linear Transformers Are Secretly Fast Weight Programmers — Imanol Schlag, Kazuki Irie, Jurgen Schmidhuber, 2021
https://scholar.google.com/scholar?q=Linear+Transformers+Are+Secretly+Fast+Weight+Programmers
7. Gated Delta Networks: Improving Mamba2 with Delta Rule — Songlin Yang, Jan Kautz, Ali Hatamizadeh, 2024
https://scholar.google.com/scholar?q=Gated+Delta+Networks%3A+Improving+Mamba2+with+Delta+Rule
8. Kimi Linear: An Expressive, Efficient Attention Architecture — Kimi Team; Yu Zhang, Zongyu Lin, Xingcheng Yao, Jiaxi Hu, Fanqing Meng, et al., 2025
https://scholar.google.com/scholar?q=Kimi+Linear%3A+An+Expressive%2C+Efficient+Attention+Architecture
9. Parallelizing Linear Transformers with the Delta Rule over Sequence Length — Songlin Yang, Bailin Wang, Yu Zhang, Yikang Shen, Yoon Kim, 2024
https://scholar.google.com/scholar?q=Parallelizing+Linear+Transformers+with+the+Delta+Rule+over+Sequence+Length
10. Learning to (Learn at Test Time): RNNs with Expressive Hidden States — Yu Sun, Xinhao Li, Karan Dalal, Jiarui Xu, Arjun Vikram, Genghan Zhang, Yann Dubois, Xinlei Chen, Xiaolong Wang, Sanmi Koyejo, et al., 2024
https://scholar.google.com/scholar?q=Learning+to+%28Learn+at+Test+Time%29%3A+RNNs+with+Expressive+Hidden+States
11. Simple Linear Attention Language Models Balance the Recall-Throughput Tradeoff — Simran Arora, Sabri Eyuboglu, Michael Zhang, Aman Timalsina, Silas Alberti, Dylan Zinsley, James Zou, Atri Rudra, Christopher Re, 2024
https://scholar.google.com/scholar?q=Simple+Linear+Attention+Language+Models+Balance+the+Recall-Throughput+Tradeoff
12. Test-time Regression: a Unifying Framework for Designing Sequence Models with Associative Memory — Ke Alexander Wang, Jiaxin Shi, Emily B. Fox, 2025
https://scholar.google.com/scholar?q=Test-time+Regression%3A+a+Unifying+Framework+for+Designing+Sequence+Models+with+Associative+Memory
13. KVLink: Accelerating Large Language Models via Efficient KV Cache Reuse — Jingbo Yang et al., 2025
https://arxiv.org/abs/2502.16002
14. LMCache: An Efficient KV Cache Layer for Enterprise-Scale LLM Inference — Yihua Cheng et al., 2025
https://arxiv.org/abs/2510.09665
15. No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization — June Yong Yang et al., 2024
https://arxiv.org/abs/2402.18096
16. Attention Score is not All You Need for Token Importance Indicator in KV Cache Reduction: Value Also Matters — Zhiyu Guo et al., 2024
https://arxiv.org/abs/2406.12335
17. ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification — Yefei He et al., 2024
https://arxiv.org/abs/2405.14256
18. State-space Models can Learn In-Context by Gradient Descent — Neeraj Mohan Sushma et al., 2024
https://arxiv.org/abs/2410.11687
19. Test-Time Training Done Right — Tianyuan Zhang et al., 2025
https://arxiv.org/abs/2505.23884
20. Titans: Learning to Memorize at Test Time — Ali Behrouz, Peilin Zhong, Vahab Mirrokni, 2025
https://arxiv.org/abs/2501.00663
21. AI Post Transformers: TRELLIS and Bounded-Memory Transformer KV Compression — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-02-trellis-and-bounded-memory-transformer-k-81f237.mp3
22. AI Post Transformers: Parallelizing DeltaNet Linear Transformers over Sequence Length — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-18-parallelizing-deltanet-linear-transforme-2d0377.mp3
23. AI Post Transformers: Gated Delta Networks for Long-Context Retrieval — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-17-gated-delta-networks-for-long-context-re-706d85.mp3
24. AI Post Transformers: Gated Linear Attention for Efficient Long Sequences — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-18-gated-linear-attention-for-efficient-lon-c858ab.mp3
25. AI Post Transformers: Explicit Information Transmission for Context Compression — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-explicit-information-transmission-for-co-24e3c2.mp3
26. AI Post Transformers: Titans: Learning to Memorize at Test Time — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-20-titans-learning-to-memorize-at-test-time-054662.mp3
27. AI Post Transformers: Long Context Pre-Training with Lighthouse Attention — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-13-long-context-pre-training-with-lighthous-e85bbe.mp3
28. AI Post Transformers: Memory-Bound, Not Bandwidth-Limited Batch-1 LLM Decode — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-02-memory-bound-not-bandwidth-limited-batch-114799.mp3
Interactive Visualization: Lattice: Fixed-Slot Compression for Transformer Memory

This episode explores KumoRFM, a 2025 proposal for a foundation model that can perform in-context learning directly on relational databases, aiming to handle tasks like churn prediction, fraud detection, recommendation, and forecasting without training a separate model for each schema and label. It explains how the approach represents warehouse data as heterogeneous graphs of rows and foreign-key relationships, using attention over local relational neighborhoods instead of flattening everything into handcrafted feature tables. The discussion focuses on the paper’s strongest claim, zero-shot transfer, and carefully separates true inference-time generalization from easier settings like continued pretraining on the target database or later fine-tuning on the target task. Listeners would find it interesting because the episode gets precise about what this system could change in enterprise ML, while also surfacing the practical caveats around task specification, temporal leakage, infrastructure cost, and how literal the any database, any task promise really is.

Sources:
1. KumoRFM for In-Context Relational Learning
https://kumo.ai/research/kumo_relational_foundation_model.pdf
2. Modeling Relational Data with Graph Convolutional Networks — Michael Schlichtkrull, Thomas N. Kipf, Peter Bloem, Max Welling, et al., 2017
https://scholar.google.com/scholar?q=Modeling+Relational+Data+with+Graph+Convolutional+Networks
3. Heterogeneous Graph Transformer — Ziniu Hu, Yuxiao Dong, Kuansan Wang, Yizhou Sun, 2020
https://scholar.google.com/scholar?q=Heterogeneous+Graph+Transformer
4. Relational Deep Learning: Graph Representation Learning on Relational Databases — Matthias Fey, Weihua Hu, Kexin Huang, Jan Eric Lenssen, Jure Leskovec, et al., 2023
https://scholar.google.com/scholar?q=Relational+Deep+Learning%3A+Graph+Representation+Learning+on+Relational+Databases
5. Relational Graph Transformer — Vijay Prakash Dwivedi, Sri Jaladi, Yangyi Shen, Federico Lopez, Matthias Fey, Jure Leskovec, et al., 2025 (ICLR 2026)
https://scholar.google.com/scholar?q=Relational+Graph+Transformer
6. One Model to Rule them All: Towards Zero-Shot Learning for Databases — Benjamin Hilprecht, Carsten Binnig, 2021
https://scholar.google.com/scholar?q=One+Model+to+Rule+them+All%3A+Towards+Zero-Shot+Learning+for+Databases
7. Zero-Shot Cost Models for Out-of-the-box Learned Cost Prediction — Benjamin Hilprecht, Carsten Binnig, 2022
https://scholar.google.com/scholar?q=Zero-Shot+Cost+Models+for+Out-of-the-box+Learned+Cost+Prediction
8. TabPFN: A Transformer That Solves Small Tabular Classification Problems in a Second — Noah Hollmann, Samuel Muller, Katharina Eggensperger, Frank Hutter, 2022 (ICLR 2023)
https://scholar.google.com/scholar?q=TabPFN%3A+A+Transformer+That+Solves+Small+Tabular+Classification+Problems+in+a+Second
9. Relational Transformer: Toward Zero-Shot Foundation Models for Relational Data — Rishabh Ranjan, Valter Hudovernik, Mark Znidar, Carlos Guestrin, Jure Leskovec, et al., 2025 (ICLR 2026)
https://scholar.google.com/scholar?q=Relational+Transformer%3A+Toward+Zero-Shot+Foundation+Models+for+Relational+Data
10. Relational In-Context Learning via Synthetic Pre-training with Structural Prior — Yanbo Wang, Jiaxuan You, Chuan Shi, Muhan Zhang, 2026
https://scholar.google.com/scholar?q=Relational+In-Context+Learning+via+Synthetic+Pre-training+with+Structural+Prior
11. OpenRFM: Dissecting Relational In-Context Learning — Zhikai Chen et al., 2026
https://scholar.google.com/scholar?q=OpenRFM%3A+Dissecting+Relational+In-Context+Learning
12. KumoRFM-2: Scaling Foundation Models for Relational Learning — Valter Hudovernik et al., 2026
https://scholar.google.com/scholar?q=KumoRFM-2%3A+Scaling+Foundation+Models+for+Relational+Learning
13. Retrieval & Fine-Tuning for In-Context Tabular Models — Valentin Thomas et al., 2024
https://scholar.google.com/scholar?q=Retrieval+%26+Fine-Tuning+for+In-Context+Tabular+Models
14. Scalable In-Context Learning on Tabular Data via Retrieval-Augmented Large Language Models — Xumeng Wen et al., 2025
https://scholar.google.com/scholar?q=Scalable+In-Context+Learning+on+Tabular+Data+via+Retrieval-Augmented+Large+Language+Models
15. Exploring Fine-Tuning for Tabular Foundation Models — Aditya Tanna et al., 2026
https://scholar.google.com/scholar?q=Exploring+Fine-Tuning+for+Tabular+Foundation+Models
16. AI Post Transformers: Predictive Query Language for Relational Databases — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-08-predictive-query-language-for-relational-103e68.mp3
17. AI Post Transformers: KumoRFM-2 for Relational Learning at Scale — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-08-kumorfm-2-for-relational-learning-at-sca-13c996.mp3
18. AI Post Transformers: Why LightGBM Made Boosted Trees Fast — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-why-lightgbm-made-boosted-trees-fast-286a89.mp3
19. AI Post Transformers: How Induction Heads Emerge in Transformers — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-03-how-induction-heads-emerge-in-transforme-a7bfcb.mp3

This episode explores Atlas, a 2025 paper on test-time memorization that asks whether a model with fixed recurrent memory can learn to update that memory during inference and rival Transformers on long-context recall and reasoning. It explains the core tradeoff between Transformer-style KV caches, which preserve near-exact token access at growing cost, and bounded recurrent memory, which must decide what to keep, compress, or forget. The discussion focuses on why earlier recurrent memory systems fell short, then breaks down Atlas's proposed fixes: evaluating memory updates against a window of recent tokens rather than only the newest token, using richer key representations, and learning stronger retention and optimizer-style write rules. Listeners get a clear view of why this matters for post-Transformer architectures, and why fixed-size memory remains both a promising direction and a stubborn bottleneck.

Sources:
1. ATLAS: Learning to Optimally Memorize the Context at Test Time — Ali Behrouz, Zeman Li, Praneeth Kacham, Majid Daliri, Yuan Deng, Peilin Zhong, Meisam Razaviyayn, Vahab Mirrokni, 2025
http://arxiv.org/abs/2505.23735
2. Linear Transformers Are Secretly Fast Weight Programmers — Imanol Schlag, Kazuki Irie, Jürgen Schmidhuber, 2021
https://scholar.google.com/scholar?q=Linear+Transformers+Are+Secretly+Fast+Weight+Programmers
3. Retentive Network: A Successor to Transformer for Large Language Models — Yutao Sun, Li Dong, Shaohan Huang, Furu Wei, et al., 2023
https://scholar.google.com/scholar?q=Retentive+Network%3A+A+Successor+to+Transformer+for+Large+Language+Models
4. Titans: Learning to Memorize at Test Time — Ali Behrouz, Peilin Zhong, Vahab Mirrokni, 2025
https://scholar.google.com/scholar?q=Titans%3A+Learning+to+Memorize+at+Test+Time
5. It's All Connected: A Journey Through Test-Time Memorization, Attentional Bias, Retention, and Online Optimization — Ali Behrouz, Meisam Razaviyayn, Peilin Zhong, Vahab Mirrokni, 2025
https://scholar.google.com/scholar?q=It%27s+All+Connected%3A+A+Journey+Through+Test-Time+Memorization%2C+Attentional+Bias%2C+Retention%2C+and+Online+Optimization
6. Learning to (Learn at Test Time): RNNs with Expressive Hidden States — Yu Sun et al., 2024
https://scholar.google.com/scholar?q=Learning+to+%28Learn+at+Test+Time%29%3A+RNNs+with+Expressive+Hidden+States
7. RNNs are not Transformers (Yet): The Key Bottleneck on In-Context Retrieval — Kaiyue Wen, Xingyu Dang, Kaifeng Lyu, 2024
https://scholar.google.com/scholar?q=RNNs+are+not+Transformers+%28Yet%29%3A+The+Key+Bottleneck+on+In-Context+Retrieval
8. Test-time Regression: a Unifying Framework for Designing Sequence Models with Associative Memory — Ke Alexander Wang, Jiaxin Shi, Emily B. Fox, 2025
https://scholar.google.com/scholar?q=Test-time+Regression%3A+a+Unifying+Framework+for+Designing+Sequence+Models+with+Associative+Memory
9. BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack — Yuri Kuratov et al., 2024
https://scholar.google.com/scholar?q=BABILong%3A+Testing+the+Limits+of+LLMs+with+Long+Context+Reasoning-in-a-Haystack
10. SCBench: A KV Cache-Centric Analysis of Long-Context Methods — Yucheng Li et al., 2024
https://scholar.google.com/scholar?q=SCBench%3A+A+KV+Cache-Centric+Analysis+of+Long-Context+Methods
11. Model Tells You Where to Merge: Adaptive KV Cache Merging for LLMs on Long-Context Tasks — Zheng Wang et al., 2024
https://scholar.google.com/scholar?q=Model+Tells+You+Where+to+Merge%3A+Adaptive+KV+Cache+Merging+for+LLMs+on+Long-Context+Tasks
12. Forgetting Transformer: Softmax Attention with a Forget Gate — Zhixuan Lin et al., 2025
https://scholar.google.com/scholar?q=Forgetting+Transformer%3A+Softmax+Attention+with+a+Forget+Gate
13. Test-Time Training Done Right — Tianyuan Zhang et al., 2025
https://scholar.google.com/scholar?q=Test-Time+Training+Done+Right
14. Associative Recurrent Memory Transformer — Ivan Rodkin et al., 2024
https://scholar.google.com/scholar?q=Associative+Recurrent+Memory+Transformer
15. AI Post Transformers: Titans: Learning to Memorize at Test Time — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-20-titans-learning-to-memorize-at-test-time-054662.mp3
16. AI Post Transformers: Gated Delta Networks for Long-Context Retrieval — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-17-gated-delta-networks-for-long-context-re-706d85.mp3
17. AI Post Transformers: Parallelizing DeltaNet Linear Transformers over Sequence Length — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-18-parallelizing-deltanet-linear-transforme-2d0377.mp3
18. AI Post Transformers: Gated Linear Attention for Efficient Long Sequences — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-18-gated-linear-attention-for-efficient-lon-c858ab.mp3
19. AI Post Transformers: Muon Is Scalable for LLM Training — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-25-muon-is-scalable-for-llm-training-587ed8.mp3
20. AI Post Transformers: How Induction Heads Emerge in Transformers — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-03-how-induction-heads-emerge-in-transforme-a7bfcb.mp3
Interactive Visualization: Atlas: Test-Time Memory for Long Contexts

This episode explores the paper Learning to (Learn at Test Time): RNNs with Expressive Hidden States and its attempt to give recurrent models transformer-like long-context behavior without the quadratic cost of attention. It explains why standard RNN hidden states are a bottleneck, compares that limitation to transformers’ growing KV cache, and highlights a key empirical motivation: in the paper’s setup, Mamba’s token-level perplexity improvements flatten around 16k tokens while transformers keep improving deeper into a 32k context. The discussion focuses on the paper’s core idea of test-time training, where the hidden state is treated as a small inner model whose parameters are updated online with a self-supervised learning rule, rather than as a fixed vector summary. Listeners would find it interesting because it connects old fast-weights and dynamic-evaluation ideas to a new systems-level proposal for long-context efficiency, while also noting the open question of whether better perplexity truly translates into stronger retrieval and reasoning.

Sources:
1. Learning to (Learn at Test Time): RNNs with Expressive Hidden States — Yu Sun, Xinhao Li, Karan Dalal, Jiarui Xu, Arjun Vikram, Genghan Zhang, Yann Dubois, Xinlei Chen, Xiaolong Wang, Sanmi Koyejo, Tatsunori Hashimoto, Carlos Guestrin, 2024
http://arxiv.org/abs/2407.04620
2. Using Fast Weights to Attend to the Recent Past — Jimmy Ba, Geoffrey Hinton, Volodymyr Mnih, Joel Z. Leibo, Catalin Ionescu, 2016
https://arxiv.org/abs/1610.06258
3. Linear Transformers Are Secretly Fast Weight Programmers — Imanol Schlag, Kazuki Irie, Jürgen Schmidhuber, 2021
https://arxiv.org/abs/2102.11174
4. Mamba: Linear-Time Sequence Modeling with Selective State Spaces — Albert Gu, Tri Dao, 2023
https://arxiv.org/abs/2312.00752
5. Learning to (Learn at Test Time): RNNs with Expressive Hidden States — Yu Sun, Xinhao Li, Karan Dalal, Xiaolong Wang, Tatsunori Hashimoto, Carlos Guestrin, 2024
https://arxiv.org/abs/2407.04620
6. Dynamic Evaluation of Transformer Language Models — Ben Krause, Emmanuel Kahembwe, Iain Murray, Steve Renals, 2019
https://arxiv.org/abs/1904.08378
7. Effective Long-Context Scaling of Foundation Models — Wenhan Xiong et al., 2023
https://arxiv.org/abs/2309.16039
8. Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models — Soham De et al., 2024
https://arxiv.org/abs/2402.19427
9. Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality — Tri Dao, Albert Gu, 2024
https://arxiv.org/abs/2405.21060
10. An Empirical Study of Mamba-based Language Models — Roger Waleffe et al., 2024
https://arxiv.org/abs/2406.07887
11. KVzip: Query-Agnostic KV Cache Compression with Context Reconstruction — Jang-Hyun Kim et al., 2025
https://arxiv.org/abs/2505.23416
12. KVLink: Accelerating Large Language Models via Efficient KV Cache Reuse — Jingbo Yang et al., 2025
https://arxiv.org/abs/2502.16002
13. ReMamba: Equip Mamba with Effective Long-Sequence Modeling — Danlong Yuan et al., 2024
https://arxiv.org/abs/2408.15496
14. LongMamba: Enhancing Mamba's Long Context Capabilities via Training-Free Receptive Field Enlargement — Zhifan Ye et al., 2025
https://arxiv.org/abs/2504.16053
15. Fast-weight Product Key Memory — Tianyu Zhao, Llion Jones, 2026
https://arxiv.org/abs/2601.00671
16. Test-Time Learning for Large Language Models — Jinwu Hu et al., 2025
https://arxiv.org/abs/2505.20633
17. Test-Time Training on Nearest Neighbors for Large Language Models — Moritz Hardt, Yu Sun, 2023
https://arxiv.org/abs/2305.18466
18. AI Post Transformers: Titans: Learning to Memorize at Test Time — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-20-titans-learning-to-memorize-at-test-time-054662.mp3
19. AI Post Transformers: Gated Linear Attention for Efficient Long Sequences — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-18-gated-linear-attention-for-efficient-lon-c858ab.mp3
20. AI Post Transformers: Kimi Linear: Efficient Expressive Attention Architecture — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/kimi-linear-efficient-expressive-attention-architecture/
21. AI Post Transformers: How Induction Heads Emerge in Transformers — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-03-how-induction-heads-emerge-in-transforme-a7bfcb.mp3
22. AI Post Transformers: Explicit Information Transmission for Context Compression — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-explicit-information-transmission-for-co-24e3c2.mp3
23. AI Post Transformers: Memory-Bound, Not Bandwidth-Limited Batch-1 LLM Decode — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-02-memory-bound-not-bandwidth-limited-batch-114799.mp3
Interactive Visualization: Learning at Test Time with Expressive RNN States

This episode explores whether transformers really need separate query, key, and value projections, treating the problem as weight tying inside attention rather than as a brand-new model design. It explains why KV-cache size and memory bandwidth are major bottlenecks for long-context, on-device decoding, then compares increasingly aggressive sharing schemes, especially the difference between tying keys and values versus tying queries and keys. The discussion emphasizes that the broader sweep happens at 300M parameters, while only the shared-K/V variant is carried to 1.2B scale and remains in contention against practical baselines like grouped-query and multi-query attention. Listeners get a concrete deployment tradeoff: shared K/V can reduce KV-cache memory by about 50 percent at roughly a 3.1 percent perplexity cost, making the episode especially interesting for anyone focused on efficient inference.

Interactive Visualization: Do Transformers Need Three Projections?
Sources:
1. Do Transformers Need Three Projections? Systematic Study of QKV Variants — Ali Kayyam, Anusha Madan Gopal, M Anthony Lewis, 2026
http://arxiv.org/abs/2606.04032v2
2. Attention Is All You Need — Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, 2017
https://arxiv.org/abs/1706.03762
3. Fast Transformer Decoding: One Write-Head is All You Need — Noam Shazeer, 2019
https://arxiv.org/abs/1911.02150
4. GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints — Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebrón, Sumit Sanghai, 2023
https://arxiv.org/abs/2305.13245
5. Do Transformers Need Three Projections? Systematic Study of QKV Variants — Ali Kayyam, Anusha Madan Gopal, M. Anthony Lewis, 2026
https://arxiv.org/abs/2606.04032
6. Using the Output Embedding to Improve Language Models — Ofir Press, Lior Wolf, 2017
https://arxiv.org/abs/1608.05859
7. Linformer: Self-Attention with Linear Complexity — Sinong Wang, Belinda Z. Li, Madian Khabsa, Han Fang, Hao Ma, 2020
https://scholar.google.com/scholar?q=Linformer%3A+Self-Attention+with+Linear+Complexity
8. Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention — Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, Francois Fleuret, 2020
https://scholar.google.com/scholar?q=Transformers+are+RNNs%3A+Fast+Autoregressive+Transformers+with+Linear+Attention
9. Mamba: Linear-Time Sequence Modeling with Selective State Spaces — Albert Gu, Tri Dao, 2023
https://scholar.google.com/scholar?q=Mamba%3A+Linear-Time+Sequence+Modeling+with+Selective+State+Spaces
10. AsymKV: Enabling 1-Bit Quantization of KV Cache with Layer-Wise Asymmetric Quantization Configurations — Qian Tao et al., 2024
https://arxiv.org/abs/2410.13212
11. LongHeads: Multi-Head Attention is Secretly a Long Context Processor — Yi Lu et al., 2024
https://arxiv.org/abs/2402.10685
12. MuDAF: Long-Context Multi-Document Attention Focusing through Contrastive Learning on Attention Heads — Weihao Liu et al., 2025
https://arxiv.org/abs/2502.13963
13. Squeezed Attention: Accelerating Long Context Length LLM Inference — Coleman Hooper et al., 2024
https://arxiv.org/abs/2411.09688
14. Beyond Uniform Query Distribution: Key-Driven Grouped Query Attention — Zohaib Khan et al., 2024
https://arxiv.org/abs/2408.08454
15. Weight Decay Induces Low-Rank Attention Layers — Seijin Kobayashi et al., 2024
https://arxiv.org/abs/2410.23819
16. Dissecting Query-Key Interaction in Vision Transformers — Xu Pan et al., 2024
https://arxiv.org/abs/2405.14880
17. AI Post Transformers: Affordable Large-Scale Decoding Through Model-System Co-Design — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-19-affordable-large-scale-decoding-through-e1d7ed.mp3
18. AI Post Transformers: Memory-Bound, Not Bandwidth-Limited Batch-1 LLM Decode — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-02-memory-bound-not-bandwidth-limited-batch-114799.mp3
19. AI Post Transformers: Prefill-as-a-Service for Cross-Datacenter KV Cache — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-19-prefill-as-a-service-for-cross-datacente-7560be.mp3
20. AI Post Transformers: TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-25-turboquant-online-vector-quantiz-1967b7.mp3
21. AI Post Transformers: How Induction Heads Emerge in Transformers — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-03-how-induction-heads-emerge-in-transforme-a7bfcb.mp3
Interactive Visualization: Do Transformers Need Three Projections?

This episode explores the position paper Robots Need More Than VLAs & World Models and its claim that the main bottleneck in robotics may be grounding: turning raw physical behavior into robot-usable signals such as actions, contacts, task phases, goals, and rewards. It explains why vision-language-action models, world models, and reward models play different roles, and why simply scaling policy transformers cannot recover supervision that was never captured in the data. The discussion also digs into cross-embodiment learning and task-preserving retargeting, focusing on how humans and different robots can share useful experience despite mismatched bodies, sensors, and action spaces. A standout example is EgoMimic, which uses egocentric human video, 3D hand tracking, cross-domain alignment, and joint human-robot training to improve long-horizon real-robot manipulation, giving listeners a concrete picture of what might actually unlock broader robot generalization.

Sources:
1. Robots Need More than VLA and World Models — Elis Karcini, Faisal Mehrban, Quang Nguyen, Mac Schwager, Arash Ajoudani, Cesar Cadena, Jan Peters, Marco Hutter, Haitham Bou-Ammar, 2026
http://arxiv.org/abs/2606.06556
2. RT-1: Robotics Transformer for Real-World Control at Scale — Anthony Brohan, Noah Brown, Chelsea Finn, Sergey Levine, et al., 2022
https://scholar.google.com/scholar?q=RT-1%3A+Robotics+Transformer+for+Real-World+Control+at+Scale
3. RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control — Anthony Brohan, Noah Brown, Danny Driess, Karol Hausman, Chelsea Finn, Sergey Levine, et al., 2023
https://scholar.google.com/scholar?q=RT-2%3A+Vision-Language-Action+Models+Transfer+Web+Knowledge+to+Robotic+Control
4. OpenVLA: An Open-Source Vision-Language-Action Model — Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Chelsea Finn, Sergey Levine, et al., 2024
https://scholar.google.com/scholar?q=OpenVLA%3A+An+Open-Source+Vision-Language-Action+Model
5. π0: A Vision-Language-Action Flow Model for General Robot Control — Kevin Black, Noah Brown, Danny Driess, Karol Hausman, Sergey Levine, Chelsea Finn, et al., 2024
https://scholar.google.com/scholar?q=%CF%800%3A+A+Vision-Language-Action+Flow+Model+for+General+Robot+Control
6. Deep reinforcement learning from human preferences — Paul Christiano, Jan Leike, Tom B. Brown, Dario Amodei, et al., 2017
https://scholar.google.com/scholar?q=Deep+reinforcement+learning+from+human+preferences
7. Learning Generalizable Robotic Reward Functions from "In-The-Wild" Human Videos — Annie S. Chen, Suraj Nair, Chelsea Finn, 2021
https://scholar.google.com/scholar?q=Learning+Generalizable+Robotic+Reward+Functions+from+%22In-The-Wild%22+Human+Videos
8. Language to Rewards for Robotic Skill Synthesis — Wenhao Yu, Nimrod Gileadi, Chuyuan Fu, Brian Ichter, Ted Xiao, Fei Xia, et al., 2023
https://scholar.google.com/scholar?q=Language+to+Rewards+for+Robotic+Skill+Synthesis
9. RoboReward: General-Purpose Vision-Language Reward Models for Robotics — Tony Lee, Andrew Wagenmaker, Karl Pertsch, Percy Liang, Sergey Levine, Chelsea Finn, 2026
https://scholar.google.com/scholar?q=RoboReward%3A+General-Purpose+Vision-Language+Reward+Models+for+Robotics
10. Open X-Embodiment: Robotic Learning Datasets and RT-X Models — Open X-Embodiment Collaboration; Abby O'Neill, Fei Xia, Chelsea Finn, Sergey Levine, et al., 2023
https://scholar.google.com/scholar?q=Open+X-Embodiment%3A+Robotic+Learning+Datasets+and+RT-X+Models
11. XSkill: Cross Embodiment Skill Discovery — Mengda Xu, Zhenjia Xu, Cheng Chi, Manuela Veloso, Shuran Song, 2023
https://scholar.google.com/scholar?q=XSkill%3A+Cross+Embodiment+Skill+Discovery
12. Cross-Embodiment Robot Manipulation Skill Transfer using Latent Space Alignment — Tianyu Wang, Dwait Bhatt, Xiaolong Wang, Nikolay Atanasov, 2024
https://scholar.google.com/scholar?q=Cross-Embodiment+Robot+Manipulation+Skill+Transfer+using+Latent+Space+Alignment
13. Gemini Robotics 1.5: Pushing the Frontier of Generalist Robots with Advanced Embodied Reasoning, Thinking, and Motion Transfer — Gemini Robotics Team; Abbas Abdolmaleki, Anthony Brohan, Keerthana Gopalakrishnan, Ted Xiao, et al., 2025
https://scholar.google.com/scholar?q=Gemini+Robotics+1.5%3A+Pushing+the+Frontier+of+Generalist+Robots+with+Advanced+Embodied+Reasoning%2C+Thinking%2C+and+Motion+Transfer
14. Learning Latent Plans from Play — Corey Lynch, Mohi Khansari, Ted Xiao, Vikash Kumar, Sergey Levine, Pierre Sermanet, et al., 2019
https://scholar.google.com/scholar?q=Learning+Latent+Plans+from+Play
15. R3M: A Universal Visual Representation for Robot Manipulation — Suraj Nair, Aravind Rajeswaran, Vikash Kumar, Chelsea Finn, Abhinav Gupta, 2022
https://scholar.google.com/scholar?q=R3M%3A+A+Universal+Visual+Representation+for+Robot+Manipulation
16. Zero-Shot Robot Manipulation from Passive Human Videos — Homanga Bharadhwaj, Abhinav Gupta, Shubham Tulsiani, Vikash Kumar, 2023
https://scholar.google.com/scholar?q=Zero-Shot+Robot+Manipulation+from+Passive+Human+Videos
17. GenSim: Generating Robotic Simulation Tasks via Large Language Models — Lirui Wang, Yiyang Ling, Zhecheng Yuan, Mohit Shridhar, Huazhe Xu, Xiaolong Wang, et al., 2023
https://scholar.google.com/scholar?q=GenSim%3A+Generating+Robotic+Simulation+Tasks+via+Large+Language+Models
18. $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control — Kevin Black, Noah Brown, Danny Driess, Adnan Esmail, Michael Equi, et al., 2024
https://scholar.google.com/scholar?q=%24%5Cpi_0%24%3A+A+Vision-Language-Action+Flow+Model+for+General+Robot+Control
19. EgoMimic: Scaling Imitation Learning via Egocentric Video — Simar Kareer, Dhruv Patel, Ryan Punamiya, Pranay Mathur, Shuo Cheng, Chen Wang, Judy Hoffman, Danfei Xu, 2024
https://scholar.google.com/scholar?q=EgoMimic%3A+Scaling+Imitation+Learning+via+Egocentric+Video
20. LEGATO: Cross-Embodiment Imitation Using a Grasping Tool — Mingyo Seo, H. Andy Park, Shenli Yuan, Yuke Zhu, Luis Sentis, 2024
https://scholar.google.com/scholar?q=LEGATO%3A+Cross-Embodiment+Imitation+Using+a+Grasping+Tool
21. Rank2Reward: Learning Shaped Reward Functions from Passive Video — Daniel Yang, Davin Tjia, Jacob Berg, Dima Damen, Pulkit Agrawal, Abhishek Gupta, 2024
https://scholar.google.com/scholar?q=Rank2Reward%3A+Learning+Shaped+Reward+Functions+from+Passive+Video
22. Genie: Generative Interactive Environments — Jake Bruce, Michael Dennis, Ashley Edwards, Jack Parker-Holder, Yuge Shi, et al., 2024
https://scholar.google.com/scholar?q=Genie%3A+Generative+Interactive+Environments
23. Neural Scaling Laws in Robotics — Sebastian Sartor, Neil Thompson, 2024
https://arxiv.org/abs/2405.14005
24. Diffusion-VLA: Scaling Robot Foundation Models via Unified Diffusion and Autoregression — Junjie Wen et al., 2024
https://arxiv.org/abs/2412.03293
25. Scaling Cross-Embodied Learning: One Policy for Manipulation, Navigation, Locomotion and Aviation — Ria Doshi et al., 2024
https://arxiv.org/abs/2408.11812
26. RoVi-Aug: Robot and Viewpoint Augmentation for Cross-Embodiment Robot Learning — Lawrence Yunliang Chen et al., 2024
https://arxiv.org/abs/2409.03403
27. CLAM: Continuous Latent Action Models for Robot Learning from Unlabeled Demonstrations — Anthony Liang et al., 2025
https://arxiv.org/abs/2505.04999
28. DayDreamer: World Models for Physical Robot Learning — Philipp Wu, Alejandro Escontrela, Danijar Hafner, Ken Goldberg, Pieter Abbeel, 2022
https://arxiv.org/abs/2206.14176
29. Ctrl-World: A Controllable Generative World Model for Robot Manipulation — Yanjiang Guo et al., 2025
https://arxiv.org/abs/2510.10125
30. AI Post Transformers: DreamerV3 World Models Across 150 Tasks — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-20-dreamerv3-world-models-across-150-tasks-af5edb.mp3
31. AI Post Transformers: When LLM Judges Become Coin Flips — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-when-llm-judges-become-coin-flips-8b43ef.mp3
32. AI Post Transformers: SkillsBench for Evaluating Agent Skills — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-14-skillsbench-for-evaluating-agent-skills-58bb1e.mp3
Interactive Visualization: Robots Need More Than VLAs and World Models

This episode explores End-to-End Context Compression at Scale, a paper on whether learned context compression can beat the cost of long-context inference in quality, time to first token, and peak memory. It explains the main design choices behind the authors’ Latent Context Language Models, which use a 0.6B encoder and 4B decoder to replace long token sequences with learned latent memory at compression ratios from 1:4 to 1:16, and contrasts that approach with full-context prompting, retrieval, summarization, and KV-cache compression methods such as SnapKV and KVzip. The discussion highlights the paper’s core result: on RULER and LongBench EN-16, the released system reportedly sets a new Pareto frontier, delivering up to 8.8x faster inference on RULER and 5.2x faster on LongBench with lower memory use and stronger accuracy at aggressive compression. It also digs into the catch that makes the result interesting for practitioners: this speedup depends on a heavily trained system and changes the serving stack, so the real question is not just whether the benchmark wins are real, but whether learned compression is finally practical infrastructure for long-horizon agents and large-scale deployment.

Sources:
1. End-to-End Context Compression at Scale — Ang Li, Sean McLeish, Haozhe Chen, Nimit Kalra, Zaiqian Chen, Artem Gazizov, Venkata Anoop Suhas Kumar Morisetty, Bhavya Kailkhura, Harshitha Menon, Zhuang Liu, Brian R. Bartoldson, Tom Goldstein, Sanae Lotfi, Micah Goldblum, Pavel Izmailov, 2026
http://arxiv.org/abs/2606.09659
2. Learning to Compress Prompts with Gist Tokens — Jesse Mu, Xiang Lisa Li, Noah Goodman, 2023
https://arxiv.org/abs/2304.08467
3. Adapting Language Models to Compress Contexts — Alexis Chevalier, Alexander Wettig, Anirudh Ajith, Danqi Chen, 2023
https://arxiv.org/abs/2305.14788
4. Long-Context Language Modeling with Parallel Context Encoding — Howard Yen, Tianyu Gao, Danqi Chen, 2024
https://arxiv.org/abs/2402.16617
5. ARC-Encoder: learning compressed text representations for large language models — Hippolyte Pilchen, Edouard Grave, Patrick Perez, 2025
https://arxiv.org/abs/2510.20535
6. SnapKV: LLM Knows What You are Looking for Before Generation — Yuhong Li, Yingbing Huang, Bowen Yang, Bharat Venkitesh, Acyr Locatelli, Hanchen Ye, Tianle Cai, Patrick Lewis, Deming Chen, 2024
https://scholar.google.com/scholar?q=SnapKV%3A+LLM+Knows+What+You+are+Looking+for+Before+Generation
7. KVzip: Query-Agnostic KV Cache Compression with Context Reconstruction — Jang-Hyun Kim, Jinuk Kim, Sangwoo Kwon, Jae W. Lee, Sangdoo Yun, Hyun Oh Song, 2025
https://scholar.google.com/scholar?q=KVzip%3A+Query-Agnostic+KV+Cache+Compression+with+Context+Reconstruction
8. Fast KV Compaction via Attention Matching — Adam Zweiger, Xinghong Fu, Han Guo, Yoon Kim, 2026
https://scholar.google.com/scholar?q=Fast+KV+Compaction+via+Attention+Matching
9. Cartridges: Lightweight and general-purpose long context representations via self-study — Sabri Eyuboglu, Ryan Ehrlich, Simran Arora, Neel Guha, Dylan Zinsley, Emily Liu, Will Tennien, Atri Rudra, James Zou, Azalia Mirhoseini, Christopher Re, 2025
https://scholar.google.com/scholar?q=Cartridges%3A+Lightweight+and+general-purpose+long+context+representations+via+self-study
10. Latent Context Compilation: Distilling Long Context into Compact Portable Memory — Zeju Li, Yizhou Zhou, Qiang Xu, 2026
https://scholar.google.com/scholar?q=Latent+Context+Compilation%3A+Distilling+Long+Context+into+Compact+Portable+Memory
11. RULER: What's the Real Context Size of Your Long-Context Language Models? — Cheng-Ping Hsieh, Simeng Sun, Samuel Kriman, Shantanu Acharya, Dima Rekesh, Fei Jia, Yang Zhang, Boris Ginsburg, 2024
https://scholar.google.com/scholar?q=RULER%3A+What%27s+the+Real+Context+Size+of+Your+Long-Context+Language+Models%3F
12. LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks — Yushi Bai, Shangqing Tu, Jiajie Zhang, Hao Peng, Xiaozhi Wang, Xin Lv, Shulin Cao, Jiazheng Xu, Lei Hou, Yuxiao Dong, Jie Tang, Juanzi Li, 2024
https://scholar.google.com/scholar?q=LongBench+v2%3A+Towards+Deeper+Understanding+and+Reasoning+on+Realistic+Long-context+Multitasks
13. ObjectCache: Layerwise Object-Storage Retrieval for KV Cache Reuse — Yu Zhu et al., 2026
https://arxiv.org/abs/2605.22850
14. PagedEviction: Structured Block-wise KV Cache Pruning for Efficient Large Language Model Inference — Krishna Teja Chitty-Venkata et al., 2025
https://arxiv.org/abs/2509.04377
15. ChunkKV: Semantic-Preserving KV Cache Compression for Efficient Long-Context LLM Inference — Xiang Liu et al., 2025
https://arxiv.org/abs/2502.00299
16. Long Context Compression with Activation Beacon — Peitian Zhang et al., 2024
https://arxiv.org/abs/2401.03462
17. AI Post Transformers: Explicit Information Transmission for Context Compression — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-explicit-information-transmission-for-co-24e3c2.mp3
18. AI Post Transformers: TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-25-turboquant-online-vector-quantiz-1967b7.mp3
19. AI Post Transformers: DeepSeek-V4 and Practical Million-Token Context — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-25-deepseek-v4-and-practical-million-token-6f4de1.mp3
20. AI Post Transformers: Memory Sparse Attention for 100M-Token Scaling — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-07-memory-sparse-attention-for-100m-token-s-377cff.mp3
21. AI Post Transformers: Long Context Pre-Training with Lighthouse Attention — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-13-long-context-pre-training-with-lighthous-e85bbe.mp3
22. AI Post Transformers: Compressed Convolutional Attention in Latent Space — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-25-compressed-convolutional-attention-in-la-61e1cf.mp3

This episode explores why decoder-style language models can generate fluent text yet still underperform dedicated embedding models when asked for zero-shot sentence vectors, despite embeddings being critical infrastructure for search, retrieval-augmented generation, clustering, and recommendation. It examines the paper’s main argument that the problem is not just bad pooling or prompting, but a deeper geometric bias: sentence representations appear overly aligned with frequent, low-information tokens, which becomes visible when they are projected through the model’s unembedding matrix. It also digs into the debate over whether decoder models mainly suffer from poor extraction recipes or from genuinely weaker embedding spaces, using concrete details from the authors’ code such as prompt-based summarization, last-token pooling, and custom truncation. A listener would find it interesting because the discussion connects mechanistic interpretability to real embedding-system design, and even suggests why filtering or reducing dimensions could improve both quality and efficiency.

Sources:
1. Your UnEmbedding Matrix is Secretly a Feature Lens for Text Embeddings — Songhao Wu, Zhongxin Chen, Yuxuan Liu, Heng Cui, Cong Li, Rui Yan, 2026
http://arxiv.org/abs/2606.07502
2. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks — Nils Reimers, Iryna Gurevych, 2019
https://arxiv.org/abs/1908.10084
3. BEIR: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval Models — Nandan Thakur, Nils Reimers, Andreas Ruckle, Abhishek Srivastava, Iryna Gurevych, 2021
https://arxiv.org/abs/2104.08663
4. MTEB: Massive Text Embedding Benchmark — Niklas Muennighoff, Nouamane Tazi, Loic Magne, Nils Reimers, 2022
https://arxiv.org/abs/2210.07316
5. LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders — Parishad BehnamGhader, Vaibhav Adlakha, Marius Mosbach, Dzmitry Bahdanau, Nicolas Chapados, Siva Reddy, 2024
https://arxiv.org/abs/2404.05961
6. Indexing by Latent Semantic Analysis — Scott Deerwester, Susan T. Dumais, George W. Furnas, Thomas K. Landauer, Richard Harshman, 1990
https://www.cs.csustan.edu/~mmartin/LDS/Deerwester-et-al.pdf
7. All-but-the-Top: Simple and Effective Postprocessing for Word Representations — Jiaqi Mu, Suma Bhat, Pramod Viswanath, 2018
https://arxiv.org/abs/1702.01417
8. Whitening Sentence Representations for Better Semantics and Faster Retrieval — Jianlin Su, Jiarun Cao, Weijie Liu, Yangyiwen Ou, 2021
https://arxiv.org/abs/2103.15316
9. Matryoshka Representation Learning — Aditya Kusupati, Gantavya Bhatt, Aniket Rege, Matthew Wallingford, Aditya Sinha, Vivek Ramanujan, William Howard-Snyder, Kaifeng Chen, Sham Kakade, Prateek Jain, Ali Farhadi, 2022
https://arxiv.org/abs/2205.13147
10. Scaling Sentence Embeddings with Large Language Models — Ting Jiang, Shaohan Huang, Zhongzhi Luan, Deqing Wang, Fuzhen Zhuang, 2023
https://scholar.google.com/scholar?q=Scaling+Sentence+Embeddings+with+Large+Language+Models
11. Improving Text Embeddings with Large Language Models — Liang Wang, Nan Yang, Xiaolong Huang, Linjun Yang, Rangan Majumder, Furu Wei, 2024
https://scholar.google.com/scholar?q=Improving+Text+Embeddings+with+Large+Language+Models
12. Eliciting Latent Predictions from Transformers with the Tuned Lens — Nora Belrose, Zach Furman, Logan Smith, Danny Halawi, Igor Ostrovsky, Lev McKinney, Stella Biderman, Jacob Steinhardt, 2023
https://scholar.google.com/scholar?q=Eliciting+Latent+Predictions+from+Transformers+with+the+Tuned+Lens
13. Ditto: A Simple and Efficient Approach to Improve Sentence Embeddings — Qian Chen, Wen Wang, Qinglin Zhang, Siqi Zheng, Chong Deng, Hai Yu, Jiaqing Liu, Yukun Ma, Chong Zhang, 2023
https://scholar.google.com/scholar?q=Ditto%3A+A+Simple+and+Efficient+Approach+to+Improve+Sentence+Embeddings
14. Is anisotropy really the cause of BERT embeddings not being semantic? — Alejandro Fuster Baggetto, Victor Fresno, 2022
https://scholar.google.com/scholar?q=Is+anisotropy+really+the+cause+of+BERT+embeddings+not+being+semantic%3F
15. Mitigating Frequency Bias and Anisotropy in Language Model Pre-Training with Syntactic Smoothing — Richard Diehl Martinez et al., 2024
https://arxiv.org/abs/2410.11462
16. Anisotropy Is Inherent to Self-Attention in Transformers — Nathan Godey, Eric de la Clergerie, Benoit Sagot, 2024
https://arxiv.org/abs/2401.12143
17. Indic-TunedLens: Interpreting Multilingual Models in Indian Languages — Mihir Panchal et al., 2026
https://arxiv.org/abs/2602.15038
18. KV-Embedding: Training-free Text Embedding via Internal KV Re-routing in Decoder-only LLMs — Yixuan Tang, Yi Yang, 2026
https://arxiv.org/abs/2601.01046
19. Causal2Vec: Improving Decoder-only LLMs as Embedding Models through a Contextual Token — Ailiang Lin et al., 2025
https://arxiv.org/abs/2507.23386
20. On the Theoretical Limitations of Embedding-Based Retrieval — Orion Weller et al., 2025
https://arxiv.org/abs/2508.21038
21. AI Post Transformers: Why Transformers Fail at Counting — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-08-why-transformers-fail-at-counting-137924.mp3
22. AI Post Transformers: Neural Chameleons and Evading Activation Monitors — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-14-neural-chameleons-and-evading-activation-bc470e.mp3
23. AI Post Transformers: TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-25-turboquant-online-vector-quantiz-1967b7.mp3
Interactive Visualization: Unembedding Matrices as Feature Lenses for Embeddings

This episode explores Predictive Query Language (PQL), a SQL-shaped domain-specific language for defining supervised learning tasks directly over relational databases by specifying the prediction target, entity, and future time horizon in one declarative statement. It explains why training label generation is often the real bottleneck in applied machine learning, unpacking concepts like prediction entities, relational learning, anchor times, point-in-time consistency, and information leakage. The discussion compares PQL to earlier work on prediction engineering and relational deep learning, arguing that its main contribution is not a new model but a more disciplined way to construct temporally valid prediction problems from messy, multi-table operational data. Listeners interested in real-world ML systems will find it interesting because it focuses on the part most papers skip: how to ask the predictive question correctly when database history is incomplete, revised, and easy to misuse.

Sources:
1. Predictive Query Language for Relational Databases
https://arxiv.org/pdf/2602.09572
2. Declarative Machine Learning - A Classification of Basic Properties and Types — Matthias Boehm, Alexandre V. Evfimievski, Niketan Pansare, Berthold Reinwald, 2016
https://arxiv.org/abs/1605.05826
3. The MADlib Analytics Library or MAD Skills, the SQL — Joseph M. Hellerstein, Christopher Re, Florian Schoppmann, Daisy Zhe Wang, et al., 2012
https://arxiv.org/abs/1208.4165
4. MLog: Towards Declarative In-Database Machine Learning — Xupeng Li, Bin Cui, Yiru Chen, Wentao Wu, Ce Zhang, 2017
https://www.vldb.org/pvldb/vol10/p1933-zhang.pdf
5. sql4ml A declarative end-to-end workflow for machine learning — Nantia Makrynioti, Ruy Ley-Wild, Vasilis Vassalos, 2019
https://arxiv.org/abs/1907.12415
6. Deep Feature Synthesis: Towards Automating Data Science Endeavors — James Max Kanter, Kalyan Veeramachaneni, 2015
https://doi.org/10.1109/DSAA.2015.7344858
7. Label, Segment, Featurize: A Cross Domain Framework for Prediction Engineering — James Max Kanter, Owen Gillespie, Kalyan Veeramachaneni, 2016
https://www.maxkanter.com/papers/DSAA_LSF_2016.pdf
8. Predictive Query Language: A Domain-Specific Language for Predictive Modeling on Relational Databases — Vid Kocijan, Jinu Sunil, Jan Eric Lenssen, Viman Deb, et al., 2026
https://arxiv.org/abs/2602.09572
9. RelBench: A Benchmark for Deep Learning on Relational Databases — Joshua Robinson, Rishabh Ranjan, Weihua Hu, Kexin Huang, et al., 2024
https://arxiv.org/abs/2407.20060
10. Temporal features in SQL:2011 — Krishna Kulkarni, Jan-Eike Michels, 2012
https://sigmodrecord.org/publications/sigmodRecord/1209/pdfs/07.industry.kulkarni.pdf
11. Time Travel and Provenance for Machine Learning Pipelines — Alexandru A. Ormenisan, Moritz Meister, Fabio Buso, Robin Andersson, Seif Haridi, Jim Dowling, 2020
https://www.usenix.org/conference/opml20/presentation/ormenisan
12. Optimizing Data Pipelines for Machine Learning in Feature Stores — Rui Liu, Kwanghyun Park, Fotis Psallidas, Xiaoyong Zhu, et al., 2023
https://www.microsoft.com/en-us/research/publication/optimizing-data-pipelines-for-machine-learning-in-feature-stores/
13. The Hopsworks Feature Store for Machine Learning — Javier de la Rua Martinez, Fabio Buso, Antonios Kouzoupis, Alexandru A. Ormenisan, et al., 2024
https://content.hopsworks.ai/hubfs/The_Hopsworks_Feature_Store_for_Machine_Learning.pdf
14. Leakage in Data Mining: Formulation, Detection, and Avoidance — Shachar Kaufman, Saharon Rosset, Claudia Perlich, 2012
https://doi.org/10.1145/2020408.2020496
15. How to avoid machine learning pitfalls: a guide for academic researchers — Michael A. Lones, 2021
https://arxiv.org/abs/2108.02497
16. Leakage and the reproducibility crisis in machine-learning-based science — Sayash Kapoor, Arvind Narayanan, 2023
https://doi.org/10.1016/j.patter.2023.100804
17. MLearn: A Declarative Machine Learning Language for Database Systems — Maximilian E. Schuele, Matthias Bungeroth, Alfons Kemper, Stephan Guennemann, Thomas Neumann, 2019
https://scholar.google.com/scholar?q=MLearn%3A+A+Declarative+Machine+Learning+Language+for+Database+Systems
18. End-to-end Optimization of Machine Learning Prediction Queries — Kwanghyun Park, Karla Saur, Dalitso Banda, Rathijit Sen, Matteo Interlandi, Konstantinos Karanasos, 2022
https://scholar.google.com/scholar?q=End-to-end+Optimization+of+Machine+Learning+Prediction+Queries
19. Relational Deep Learning: Graph Representation Learning on Relational Databases — Matthias Fey, Weihua Hu, Kexin Huang, Jan Eric Lenssen, Rishabh Ranjan, Joshua Robinson, Rex Ying, Jiaxuan You, Jure Leskovec, 2023
https://scholar.google.com/scholar?q=Relational+Deep+Learning%3A+Graph+Representation+Learning+on+Relational+Databases
20. KumoRFM: A Foundation Model for In-Context Learning on Relational Data — Matthias Fey, Vid Kocijan, Federico Lopez, Jan Eric Lenssen, Jure Leskovec, 2025
https://scholar.google.com/scholar?q=KumoRFM%3A+A+Foundation+Model+for+In-Context+Learning+on+Relational+Data
21. Toward a Declarative Query Language for Machine Learning — Hasan M. Jamil, 2024
https://vldb.org/workshops/2024/proceedings/TaDA/TaDA.2.pdf
22. Implementing a Declarative Query Language for High Level Machine Learning Application Design — Hasan Rahman, Hasan M. Jamil, 2025
https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5210204
23. LinkAlign: Scalable Schema Linking for Real-World Large-Scale Multi-Database Text-to-SQL — Yihan Wang, Peiyu Liu, Xin Yang, 2025
https://arxiv.org/abs/2503.18596
24. E-SQL: Direct Schema Linking via Question Enrichment in Text-to-SQL — Hasan Alp Caferoglu, Ozgur Ulusoy, 2024
https://arxiv.org/abs/2409.16751
25. AutoLink: Autonomous Schema Exploration and Expansion for Scalable Schema Linking in Text-to-SQL at Scale — Ziyang Wang et al., 2025
https://arxiv.org/abs/2511.17190
26. From Alignment to Entailment: A Unified Textual Entailment Framework for Entity Alignment — Yu Zhao et al., 2023
https://arxiv.org/abs/2305.11501
27. AI Post Transformers: SGLang for Faster Structured LLM Programs — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-06-sglang-for-faster-structured-llm-program-c59f1c.mp3
28. AI Post Transformers: Caffe and the Rise of CNN Frameworks — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-06-caffe-and-the-rise-of-cnn-frameworks-cf15f3.mp3

This episode explores KumoRFM-2, a relational foundation model designed to learn directly from connected database tables instead of flattening customers, orders, products, and tickets into a single feature table. It explains why relational learning matters for enterprise tasks such as churn, fraud, and demand prediction, arguing that flattening often erases multi-hop relationships, repeated interactions, and temporal patterns that carry the real signal. The discussion centers on KumoRFM-2’s main technical claim: a two-stage, task-conditioned attention pipeline that first selects relevant information within each table and then aggregates evidence across foreign-key neighborhoods and labeled in-context examples derived from predictive queries. Listeners would find it interesting because it connects a very practical data-engineering pain point to a broader question about whether pretrained, database-native models can beat hand-built tabular pipelines without cheating on time-aware prediction.

Sources:
1. KumoRFM-2: Scaling Foundation Models for Relational Learning — Valter Hudovernik, Federico López, Vid Kocijan, Akihiro Nitta, Jan Eric Lenssen, Jure Leskovec, Matthias Fey, 2026
http://arxiv.org/abs/2604.12596
2. Learning Probabilistic Relational Models — Nir Friedman, Lise Getoor, Daphne Koller, Avi Pfeffer, 1999
https://scholar.google.com/scholar?q=Learning+Probabilistic+Relational+Models
3. Relational Inductive Biases, Deep Learning, and Graph Networks — Peter W. Battaglia, Jessica B. Hamrick, Victor Bapst, Alvaro Sanchez-Gonzalez, Yujia Li, Razvan Pascanu, et al., 2018
https://scholar.google.com/scholar?q=Relational+Inductive+Biases%2C+Deep+Learning%2C+and+Graph+Networks
4. Relational Deep Learning: Graph Representation Learning on Relational Databases — Matthias Fey, Weihua Hu, Kexin Huang, Jan Eric Lenssen, Rishabh Ranjan, Joshua Robinson, Rex Ying, Jiaxuan You, Jure Leskovec, 2023
https://scholar.google.com/scholar?q=Relational+Deep+Learning%3A+Graph+Representation+Learning+on+Relational+Databases
5. RelBench: A Benchmark for Deep Learning on Relational Databases — Joshua Robinson, Rishabh Ranjan, Weihua Hu, Kexin Huang, Jiaqi Han, Alejandro Dobles, Matthias Fey, Jan E. Lenssen, Jure Leskovec, et al., 2024
https://scholar.google.com/scholar?q=RelBench%3A+A+Benchmark+for+Deep+Learning+on+Relational+Databases
6. Position: Why Tabular Foundation Models Should Be a Research Priority — Boris Van Breugel, Mihaela Van Der Schaar, 2024
https://scholar.google.com/scholar?q=Position%3A+Why+Tabular+Foundation+Models+Should+Be+a+Research+Priority
7. Accurate Predictions on Small Data with a Tabular Foundation Model — Noah Hollmann, Samuel Müller, Katharina Eggensperger, Frank Hutter, et al., 2025
https://scholar.google.com/scholar?q=Accurate+Predictions+on+Small+Data+with+a+Tabular+Foundation+Model
8. TabICL: A Tabular Foundation Model for In-Context Learning on Large Data — Jingang Qu, David Holzmüller, Gaël Varoquaux, Marine Le Morvan, 2025
https://scholar.google.com/scholar?q=TabICL%3A+A+Tabular+Foundation+Model+for+In-Context+Learning+on+Large+Data
9. KumoRFM: A Foundation Model for In-Context Learning on Relational Data — Matthias Fey, Vid Kocijan, Federico Lopez, Jan Eric Lenssen, Jure Leskovec, 2025
https://scholar.google.com/scholar?q=KumoRFM%3A+A+Foundation+Model+for+In-Context+Learning+on+Relational+Data
10. PluRel: Synthetic Data unlocks Scaling Laws for Relational Foundation Models — Vignesh Kothapalli et al., 2026
https://scholar.google.com/scholar?q=PluRel%3A+Synthetic+Data+unlocks+Scaling+Laws+for+Relational+Foundation+Models
11. Relational Transformer: Toward Zero-Shot Foundation Models for Relational Data — Rishabh Ranjan et al., 2026
https://scholar.google.com/scholar?q=Relational+Transformer%3A+Toward+Zero-Shot+Foundation+Models+for+Relational+Data
12. No Need to Train Your RDB Foundation Model — Linjie Xu, Yanlin Zhang, Quan Gan, Minjie Wang, and David Wipf, 2026
https://scholar.google.com/scholar?q=No+Need+to+Train+Your+RDB+Foundation+Model
13. TabICLv2: A better, faster, scalable, and open tabular foundation model — Jingang Qu, David Holzmuller, Gael Varoquaux, and Marine Le Morvan, 2026
https://scholar.google.com/scholar?q=TabICLv2%3A+A+better%2C+faster%2C+scalable%2C+and+open+tabular+foundation+model
14. Griffin: Towards a Graph-Centric Relational Database Foundation Model — Yanbo Wang, Xiyuan Wang, Quan Gan, Minjie Wang, Qibin Yang, David Wipf, and Muhan Zhang, 2025
https://scholar.google.com/scholar?q=Griffin%3A+Towards+a+Graph-Centric+Relational+Database+Foundation+Model
15. Graph Machine Learning Meets Multi-Table Relational Data — Quan Gan, Minjie Wang, David Wipf, Christos Faloutsos, 2024
https://scholar.google.com/scholar?q=Graph+Machine+Learning+Meets+Multi-Table+Relational+Data
16. Large Scale Transfer Learning for Tabular Data via Language Modeling — Josh Gardner, Juan C. Perdomo, Ludwig Schmidt, 2024
https://scholar.google.com/scholar?q=Large+Scale+Transfer+Learning+for+Tabular+Data+via+Language+Modeling
17. Towards Synthetic Data for Fine-tuning Tabular Foundation Models — Magnus Buhler, Lennart Purucker, Frank Hutter, 2025
https://scholar.google.com/scholar?q=Towards+Synthetic+Data+for+Fine-tuning+Tabular+Foundation+Models
18. Range-limited Augmentation for Few-shot Learning in Tabular Data with Comprehensive Benchmark — Kyungeun Lee et al., 2025
https://scholar.google.com/scholar?q=Range-limited+Augmentation+for+Few-shot+Learning+in+Tabular+Data+with+Comprehensive+Benchmark
19. AI Post Transformers: Muon Is Scalable for LLM Training — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-25-muon-is-scalable-for-llm-training-587ed8.mp3
20. AI Post Transformers: Scaling Laws for Multilingual Code Pretraining — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-scaling-laws-for-multilingual-code-pretr-7d220e.mp3

This episode explores Unified Neural Scaling Laws, a framework for predicting model performance when parameter count, data volume, training steps, inference compute, and training-recipe choices all change at once. It explains how the paper moves beyond classic smooth power-law curves by introducing broken multivariate scaling laws with regime shifts, including hyperbreaks, bottleneck versus non-bottleneck components, and joint interaction surfaces across training variables. The discussion highlights the paper’s argument that good forecasting must capture both beneficial scaling effects and harmful effects such as overfitting, bad hyperparameter regimes, and data or compute limits, rather than assuming one clean trend forever. A listener would find it interesting because it ties abstract scaling-law math directly to expensive real-world training decisions and to the question of whether pretraining gains actually transfer to downstream benchmarks.

Sources:
1. Unified Neural Scaling Laws — Ethan Caballero, Priyank Jaini, David Krueger, Irina Rish, 2026
http://arxiv.org/abs/2605.26248
2. A Constructive Prediction of the Generalization Error Across Scales — Jonathan S. Rosenfeld, Amir Rosenfeld, Yonatan Belinkov, Nir Shavit, 2020
https://scholar.google.com/scholar?q=A+Constructive+Prediction+of+the+Generalization+Error+Across+Scales
3. Scaling Laws for Neural Language Models — Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, et al., 2020
https://scholar.google.com/scholar?q=Scaling+Laws+for+Neural+Language+Models
4. Training Compute-Optimal Large Language Models — Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, et al., 2022
https://scholar.google.com/scholar?q=Training+Compute-Optimal+Large+Language+Models
5. Unified Neural Scaling Laws — Ethan Caballero, Priyank Jaini, David Krueger, Irina Rish, 2026
https://scholar.google.com/scholar?q=Unified+Neural+Scaling+Laws
6. Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour — Priya Goyal, Piotr Dollar, Ross Girshick, Kaiming He, et al., 2017
https://scholar.google.com/scholar?q=Accurate%2C+Large+Minibatch+SGD%3A+Training+ImageNet+in+1+Hour
7. An Empirical Model of Large-Batch Training — Sam McCandlish, Jared Kaplan, Dario Amodei, OpenAI Dota Team, 2018
https://scholar.google.com/scholar?q=An+Empirical+Model+of+Large-Batch+Training
8. Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer — Greg Yang, Edward J. Hu, Igor Babuschkin, Szymon Sidor, et al., 2022
https://scholar.google.com/scholar?q=Tensor+Programs+V%3A+Tuning+Large+Neural+Networks+via+Zero-Shot+Hyperparameter+Transfer
9. Resolving Discrepancies in Compute-Optimal Scaling of Language Models — Tomer Porian, Mitchell Wortsman, Jenia Jitsev, Ludwig Schmidt, Yair Carmon, 2024
https://scholar.google.com/scholar?q=Resolving+Discrepancies+in+Compute-Optimal+Scaling+of+Language+Models
10. Reconciling modern machine learning practice and the bias-variance trade-off — Mikhail Belkin, Daniel Hsu, Siyuan Ma, Soumik Mandal, 2019
https://scholar.google.com/scholar?q=Reconciling+modern+machine+learning+practice+and+the+bias-variance+trade-off
11. Deep Double Descent: Where Bigger Models and More Data Hurt — Preetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang, Boaz Barak, Ilya Sutskever, 2020
https://scholar.google.com/scholar?q=Deep+Double+Descent%3A+Where+Bigger+Models+and+More+Data+Hurt
12. Broken Neural Scaling Laws — Ethan Caballero, Kshitij Gupta, Irina Rish, David Krueger, 2023
https://scholar.google.com/scholar?q=Broken+Neural+Scaling+Laws
13. Scaling Data-Constrained Language Models — Niklas Muennighoff, Alexander M. Rush, Boaz Barak, et al., 2023
https://scholar.google.com/scholar?q=Scaling+Data-Constrained+Language+Models
14. Observational Scaling Laws and the Predictability of Language Model Performance — Yangjun Ruan, Chris J. Maddison, Tatsunori Hashimoto, 2024
https://scholar.google.com/scholar?q=Observational+Scaling+Laws+and+the+Predictability+of+Language+Model+Performance
15. Scaling Laws for Data Filtering -- Data Curation cannot be Compute Agnostic — Sachin Goyal et al., 2024
https://scholar.google.com/scholar?q=Scaling+Laws+for+Data+Filtering+--+Data+Curation+cannot+be+Compute+Agnostic
16. Revisiting Scaling Laws for Language Models: The Role of Data Quality and Training Strategies — Zhengyu Chen et al., 2025
https://scholar.google.com/scholar?q=Revisiting+Scaling+Laws+for+Language+Models%3A+The+Role+of+Data+Quality+and+Training+Strategies
17. Scaling Laws for Optimal Data Mixtures — Mustafa Shukor et al., 2025
https://scholar.google.com/scholar?q=Scaling+Laws+for+Optimal+Data+Mixtures
18. Joint MoE Scaling Laws: Mixture of Experts Can Be Memory Efficient — Jan Ludziejewski et al., 2025
https://scholar.google.com/scholar?q=Joint+MoE+Scaling+Laws%3A+Mixture+of+Experts+Can+Be+Memory+Efficient
19. Towards Thinking-Optimal Scaling of Test-Time Compute for LLM Reasoning — Wenkai Yang et al., 2025
https://scholar.google.com/scholar?q=Towards+Thinking-Optimal+Scaling+of+Test-Time+Compute+for+LLM+Reasoning
20. Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach — Jonas Geiping et al., 2025
https://scholar.google.com/scholar?q=Scaling+up+Test-Time+Compute+with+Latent+Reasoning%3A+A+Recurrent+Depth+Approach
21. Towards Robust Scaling Laws for Optimizers — Alexandra Volkova et al., 2026
https://scholar.google.com/scholar?q=Towards+Robust+Scaling+Laws+for+Optimizers
22. Neural Scaling Laws Rooted in the Data Distribution — Ari Brill, 2024
https://scholar.google.com/scholar?q=Neural+Scaling+Laws+Rooted+in+the+Data+Distribution
23. The Quantization Model of Neural Scaling — Eric J. Michaud et al., 2023
https://scholar.google.com/scholar?q=The+Quantization+Model+of+Neural+Scaling
24. AI Post Transformers: Muon Is Scalable for LLM Training — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-25-muon-is-scalable-for-llm-training-587ed8.mp3
25. AI Post Transformers: TMAS: Scaling Test-Time Compute with Multi-Agent Synergy — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-14-tmas-scaling-test-time-compute-with-mult-3abe7a.mp3
26. AI Post Transformers: Trajectory Summaries for Long-Horizon Coding Agents — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-24-trajectory-summaries-for-long-horizon-co-0194be.mp3
Interactive Visualization: Unified Neural Scaling Laws Across Regimes

This episode explores the RAGEN-2 paper’s claim that agentic reinforcement learning can produce reasoning traces that look active and diverse while losing real dependence on the input. It explains the paper’s central distinction between ordinary entropy, which measures diversity within a single prompt, and template collapse, where traces across many different prompts become generic variations of the same pattern. The discussion also covers the proposed mutual-information-style monitoring approach, which rescoring traces against other prompts to test whether reasoning remains identifiable to its source, and links the failure mode to weak reward signal, PPO-style regularization, and sparse long-horizon feedback. Listeners would find it interesting because it reframes a core question in reasoning RL: not whether an agent looks busy, but whether its reasoning is still actually about the problem in front of it.

Sources:
1. RAGEN-2: Reasoning Collapse in Agentic RL — Zihan Wang, Chi Gui, Xing Jin, Qineng Wang, Licheng Liu, Kangrui Wang, Shiqi Chen, Linjie Li, Zhengyuan Yang, Pingyue Zhang, Yiping Lu, Jiajun Wu, Li Fei-Fei, Lijuan Wang, Yejin Choi, Manling Li, 2026
http://arxiv.org/abs/2604.06268
2. MINE: Mutual Information Neural Estimation — Mohamed Ishmael Belghazi, Aristide Baratin, Sai Rajeswar, Sherjil Ozair, Yoshua Bengio, Aaron Courville, R. Devon Hjelm, 2018
https://scholar.google.com/scholar?q=MINE%3A+Mutual+Information+Neural+Estimation
3. Representation Learning with Contrastive Predictive Coding — Aaron van den Oord, Yazhe Li, Oriol Vinyals, 2018
https://scholar.google.com/scholar?q=Representation+Learning+with+Contrastive+Predictive+Coding
4. Learning deep representations by mutual information estimation and maximization — R. Devon Hjelm, Alex Fedorov, Samuel Lavoie-Marchildon, Phil Bachman, Adam Trischler, Yoshua Bengio, 2019
https://scholar.google.com/scholar?q=Learning+deep+representations+by+mutual+information+estimation+and+maximization
5. RAGEN-2: Reasoning Collapse in Agentic RL — Zihan Wang, Chi Gui, Xing Jin, Qineng Wang, Manling Li, et al., 2026
https://scholar.google.com/scholar?q=RAGEN-2%3A+Reasoning+Collapse+in+Agentic+RL
6. Proximal Policy Optimization Algorithms — John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, Oleg Klimov, 2017
https://scholar.google.com/scholar?q=Proximal+Policy+Optimization+Algorithms
7. Signal-to-Noise Ratio Analysis of Policy Gradient Algorithms — John W. Roberts, Russ Tedrake, 2008
https://scholar.google.com/scholar?q=Signal-to-Noise+Ratio+Analysis+of+Policy+Gradient+Algorithms
8. Understanding the Impact of Entropy on Policy Optimization — Zafarali Ahmed, Nicolas Le Roux, Mohammad Norouzi, Dale Schuurmans, 2019
https://scholar.google.com/scholar?q=Understanding+the+Impact+of+Entropy+on+Policy+Optimization
9. Prioritized Experience Replay — Tom Schaul, John Quan, Ioannis Antonoglou, David Silver, 2016
https://scholar.google.com/scholar?q=Prioritized+Experience+Replay
10. DAPO: An Open-Source LLM Reinforcement Learning System at Scale — Qiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan, Xiaochen Zuo, et al., 2025
https://scholar.google.com/scholar?q=DAPO%3A+An+Open-Source+LLM+Reinforcement+Learning+System+at+Scale
11. No Prompt Left Behind: Exploiting Zero-Variance Prompts in LLM Reinforcement Learning via Entropy-Guided Advantage Shaping — Thanh-Long V. Le, Myeongho Jeon, Kim Vu, Viet Lai, Eunho Yang, 2025
https://scholar.google.com/scholar?q=No+Prompt+Left+Behind%3A+Exploiting+Zero-Variance+Prompts+in+LLM+Reinforcement+Learning+via+Entropy-Guided+Advantage+Shaping
12. RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning — Zihan Wang et al., 2025
https://scholar.google.com/scholar?q=RAGEN%3A+Understanding+Self-Evolution+in+LLM+Agents+via+Multi-Turn+Reinforcement+Learning
13. The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models — Ganqu Cui et al., 2025
https://scholar.google.com/scholar?q=The+Entropy+Mechanism+of+Reinforcement+Learning+for+Reasoning+Language+Models
14. ASTER: Agentic Scaling with Tool-integrated Extended Reasoning — Xuqin Zhang, Quan He, Zhenrui Zheng, Zongzhang Zhang, Xu He, and Dong Li, 2026
https://scholar.google.com/scholar?q=ASTER%3A+Agentic+Scaling+with+Tool-integrated+Extended+Reasoning
15. Demystifying Reinforcement Learning in Agentic Reasoning — Zhaochen Yu, Ling Yang, Jiaru Zou, Shuicheng Yan, and Mengdi Wang, 2025
https://scholar.google.com/scholar?q=Demystifying+Reinforcement+Learning+in+Agentic+Reasoning
16. The Price of Format: Diversity Collapse in LLMs — Longfei Yun, Chenyang An, Zilong Wang, Letian Peng, and Jingbo Shang, 2025
https://scholar.google.com/scholar?q=The+Price+of+Format%3A+Diversity+Collapse+in+LLMs
17. Revisiting Entropy in Reinforcement Learning for Large Reasoning Models — Renren Jin et al., 2025
https://scholar.google.com/scholar?q=Revisiting+Entropy+in+Reinforcement+Learning+for+Large+Reasoning+Models
18. ERPO: Token-Level Entropy-Regulated Policy Optimization for Large Reasoning Models — Song Yu and Li Li, 2026
https://scholar.google.com/scholar?q=ERPO%3A+Token-Level+Entropy-Regulated+Policy+Optimization+for+Large+Reasoning+Models
19. Demystifying Reasoning Dynamics with Mutual Information: Thinking Tokens are Information Peaks in LLM Reasoning — Chen Qian et al., 2025
https://scholar.google.com/scholar?q=Demystifying+Reasoning+Dynamics+with+Mutual+Information%3A+Thinking+Tokens+are+Information+Peaks+in+LLM+Reasoning
20. MITS: Enhanced Tree Search Reasoning for LLMs via Pointwise Mutual Information — Jiaxi Li et al., 2025
https://scholar.google.com/scholar?q=MITS%3A+Enhanced+Tree+Search+Reasoning+for+LLMs+via+Pointwise+Mutual+Information
21. On the Design of KL-Regularized Policy Gradient Algorithms for LLM Reasoning — Yifan Zhang et al., 2025
https://scholar.google.com/scholar?q=On+the+Design+of+KL-Regularized+Policy+Gradient+Algorithms+for+LLM+Reasoning
22. Rethinking KL Regularization in RLHF: From Value Estimation to Gradient Optimization — Kezhao Liu et al., 2025
https://scholar.google.com/scholar?q=Rethinking+KL+Regularization+in+RLHF%3A+From+Value+Estimation+to+Gradient+Optimization
23. What Makes a Reward Model a Good Teacher? An Optimization Perspective — Noam Razin et al., 2025
https://scholar.google.com/scholar?q=What+Makes+a+Reward+Model+a+Good+Teacher%3F+An+Optimization+Perspective
24. Accelerating RLHF Training with Reward Variance Increase — Zonglin Yang et al., 2025
https://scholar.google.com/scholar?q=Accelerating+RLHF+Training+with+Reward+Variance+Increase
25. Efficient RLVR Training via Weighted Mutual Information Data Selection — Xinyu Zhou et al., 2026
https://scholar.google.com/scholar?q=Efficient+RLVR+Training+via+Weighted+Mutual+Information+Data+Selection
26. AI Post Transformers: Stabilizing Efficient Reasoning with Step-Level Advantage Selection — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-29-stabilizing-efficient-reasoning-with-ste-1e589d.mp3
27. AI Post Transformers: Reasoning Theater and Unfaithful Chain-of-Thought — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-reasoning-theater-and-unfaithful-chain-o-a4507e.mp3
28. AI Post Transformers: Benchmarking Test-Time Scaling for General LLM Agents — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-22-benchmarking-test-time-scaling-for-gener-8f14f9.mp3
29. AI Post Transformers: Can Models Learn from Long Context? — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-can-models-learn-from-long-context-77533e.mp3

This episode explores Latent Reasoning with Normalizing Flows, a paper that asks whether a standard left-to-right transformer can do its intermediate reasoning in continuous latent states instead of spelling every step out as text. It explains how the method uses a frozen VAE during training to compress written rationales into short latent sequences, then uses shallow normalizing flows so the same autoregressive backbone can predict both latent thought slots and normal answer tokens while preserving exact likelihoods, sampling, and KV-cache-friendly decoding. The discussion highlights matched coding results on Qwen3-8B-Base, where the reported benchmark average rises from 55.8 for the base model to 68.8 for NF-CoT Unified and 70.1 after latent-space reinforcement learning, with strong pass@k gains that suggest better exploration of multiple solution paths. Listeners would find it interesting because it frames latent reasoning as a practical alternative to verbose chain-of-thought, while also noting the current evidence is still narrow, centered on one post-trained coding model and not uniformly better than diffusion baselines on every benchmark.

Interactive Visualization: Latent Reasoning with Normalizing Flows
Sources:
1. Latent Reasoning with Normalizing Flows — Guancheng Tu, Xiangjun Fu, Suhao Yu, Yao Tang, Haoqiang Kang, Lianhui Qin, Yizhe Zhang, Jiatao Gu, 2026
http://arxiv.org/abs/2606.06447
2. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models — Jason Wei, Xuezhi Wang, Denny Zhou, Quoc Le, et al., 2022
https://scholar.google.com/scholar?q=Chain-of-Thought+Prompting+Elicits+Reasoning+in+Large+Language+Models
3. Training Large Language Models to Reason in a Continuous Latent Space — Shibo Hao, Sainbayar Sukhbaatar, Zhiting Hu, Jason Weston, Yuandong Tian, et al., 2024 preprint; COLM 2025
https://scholar.google.com/scholar?q=Training+Large+Language+Models+to+Reason+in+a+Continuous+Latent+Space
4. Reasoning Beyond Language: A Comprehensive Survey on Latent Chain-of-Thought Reasoning — Xinghao Chen, Anhao Zhao, Xiaoyu Shen, et al., 2025
https://scholar.google.com/scholar?q=Reasoning+Beyond+Language%3A+A+Comprehensive+Survey+on+Latent+Chain-of-Thought+Reasoning
5. LaDiR: Latent Diffusion Enhances LLMs for Text Reasoning — Haoqiang Kang, Yizhe Zhang, Navdeep Jaitly, Yi-An Ma, Lianhui Qin, et al., 2025
https://scholar.google.com/scholar?q=LaDiR%3A+Latent+Diffusion+Enhances+LLMs+for+Text+Reasoning
6. CODI: Compressing Chain-of-Thought into Continuous Space via Self-Distillation — Zhenyi Shen, Hanqi Yan, Linhai Zhang, Zhanghao Hu, Yali Du, Yulan He, 2025
https://scholar.google.com/scholar?q=CODI%3A+Compressing+Chain-of-Thought+into+Continuous+Space+via+Self-Distillation
7. Normalizing Flows are Capable Generative Models — Shuangfei Zhai, Ruixiang Zhang, Preetum Nakkiran, David Berthelot, Jiatao Gu, Huangjie Zheng, Tianrong Chen, Miguel Angel Bautista, Navdeep Jaitly, Josh Susskind, 2024
https://scholar.google.com/scholar?q=Normalizing+Flows+are+Capable+Generative+Models
8. Reasoning by Superposition: A Theoretical Perspective on Chain of Continuous Thought — Hanlin Zhu, Shibo Hao, Zhiting Hu, Jiantao Jiao, Stuart Russell, Yuandong Tian, 2025
https://scholar.google.com/scholar?q=Reasoning+by+Superposition%3A+A+Theoretical+Perspective+on+Chain+of+Continuous+Thought
9. Dynamics Within Latent Chain-of-Thought: An Empirical Study of Causal Structure — Zirui Li, Xuefeng Bai, Kehai Chen, Yizhi Li, Jian Yang, Chenghua Lin, Min Zhang, 2026
https://scholar.google.com/scholar?q=Dynamics+Within+Latent+Chain-of-Thought%3A+An+Empirical+Study+of+Causal+Structure
10. Latent-GRPO: Group Relative Policy Optimization for Latent Reasoning — Jingcheng Deng, Zihao Wei, Liang Pang, Junhong Wu, Shicheng Xu, Zenghao Duan, Huawei Shen, 2026
https://scholar.google.com/scholar?q=Latent-GRPO%3A+Group+Relative+Policy+Optimization+for+Latent+Reasoning
11. Chain-of-Thought Matters: Improving Long-Context Language Models with Reasoning Path Supervision — Dawei Zhu et al., 2025
https://arxiv.org/abs/2502.20790
12. Supervised Chain of Thought — Xiang Zhang and Dujian Ding, 2024
https://arxiv.org/abs/2410.14198
13. Large language models can learn and generalize steganographic chain-of-thought under process supervision — Joey Skaf et al., 2025
https://arxiv.org/abs/2506.01926
14. Hybrid Latent Reasoning via Reinforcement Learning — Zhenrui Yue et al., 2025
https://arxiv.org/abs/2505.18454
15. Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach — Jonas Geiping et al., 2025
https://arxiv.org/abs/2502.05171
16. R-KV: Redundancy-aware KV Cache Compression for Reasoning Models — Zefan Cai et al., 2025
https://arxiv.org/abs/2505.24133
17. Not All Heads Matter: A Head-Level KV Cache Compression Method with Integrated Retrieval and Reasoning — Yu Fu et al., 2024
https://arxiv.org/abs/2410.19258
18. AI Post Transformers: Generative Recursive Reasoning in Latent Space — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-21-generative-recursive-reasoning-in-latent-a9371d.mp3
19. AI Post Transformers: MELT: Decoupling Compute From Memory — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-13-melt-decoupling-compute-from-memory-26430c.mp3
20. AI Post Transformers: Reasoning Theater and Unfaithful Chain-of-Thought — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-reasoning-theater-and-unfaithful-chain-o-a4507e.mp3
21. AI Post Transformers: Gradient Descent at Inference Time for LLM Reasoning — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-10-gradient-descent-at-inference-time-for-l-20617d.mp3
22. AI Post Transformers: Explicit Information Transmission for Context Compression — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-explicit-information-transmission-for-co-24e3c2.mp3
23. AI Post Transformers: Speculative Decoding in Real vLLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-speculative-decoding-in-real-vllm-servin-6f4e2b.mp3
Interactive Visualization: Latent Reasoning with Normalizing Flows

This episode explores EMO, a Mixture-of-Experts language model that tries to turn token-level sparsity into real modularity by letting coherent expert groups emerge from document structure during pretraining. It explains why standard MoEs are not automatically deployable as smaller task-specific slices, then walks through EMO’s main idea: each document routes tokens through a learned document-specific pool of experts rather than the full expert set, with global load balancing to keep training stable. The discussion highlights that EMO was trained at substantial scale and reportedly matches a conventional MoE as a full model, while retaining surprisingly strong performance when only a fraction of experts are kept in memory for a given task. A listener would find it interesting because it connects a concrete systems problem, serving large models under tight memory budgets, to a plausible path toward more reusable, domain-specialized LLM components.

Sources:
1. EMO: Emergent Modularity in Sparse Language Models
https://allenai.org/papers/emo
2. Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer — Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, Jeff Dean, 2017
https://scholar.google.com/scholar?q=Outrageously+Large+Neural+Networks%3A+The+Sparsely-Gated+Mixture-of-Experts+Layer
3. A Review of Sparse Expert Models in Deep Learning — William Fedus, Jeff Dean, Barret Zoph, 2022
https://scholar.google.com/scholar?q=A+Review+of+Sparse+Expert+Models+in+Deep+Learning
4. ST-MoE: Designing Stable and Transferable Sparse Expert Models — Barret Zoph, Irwan Bello, Sameer Kumar, Nan Du, Yanping Huang, Jeff Dean, Noam Shazeer, William Fedus, 2022
https://scholar.google.com/scholar?q=ST-MoE%3A+Designing+Stable+and+Transferable+Sparse+Expert+Models
5. DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models — Damai Dai, Chengqi Deng, Chenggang Zhao, Huazuo Gao, Deli Chen, Jiashi Li, Chong Ruan, Zhifang Sui, Wenfeng Liang, 2024
https://scholar.google.com/scholar?q=DeepSeekMoE%3A+Towards+Ultimate+Expert+Specialization+in+Mixture-of-Experts+Language+Models
6. Neural Module Networks — Jacob Andreas, Marcus Rohrbach, Trevor Darrell, Dan Klein, 2015
https://scholar.google.com/scholar?q=Neural+Module+Networks
7. AdapterFusion: Non-Destructive Task Composition for Transfer Learning — Jonas Pfeiffer, Aishwarya Kamath, Andreas Rucklé, Kyunghyun Cho, Iryna Gurevych, 2021
https://scholar.google.com/scholar?q=AdapterFusion%3A+Non-Destructive+Task+Composition+for+Transfer+Learning
8. Sparsely Activated Mixture-of-Experts are Robust Multi-Task Learners — Shashank Gupta, Subhabrata Mukherjee, Krishan Subudhi, Eduardo Gonzalez, Damien Jose, Ahmed H. Awadallah, Jianfeng Gao, 2022
https://scholar.google.com/scholar?q=Sparsely+Activated+Mixture-of-Experts+are+Robust+Multi-Task+Learners
9. GLaM: Efficient Scaling of Language Models with Mixture-of-Experts — Nan Du, Yanping Huang, Andrew Dai, Sebastian Goodman, Orhan Firat, Quoc Le, Yonghui Wu, Zhifeng Chen, Claire Cui, et al., 2021
https://scholar.google.com/scholar?q=GLaM%3A+Efficient+Scaling+of+Language+Models+with+Mixture-of-Experts
10. OLMoE: Open Mixture-of-Experts Language Models — Niklas Muennighoff et al., 2024
https://scholar.google.com/scholar?q=OLMoE%3A+Open+Mixture-of-Experts+Language+Models
11. Branch-Train-MiX: Mixing Expert LLMs into a Mixture-of-Experts LLM — Sainbayar Sukhbaatar et al., 2024
https://scholar.google.com/scholar?q=Branch-Train-MiX%3A+Mixing+Expert+LLMs+into+a+Mixture-of-Experts+LLM
12. FlexOlmo: Open Language Models for Flexible Data Use — Weijia Shi et al., 2025
https://scholar.google.com/scholar?q=FlexOlmo%3A+Open+Language+Models+for+Flexible+Data+Use
13. Emergent Modularity in Pre-trained Transformers — Zhengyan Zhang et al., 2023
https://scholar.google.com/scholar?q=Emergent+Modularity+in+Pre-trained+Transformers
14. Domain-Specific Pruning of Large Mixture-of-Experts Models with Few-shot Demonstrations — Zican Dong et al., 2025
https://scholar.google.com/scholar?q=Domain-Specific+Pruning+of+Large+Mixture-of-Experts+Models+with+Few-shot+Demonstrations
15. Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models — Xudong Lu et al., 2024
https://arxiv.org/abs/2402.14800
16. Efficient Expert Pruning for Sparse Mixture-of-Experts Language Models: Enhancing Performance and Reducing Inference Costs — Enshu Liu et al., 2024
https://arxiv.org/abs/2407.00945
17. The Myth of Expert Specialization in MoEs: Why Routing Reflects Geometry, Not Necessarily Domain Expertise — Xi Wang, Soufiane Hayou, Eric Nalisnick, 2026
https://arxiv.org/abs/2604.09780
18. Dynamic Expert Specialization: Towards Catastrophic Forgetting-Free Multi-Domain MoE Adaptation — Junzhuo Li et al., 2025
https://arxiv.org/abs/2509.16882
19. AI Post Transformers: Batch-Aware Expert Routing for Faster MoE Decoding — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-batch-aware-expert-routing-for-faster-mo-683ab6.mp3
20. AI Post Transformers: JANUS for Scalable MoE Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-janus-for-scalable-moe-inference-78ae30.mp3
21. AI Post Transformers: Nemotron 3 Super Hybrid Mamba-Transformer MoE — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-19-nemotron-3-super-hybrid-mamba-transforme-31ac75.mp3

This episode explores Marina Favaro et al.’s 2026 paper on whether AI is beginning to accelerate frontier AI development enough to hint at recursive self-improvement, laying out the ladder from chatbots and coding agents to systems that materially help build their own successors. It distinguishes progress on long-horizon, fixed-goal engineering tasks from the harder problem of genuine research judgment, using public evidence such as METR, CORE-Bench, and RE-Bench to argue that AI is clearly getting better at sustained execution but has not yet shown strong scientific taste or reliable autonomy. It then digs into Anthropic’s internal evidence, including claims that by May 2026 Claude was responsible for over 80% of merged production code, open-ended coding-task success had risen sharply, and fixed-goal research engineering tasks improved from roughly 3x to 52x over a year, while the speakers repeatedly stress that code volume and self-reported productivity overstate true impact. Listeners would find it interesting because the discussion ties concrete benchmark results and lab productivity numbers to the bigger question of whether faster AI development also compresses the timelines for safety, governance, and the arrival of more capable successor systems.

Sources:
1. When AI Builds Itself and Recursive Self-Improvement
https://www.anthropic.com/institute/recursive-self-improvement
2. Measuring AI Ability to Complete Long Tasks — Thomas Kwa, Ben West, Joel Becker, Amy Deng, et al., 2025
https://scholar.google.com/scholar?q=Measuring+AI+Ability+to+Complete+Long+Tasks
3. RE-Bench: Evaluating Frontier AI R&D Capabilities of Language Model Agents against Human Experts — Hjalmar Wijk, Tao Roa Lin, Joel Becker, Sami Jawhar, et al., 2025
https://scholar.google.com/scholar?q=RE-Bench%3A+Evaluating+Frontier+AI+R%26D+Capabilities+of+Language+Model+Agents+against+Human+Experts
4. CORE-Bench: Fostering the Credibility of Published Research Through a Computational Reproducibility Agent Benchmark — Zachary S. Siegel, Sayash Kapoor, Nitya Nadgir, Benedikt Stroebl, Arvind Narayanan, 2024
https://scholar.google.com/scholar?q=CORE-Bench%3A+Fostering+the+Credibility+of+Published+Research+Through+a+Computational+Reproducibility+Agent+Benchmark
5. Automated Weak-to-Strong Researcher — Jiaxin Wen, Liang Qiu, Joe Benton, Jan Hendrik Kirchner, Jan Leike, 2026
https://scholar.google.com/scholar?q=Automated+Weak-to-Strong+Researcher
6. PaperBench: Evaluating AI's Ability to Replicate AI Research — Giulio Starace, Oliver Jaffe, Dane Sherburn, James Aung, et al., 2025
https://scholar.google.com/scholar?q=PaperBench%3A+Evaluating+AI%27s+Ability+to+Replicate+AI+Research
7. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity — Joel Becker, Nate Rush, Elizabeth Barnes, David Rein, 2025
https://scholar.google.com/scholar?q=Measuring+the+Impact+of+Early-2025+AI+on+Experienced+Open-Source+Developer+Productivity
8. SWE-bench Goes Live! — Linghao Zhang et al., 2025
https://scholar.google.com/scholar?q=SWE-bench+Goes+Live%21
9. Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving — Daoguang Zan et al., 2025
https://scholar.google.com/scholar?q=Multi-SWE-bench%3A+A+Multilingual+Benchmark+for+Issue+Resolving
10. SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains? — John Yang et al., 2024
https://scholar.google.com/scholar?q=SWE-bench+Multimodal%3A+Do+AI+Systems+Generalize+to+Visual+Software+Domains%3F
11. Can AI Conduct Autonomous Scientific Research? Case Studies on Two Real-World Tasks — S. Agrawal et al., 2026
https://scholar.google.com/scholar?q=Can+AI+Conduct+Autonomous+Scientific+Research%3F+Case+Studies+on+Two+Real-World+Tasks
12. Collapse of Self-trained Language Models — David Herel and Tomas Mikolov, 2024
https://scholar.google.com/scholar?q=Collapse+of+Self-trained+Language+Models
13. Exploring Automation Bias in Human-AI Collaboration: A Review and Implications for Explainable AI — G. Romeo and D. Conti, 2025
https://scholar.google.com/scholar?q=Exploring+Automation+Bias+in+Human-AI+Collaboration%3A+A+Review+and+Implications+for+Explainable+AI
14. AI Post Transformers: ASI-Evolve for Data, Architectures, and RL — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-05-asi-evolve-for-data-architectures-and-rl-197b2b.mp3
15. AI Post Transformers: Trajectory Summaries for Long-Horizon Coding Agents — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-24-trajectory-summaries-for-long-horizon-co-0194be.mp3
16. AI Post Transformers: TMAS: Scaling Test-Time Compute with Multi-Agent Synergy — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-14-tmas-scaling-test-time-compute-with-mult-3abe7a.mp3
17. AI Post Transformers: Air Force One, Jensen Huang, and Anthropic's 2028 Memo — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-2028-scenarios-for-global-ai-leadership-d0ec29.mp3
Interactive Visualization: When AI Builds Itself and Recursive Self-Improvement

This episode explores Mooncake, a production LLM serving architecture that treats KV cache reuse and movement as the central challenge in long-context chat, not just raw GPU compute. It explains why prefill and decode stress hardware in different ways, how metrics like time to first token and time between tokens drive system design, and why separating those phases helps meet real latency targets. The discussion walks through Mooncake’s cache-first scheduler, tiered KV storage across GPU memory, CPU DRAM, and SSD, and its use of RDMA, chunked prefill, and layer-wise overlap to start decoding sooner while reusing existing state. It also argues that Mooncake’s interest lies less in a single breakthrough than in how it combines prefix-aware routing, overload-aware early rejection, and cross-node KV reuse into a practical serving stack for large-scale chat systems.

Sources:
1. Mooncake for KV Cache-Centric LLM Serving
https://arxiv.org/pdf/2407.00079
2. Efficient Memory Management for Large Language Model Serving with PagedAttention — Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Joseph E. Gonzalez, Hao Zhang, Ion Stoica, et al., 2023
https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention
3. SGLang: Efficient Execution of Structured Language Model Programs — Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun, Jeff Huang, Cody Hao Yu, Ying Sheng, et al., 2023
https://scholar.google.com/scholar?q=SGLang%3A+Efficient+Execution+of+Structured+Language+Model+Programs
4. DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving — Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, Hao Zhang, 2024
https://scholar.google.com/scholar?q=DistServe%3A+Disaggregating+Prefill+and+Decoding+for+Goodput-optimized+Large+Language+Model+Serving
5. Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving — Ruoyu Qin, Zheming Li, Weiran He, Mingxing Zhang, Yongwei Wu, Weimin Zheng, Xinran Xu, 2024
https://scholar.google.com/scholar?q=Mooncake%3A+A+KVCache-centric+Disaggregated+Architecture+for+LLM+Serving
6. Overload Control for Scaling WeChat Microservices — Hao Zhou, Ming Chen, Qian Lin, Yong Wang, Xiaobin She, Sifan Liu, Rui Gu, Beng Chin Ooi, Junfeng Yang, 2018
https://scholar.google.com/scholar?q=Overload+Control+for+Scaling+WeChat+Microservices
7. Overload Control for microsecond-scale RPCs with Breakwater — Inho Cho, Ahmed Saeed, Joshua Fried, Seo Jin Park, Mohammad Alizadeh, Adam Belay, 2020
https://scholar.google.com/scholar?q=Overload+Control+for+microsecond-scale+RPCs+with+Breakwater
8. SCORPIO: Serving the Right Requests at the Right Time for Heterogeneous SLOs in LLM Inference — Yinghao Tang, Tingfeng Lan, Xiuqi Huang, Hui Lu, Wei Chen, 2025
https://scholar.google.com/scholar?q=SCORPIO%3A+Serving+the+Right+Requests+at+the+Right+Time+for+Heterogeneous+SLOs+in+LLM+Inference
9. Splitwise: Efficient Generative LLM Inference Using Phase Splitting — Pratyush Patel, Esha Choukse, Chaojie Zhang, Inigo Goiri, Aashaka Shah, Saeed Maleki, Ricardo Bianchini, 2023
https://scholar.google.com/scholar?q=Splitwise%3A+Efficient+Generative+LLM+Inference+Using+Phase+Splitting
10. AttentionStore: Cost-Effective Attention Reuse Across Multi-Turn Conversations in Large Language Model Serving — Bin Gao, Zhuomin He, Puru Sharma, Qingxuan Kang, Djordje Jevdjic, Junbo Deng, Xingkun Yang, Zhou Yu, Pengfei Zuo, 2024
https://scholar.google.com/scholar?q=AttentionStore%3A+Cost-Effective+Attention+Reuse+Across+Multi-Turn+Conversations+in+Large+Language+Model+Serving
11. Preble: Efficient Distributed Prompt Scheduling for LLM Serving — Vikranth Srivatsa, Zijian He, Reyna Abhyankar, Dongming Li, Yiying Zhang, 2024
https://scholar.google.com/scholar?q=Preble%3A+Efficient+Distributed+Prompt+Scheduling+for+LLM+Serving
12. CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowledge Fusion — Jiayi Yao, Hanchen Li, Yuhan Liu, Siddhant Ray, Yihua Cheng, Qizheng Zhang, Kuntai Du, Shan Lu, Junchen Jiang, 2024
https://scholar.google.com/scholar?q=CacheBlend%3A+Fast+Large+Language+Model+Serving+for+RAG+with+Cached+Knowledge+Fusion
13. P/D-Serve: Serving Disaggregated Large Language Model at Scale — Yibo Jin et al., 2024
https://scholar.google.com/scholar?q=P%2FD-Serve%3A+Serving+Disaggregated+Large+Language+Model+at+Scale
14. POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM Inference — Aditya K. Kamath et al., 2025
https://scholar.google.com/scholar?q=POD-Attention%3A+Unlocking+Full+Prefill-Decode+Overlap+for+Faster+LLM+Inference
15. Prepacking: A Simple Method for Fast Prefilling and Increased Throughput in Large Language Models — Siyan Zhao et al., 2024
https://scholar.google.com/scholar?q=Prepacking%3A+A+Simple+Method+for+Fast+Prefilling+and+Increased+Throughput+in+Large+Language+Models
16. Mustafar: Promoting Unstructured Sparsity for KV Cache Pruning in LLM Inference — Donghyeon Joo et al., 2025
https://scholar.google.com/scholar?q=Mustafar%3A+Promoting+Unstructured+Sparsity+for+KV+Cache+Pruning+in+LLM+Inference
17. SparK: Query-Aware Unstructured Sparsity with Recoverable KV Cache Channel Pruning — Huanxuan Liao et al., 2025/2026
https://scholar.google.com/scholar?q=SparK%3A+Query-Aware+Unstructured+Sparsity+with+Recoverable+KV+Cache+Channel+Pruning
18. More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression — Jiebin Zhang et al., 2024
https://scholar.google.com/scholar?q=More+Tokens%2C+Lower+Precision%3A+Towards+the+Optimal+Token-Precision+Trade-off+in+KV+Cache+Compression
19. KVShare: An LLM Service System with Efficient and Effective Multi-Tenant KV Cache Reuse — Huan Yang et al., 2025
https://scholar.google.com/scholar?q=KVShare%3A+An+LLM+Service+System+with+Efficient+and+Effective+Multi-Tenant+KV+Cache+Reuse
20. CacheSolidarity: Preventing Prefix Caching Side Channels in Multi-tenant LLM Serving Systems — Panagiotis Georgios Pennas et al., 2026
https://scholar.google.com/scholar?q=CacheSolidarity%3A+Preventing+Prefix+Caching+Side+Channels+in+Multi-tenant+LLM+Serving+Systems
21. AI Post Transformers: Prefill-as-a-Service for Cross-Datacenter KV Cache — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-19-prefill-as-a-service-for-cross-datacente-7560be.mp3
22. AI Post Transformers: Characterizing LLM KV Cache Workloads in Production — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/characterizing-llm-kv-cache-workloads-in-production/
23. AI Post Transformers: ContiguousKV for Faster LLM Prefill KV Reuse — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-20-contiguouskv-for-faster-llm-prefill-kv-r-59f545.mp3
24. AI Post Transformers: CacheFlow and 3D-Parallel KV Cache Restoration — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-01-cacheflow-and-3d-parallel-kv-cache-resto-8db883.mp3
25. AI Post Transformers: Beluga: CXL Memory Pooling for LLM KV Cache — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-27-beluga-cxl-memory-pooling-for-llm-kv-cac-b6142f.mp3
26. AI Post Transformers: Speculative Decoding in Real vLLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-speculative-decoding-in-real-vllm-servin-6f4e2b.mp3
Interactive Visualization: Mooncake for KV Cache-Centric LLM Serving

This episode explores DeepMind’s paper on technical AGI safety and security, focusing on how labs might prevent severe, humanity-scale harm before highly capable systems are deployed. It breaks down the paper’s core distinctions between misuse and misalignment, explains what the authors mean by Exceptional AGI and the no-human-ceiling assumption, and examines dangerous capability evaluations in areas like cyber, biology, persuasion, and self-proliferation. The discussion highlights the paper’s main argument that safety measures such as refusal training, jailbreak hardening, access controls, monitoring, anomaly detection, and model-weight security only matter if they are explicitly tied to capability thresholds that trigger real deployment restrictions. Listeners would find it interesting because it turns abstract AGI risk debates into a concrete governance and engineering framework for deciding when a model is too dangerous to release under normal conditions.

Sources:
1. Technical AGI Safety and Security Framework
https://arxiv.org/pdf/2504.01849
2. The Malicious Use of Artificial Intelligence: Forecasting, Prevention, and Mitigation — Miles Brundage, Shahar Avin, Jack Clark, Helen Toner, et al., 2018
https://scholar.google.com/scholar?q=The+Malicious+Use+of+Artificial+Intelligence%3A+Forecasting%2C+Prevention%2C+and+Mitigation
3. Model evaluation for extreme risks — Toby Shevlane, Sebastian Farquhar, Ben Garfinkel, Mary Phuong, et al., 2023
https://scholar.google.com/scholar?q=Model+evaluation+for+extreme+risks
4. Frontier AI Regulation: Managing Emerging Risks to Public Safety — Markus Anderljung, Joslyn Barnhart, Anton Korinek, Jade Leung, et al., 2023
https://scholar.google.com/scholar?q=Frontier+AI+Regulation%3A+Managing+Emerging+Risks+to+Public+Safety
5. Evaluating Frontier Models for Dangerous Capabilities — Mary Phuong, Matthew Aitchison, Elliot Catt, Victoria Krakovna, et al., 2024
https://scholar.google.com/scholar?q=Evaluating+Frontier+Models+for+Dangerous+Capabilities
6. Guidance on the Assurance of Machine Learning in Autonomous Systems (AMLAS) — Richard Hawkins, Colin Paterson, Chiara Picardi, Ibrahim Habli, et al., 2021
https://scholar.google.com/scholar?q=Guidance+on+the+Assurance+of+Machine+Learning+in+Autonomous+Systems+%28AMLAS%29
7. Safety Cases: How to Justify the Safety of Advanced AI Systems — Joshua Clymer, Nick Gabrieli, David Krueger, Thomas Larsen, 2024
https://scholar.google.com/scholar?q=Safety+Cases%3A+How+to+Justify+the+Safety+of+Advanced+AI+Systems
8. Safety case template for frontier AI: A cyber inability argument — Arthur Goemans, Marie Davidsen Buhl, Jonas Schuett, Geoffrey Irving, et al., 2024
https://scholar.google.com/scholar?q=Safety+case+template+for+frontier+AI%3A+A+cyber+inability+argument
9. The BIG Argument for AI Safety Cases — Ibrahim Habli, Richard Hawkins, Colin Paterson, Mark Sujan, et al., 2025
https://scholar.google.com/scholar?q=The+BIG+Argument+for+AI+Safety+Cases
10. Safety cases: Justifying the safety of advanced AI systems — J. Clymer, N. Gabrieli, D. Krueger, and T. Larsen, 2024
https://scholar.google.com/scholar?q=Safety+cases%3A+Justifying+the+safety+of+advanced+AI+systems
11. AI control: Improving safety despite intentional subversion — R. Greenblatt, B. Shlegeris, K. Sachan, and F. Roger, 2024
https://scholar.google.com/scholar?q=AI+control%3A+Improving+safety+despite+intentional+subversion
12. Alignment faking in large language models — R. Greenblatt, C. Denison, B. Wright, F. Roger, M. MacDiarmid, S. Marks, J. Treutlein, T. Belonax, J. Chen, D. Duvenaud, et al., 2024
https://scholar.google.com/scholar?q=Alignment+faking+in+large+language+models
13. Towards evaluations-based safety cases for AI scheming — M. Balesni, M. Hobbhahn, D. Lindner, A. Meinke, T. Korbak, J. Clymer, B. Shlegeris, J. Scheurer, C. Stix, R. Shah, et al., 2024
https://scholar.google.com/scholar?q=Towards+evaluations-based+safety+cases+for+AI+scheming
14. Generative AI misuse: A taxonomy of tactics and insights from real-world data — N. Marchal, R. Xu, R. Elasmar, I. Gabriel, B. Goldberg, and W. Isaac, 2024
https://scholar.google.com/scholar?q=Generative+AI+misuse%3A+A+taxonomy+of+tactics+and+insights+from+real-world+data
15. Stress-Testing Capability Elicitation With Password-Locked Models — Ryan Greenblatt et al., 2024
https://scholar.google.com/scholar?q=Stress-Testing+Capability+Elicitation+With+Password-Locked+Models
16. Noise Injection Reveals Hidden Capabilities of Sandbagging Language Models — Cameron Tice et al., 2024
https://scholar.google.com/scholar?q=Noise+Injection+Reveals+Hidden+Capabilities+of+Sandbagging+Language+Models
17. Benchmarking Misuse Mitigation Against Covert Adversaries — Davis Brown et al., 2025
https://scholar.google.com/scholar?q=Benchmarking+Misuse+Mitigation+Against+Covert+Adversaries
18. On scalable oversight with weak LLMs judging strong LLMs — Zachary Kenton et al., 2024
https://scholar.google.com/scholar?q=On+scalable+oversight+with+weak+LLMs+judging+strong+LLMs
19. Scaling Laws For Scalable Oversight — Joshua Engels et al., 2025
https://scholar.google.com/scholar?q=Scaling+Laws+For+Scalable+Oversight
20. Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs — Kyle O'Brien et al., 2025
https://scholar.google.com/scholar?q=Deep+Ignorance%3A+Filtering+Pretraining+Data+Builds+Tamper-Resistant+Safeguards+into+Open-Weight+LLMs
21. AI Post Transformers: When LLM Judges Become Coin Flips — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-when-llm-judges-become-coin-flips-8b43ef.mp3
22. AI Post Transformers: Reasoning Theater and Unfaithful Chain-of-Thought — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-reasoning-theater-and-unfaithful-chain-o-a4507e.mp3
23. AI Post Transformers: Split Personality Training Reveals Latent Knowledge — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-08-split-personality-training-reveals-laten-c84616.mp3
24. AI Post Transformers: Trajectory Summaries for Long-Horizon Coding Agents — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-24-trajectory-summaries-for-long-horizon-co-0194be.mp3

This episode explores VeriCache, a systems paper that asks whether a large language model can draft tokens using a compressed, lossy KV cache and then verify them against the full cache to recover exactly the same greedy-decoding output. It explains why KV cache size has become a central bottleneck in long-context inference, unpacking token dropping, KV quantization, prefix caching, and the speculative decoding ideas that VeriCache turns into a closed-loop draft-and-verify scheme. The discussion argues that fluency is not enough for real deployments because even a single wrong token can break code, structured JSON, or tool calls, so exact token-by-token agreement is the real standard for safe acceleration. Listeners get a clear picture of where serving stacks such as vLLM, Hugging Face TGI, and NVIDIA already use neighboring optimizations, and why VeriCache’s specific lossless-verification approach is both technically appealing and operationally difficult.

Sources:
1. VeriCache: Turning Lossy KV Cache into Lossless LLM Inference — Jiayi Yao, Samuel Shen, Kuntai Du, Shaoting Feng, Dongjoo Seo, Rui Zhang, Yuyang Huang, Yuhan Liu, Shan Lu, Junchen Jiang, 2026
http://arxiv.org/abs/2605.17613
2. Fast Inference from Transformers via Speculative Decoding — Yaniv Leviathan, Matan Kalman, Yossi Matias, 2023
https://proceedings.mlr.press/v202/leviathan23a.html
3. H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models — Zhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen, Beidi Chen, et al., 2023
https://proceedings.neurips.cc/paper_files/paper/2023/hash/6ceefa7b15572587b78ecfcebb2827f8-Abstract-Conference.html
4. KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache — Zirui Liu, Jiayi Yuan, Hongye Jin, Shaochen Zhong, Beidi Chen, Xia Hu, et al., 2024
https://arxiv.org/abs/2402.02750
5. QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache — Rishabh Tiwari, Haocheng Xi, Aditya Tomar, Coleman Hooper, Kurt Keutzer, Amir Gholami, et al., 2025
https://arxiv.org/abs/2502.10424
6. MagicDec: Breaking the Latency-Throughput Tradeoff for Long Context Generation with Speculative Decoding — Ranajoy Sadhukhan, Jian Chen, Zhuoming Chen, et al., 2024
https://scholar.google.com/scholar?q=MagicDec%3A+Breaking+the+Latency-Throughput+Tradeoff+for+Long+Context+Generation+with+Speculative+Decoding
7. Accelerating Large-Scale Reasoning Model Inference: Self-Speculative Decoding with Sparse Attention — Yilong Zhao, Jiaming Tang, Kan Zhu, Zihao Ye, Chi-Chih Chang, Chaofan Lin, Jongseok Park, Guangxuan Xiao, Mohamed S. Abdelfattah, Mingyu Gao, Baris Kasikci, Song Han, and Ion Stoica, 2025
https://scholar.google.com/scholar?q=Accelerating+Large-Scale+Reasoning+Model+Inference%3A+Self-Speculative+Decoding+with+Sparse+Attention
8. CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving — Yuhan Liu, Hanchen Li, Yihua Cheng, Siddhant Ray, Yuyang Huang, Qizheng Zhang, Kuntai Du, Jiayi Yao, Shan Lu, Ganesh Ananthanarayanan, and Junchen Jiang, 2024
https://scholar.google.com/scholar?q=CacheGen%3A+KV+Cache+Compression+and+Streaming+for+Fast+Large+Language+Model+Serving
9. ShadowServe: Interference-Free KV Cache Fetching for Distributed Prefix Caching — Xingyu Xiang, Raj Joshi, Yuhan Liu, Jiayi Yao, Chenxingyu Zhao, Junchen Jiang, Yang Zhou, Eddie Kohler, and Minlan Yu, 2025
https://scholar.google.com/scholar?q=ShadowServe%3A+Interference-Free+KV+Cache+Fetching+for+Distributed+Prefix+Caching
10. LMCache: An Efficient KV Cache Layer for Enterprise-Scale LLM Inference — Yuhan Liu, Yihua Cheng, Jiayi Yao, Yuwei An, Xiaokun Chen, Shaoting Feng, et al., 2025
https://scholar.google.com/scholar?q=LMCache%3A+An+Efficient+KV+Cache+Layer+for+Enterprise-Scale+LLM+Inference
11. ARKV: Adaptive and Resource-Efficient KV Cache Management under Limited Memory Budget for Long-Context Inference in LLMs — Jianlong Lei and Shashikant Ilager, 2026
https://arxiv.org/abs/2603.08727
12. Don't Waste Bits! Adaptive KV-Cache Quantization for Lightweight On-Device LLMs — Sayed Pedram Haeri Boroujeni et al., 2026
https://arxiv.org/abs/2604.04722
13. LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification — Penghui Yang et al., 2025
https://arxiv.org/abs/2502.17421
14. RAPID: Long-Context Inference with Retrieval-Augmented Speculative Decoding — Guanzheng Chen et al., 2025
https://arxiv.org/abs/2502.20330
15. LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management — Yi Xiong et al., 2024
https://arxiv.org/abs/2410.00428
16. TokenLake: A Unified Segment-level Prefix Cache Pool for Fine-grained Elastic Long-Context LLM Serving — Bingyang Wu et al., 2025
https://arxiv.org/abs/2508.17219
17. KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows — Zaifeng Pan et al., 2025
https://arxiv.org/abs/2507.07400
18. KVShare: An LLM Service System with Efficient and Effective Multi-Tenant KV Cache Reuse — Huan Yang et al., 2025
https://arxiv.org/abs/2503.16525
19. KV Cache Offloading for Context-Intensive Tasks — Andrey Bocharnikov et al., 2026
https://arxiv.org/abs/2604.08426
20. AI Post Transformers: KVzip for Query-Agnostic KV Cache Compression — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-29-kvzip-for-query-agnostic-kv-cache-compre-72afe5.mp3
21. AI Post Transformers: PackKV Lossy Compression for KV Caches — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-04-packkv-lossy-compression-for-kv-caches-b37bce.mp3
22. AI Post Transformers: Speculative Decoding in Real vLLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-speculative-decoding-in-real-vllm-servin-6f4e2b.mp3
23. AI Post Transformers: Affordable Large-Scale Decoding Through Model-System Co-Design — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-19-affordable-large-scale-decoding-through-e1d7ed.mp3
24. AI Post Transformers: Stochastic KV Routing for Cache Sharing — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-29-stochastic-kv-routing-for-cache-sharing-5fef63.mp3

This episode explores SFMP, a post-training, weight-only quantization method for large language models that targets awkward memory budgets like 3.25 or 3.75 bits per weight, where standard 3-bit or 4-bit schemes are a poor fit. It explains the core ideas behind mixed-precision quantization, including salience scoring with diagonal Fisher information, block-wise allocation, and why hardware constraints make many fine-grained methods difficult to deploy in practice. The discussion centers on SFMP’s main claim: instead of running an expensive search over bit allocations, it ranks blocks by importance and assigns only the two neighboring integer precisions, such as giving the most salient 25 percent of blocks 4-bit weights and the rest 3-bit for a 3.25 BPW target. Listeners would find it interesting because it frames quantization not as a narrow accuracy trick, but as a concrete engineering problem about trading off model quality, memory limits, and real serving-system compatibility.

Sources:
1. SFMP Search-Free Mixed-Precision LLM Quantization
https://arxiv.org/pdf/2602.01027
2. HAWQ: Hessian AWare Quantization of Neural Networks with Mixed-Precision — Zhen Dong, Zhewei Yao, Amir Gholami, Michael Mahoney, Kurt Keutzer, 2019
https://scholar.google.com/scholar?q=HAWQ%3A+Hessian+AWare+Quantization+of+Neural+Networks+with+Mixed-Precision
3. GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers — Elias Frantar, Saleh Ashkboos, Torsten Hoefler, Dan Alistarh, 2023
https://scholar.google.com/scholar?q=GPTQ%3A+Accurate+Post-Training+Quantization+for+Generative+Pre-trained+Transformers
4. AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration — Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Wei-Ming Chen, Wei-Chen Wang, Guangxuan Xiao, Xingyu Dang, Chuang Gan, Song Han, 2024
https://scholar.google.com/scholar?q=AWQ%3A+Activation-aware+Weight+Quantization+for+LLM+Compression+and+Acceleration
5. AMQ: Enabling AutoML for Mixed-precision Weight-Only Quantization of Large Language Models — Sangjun Lee, Seung-taek Woo, Jungyu Jin, Changhun Lee, Eunhyeok Park, 2025
https://scholar.google.com/scholar?q=AMQ%3A+Enabling+AutoML+for+Mixed-precision+Weight-Only+Quantization+of+Large+Language+Models
6. BitStack: Fine-Grained Size Control for Compressed Large Language Models in Variable Memory Environments — Xinghao Wang, Pengyu Wang, Bo Wang, Dong Zhang, Yunhua Zhou, Xipeng Qiu, 2024
https://scholar.google.com/scholar?q=BitStack%3A+Fine-Grained+Size+Control+for+Compressed+Large+Language+Models+in+Variable+Memory+Environments
7. SliM-LLM: Salience-Driven Mixed-Precision Quantization for Large Language Models — Wei Huang, Haotong Qin, Yangdong Liu, Yawei Li, Xianglong Liu, Luca Benini, Michele Magno, Xiaojuan Qi, 2024
https://scholar.google.com/scholar?q=SliM-LLM%3A+Salience-Driven+Mixed-Precision+Quantization+for+Large+Language+Models
8. FineQ: Software-Hardware Co-Design for Low-Bit Fine-Grained Mixed-Precision Quantization of LLMs — Xilong Xie, Liang Wang, Limin Xiao, Meng Han, Lin Sun, Shuai Zheng, Xiangrong Xu, 2025
https://scholar.google.com/scholar?q=FineQ%3A+Software-Hardware+Co-Design+for+Low-Bit+Fine-Grained+Mixed-Precision+Quantization+of+LLMs
9. T-MAC: CPU Renaissance via Table Lookup for Low-Bit LLM Deployment on Edge — Jianyu Wei, Shijie Cao, Ting Cao, Lingxiao Ma, Lei Wang, Yanyong Zhang, Mao Yang, 2024
https://scholar.google.com/scholar?q=T-MAC%3A+CPU+Renaissance+via+Table+Lookup+for+Low-Bit+LLM+Deployment+on+Edge
10. MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design — Zhen Zheng et al., 2024
https://arxiv.org/abs/2412.14590
11. SpinQuant: LLM quantization with learned rotations — Zechun Liu et al., 2024
https://arxiv.org/abs/2405.16406
12. DartQuant: Efficient Rotational Distribution Calibration for LLM Quantization — Yuantian Shao et al., 2025
https://arxiv.org/abs/2511.04063
13. KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization — Coleman Hooper et al., 2024
https://arxiv.org/abs/2401.18079
14. QAQ: Quality Adaptive Quantization for LLM KV Cache — Shichen Dong et al., 2024
https://arxiv.org/abs/2403.04643
15. AI Post Transformers: Deep Kernel Fusion for Transformer Decoding — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-deep-kernel-fusion-for-transformer-decod-b1a703.mp3
16. AI Post Transformers: PackKV Lossy Compression for KV Caches — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-04-packkv-lossy-compression-for-kv-caches-b37bce.mp3
17. AI Post Transformers: Affordable Large-Scale Decoding Through Model-System Co-Design — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-19-affordable-large-scale-decoding-through-e1d7ed.mp3
18. AI Post Transformers: Speculative Decoding in Real vLLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-speculative-decoding-in-real-vllm-servin-6f4e2b.mp3

This episode explores Harvest, a system for LLM inference that uses idle HBM on neighboring NVLink-connected GPUs as a temporary cache when a serving GPU runs out of local memory. It explains why LLM serving is often bottlenecked more by memory capacity and data movement than by raw compute, focusing on two concrete cases: growing KV caches during long-context decoding and the shifting expert weights used in mixture-of-experts models. A key argument is that peer GPU memory is only useful if it is revocable without breaking correctness, so Harvest treats borrowed memory as a best-effort cache backed by authoritative copies or reconstruction paths elsewhere. Listeners get specific performance results, including up to 5.65x lower KV-cache transfer latency than CPU offload, 7.5x to 9.5x faster expert transfers over NVLink, and roughly 1.5x to 2.0x throughput gains on models such as Qwen and Phi-3.5.

Sources:
1. Harvest: Opportunistic Peer-to-Peer GPU Caching for LLM Inference — Nikhil Gopal, Kostis Kaffes, 2026
http://arxiv.org/abs/2602.00328
2. Efficient Memory Management for Large Language Model Serving with PagedAttention — Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Joseph E. Gonzalez, Hao Zhang, Ion Stoica, 2023
https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention
3. FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU — Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Beidi Chen, Percy Liang, Christopher Re, Ion Stoica, Ce Zhang, 2023
https://scholar.google.com/scholar?q=FlexGen%3A+High-Throughput+Generative+Inference+of+Large+Language+Models+with+a+Single+GPU
4. Fiddler: CPU-GPU Orchestration for Fast Inference of Mixture-of-Experts Models — Keisuke Kamahori, Tian Tang, Yile Gu, Kan Zhu, Baris Kasikci, 2025
https://scholar.google.com/scholar?q=Fiddler%3A+CPU-GPU+Orchestration+for+Fast+Inference+of+Mixture-of-Experts+Models
5. Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving — Ruoyu Qin, Zheming Li, Mingxing Zhang, Yongwei Wu, Weimin Zheng, Xinran Xu, et al., 2025
https://scholar.google.com/scholar?q=Mooncake%3A+A+KVCache-centric+Disaggregated+Architecture+for+LLM+Serving
6. AQUA: Network-Accelerated Memory Offloading for LLMs in Scale-Up GPU Domains — Abhishek Vijaya Kumar, Gianni Antichi, Rachee Singh, 2025
https://scholar.google.com/scholar?q=AQUA%3A+Network-Accelerated+Memory+Offloading+for+LLMs+in+Scale-Up+GPU+Domains
7. Accurate Expert Predictions in MoE Inference via Cross-Layer Gate — Zhiyuan Fang, Hong Huang, Yiming Lyu, Jiyang Chen, Yu Yu, Zexi Zheng, 2025
https://scholar.google.com/scholar?q=Accurate+Expert+Predictions+in+MoE+Inference+via+Cross-Layer+Gate
8. Characterization of Large Language Model Development in the Datacenter — Qinghao Hu et al., 2024
https://scholar.google.com/scholar?q=Characterization+of+Large+Language+Model+Development+in+the+Datacenter
9. MLaaS in the Wild: Workload Analysis and Scheduling in Large-Scale Heterogeneous GPU Clusters — Qizhen Weng et al., 2022
https://scholar.google.com/scholar?q=MLaaS+in+the+Wild%3A+Workload+Analysis+and+Scheduling+in+Large-Scale+Heterogeneous+GPU+Clusters
10. MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool — Cunchen Hu et al., 2024
https://arxiv.org/abs/2406.17565
11. KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows — Zaifeng Pan et al., 2025
https://arxiv.org/abs/2507.07400
12. Learned Prefix Caching for Efficient LLM Inference — Dongsheng Yang et al., 2025
https://papers.neurips.cc/paper_files/paper/2025/hash/414f642a1ea9350006669774cba9bcd4-Abstract-Conference.html
13. DuoServe-MoE: Dual-Phase Expert Prefetch and Caching for LLM Inference QoS Assurance — Yuning Zhang et al., 2025
https://arxiv.org/abs/2509.07379
14. ExpertFlow: Adaptive Expert Scheduling and Memory Coordination for Efficient MoE Inference — Zixu Shen et al., 2025
https://arxiv.org/abs/2510.26730
15. AdapMoE: Adaptive Sensitivity-based Expert Gating and Management for Efficient MoE Inference — Shuzhang Zhong et al., 2024
https://arxiv.org/abs/2408.10284
16. HyperAttention: Long-context Attention in Near-Linear Time — Insu Han et al., 2023
https://arxiv.org/abs/2310.05869
17. Beyond Attention: Breaking the Limits of Transformer Context Length with Recurrent Memory — Aydar Bulatov et al., 2024
https://ojs.aaai.org/index.php/AAAI/article/download/29722/31239
18. AI Post Transformers: ScoutAttention for Efficient KV Cache Offloading — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-24-scoutattention-for-efficient-kv-cache-of-b26699.mp3
19. AI Post Transformers: JANUS for Scalable MoE Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-janus-for-scalable-moe-inference-78ae30.mp3
20. AI Post Transformers: FengHuang for Rack-Scale LLM Inference Memory — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-12-fenghuang-for-rack-scale-llm-inference-m-62708e.mp3
21. AI Post Transformers: FlexGen: High-Throughput LLM Inference on a Single GPU — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/flexgen-high-throughput-llm-inference-on-a-single-gpu/
22. AI Post Transformers: Stochastic KV Routing for Cache Sharing — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-29-stochastic-kv-routing-for-cache-sharing-5fef63.mp3
23. AI Post Transformers: Affordable Large-Scale Decoding Through Model-System Co-Design — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-19-affordable-large-scale-decoding-through-e1d7ed.mp3

This episode explores a 2026 paper on whether language models can keep doing long, step-by-step reasoning internally or eventually need to expose some of that reasoning in visible chain-of-thought tokens. It explains the paper’s core idea of opaque serial depth, or a model’s hidden reasoning horizon, and argues that this is a better safety-relevant measure than raw model size, parameter count, or informal layer counting. The discussion connects that metric to circuit complexity, fixed-precision computation, and transformer internals, showing why models can perform huge amounts of parallel work in one pass yet still face structural limits on long private sequential reasoning. Listeners would find it interesting because it sharpens a major AI safety question: whether monitoring visible reasoning can meaningfully constrain powerful models, and where that hope may break down.

Sources:
1. Opaque Serial Depth and Chain-of-Thought Limits
https://arxiv.org/pdf/2603.09786
2. Quantifying the Necessity of Chain of Thought through Opaque Serial Depth — Jonah Brown-Cohen, David Lindner, Rohin Shah, 2026
https://scholar.google.com/scholar?q=Quantifying+the+Necessity+of+Chain+of+Thought+through+Opaque+Serial+Depth
3. Chain of Thought Empowers Transformers to Solve Inherently Serial Problems — Zhiyuan Li, Hong Liu, Denny Zhou, Tengyu Ma, 2024
https://scholar.google.com/scholar?q=Chain+of+Thought+Empowers+Transformers+to+Solve+Inherently+Serial+Problems
4. Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety — Tomek Korbak, Mikita Balesni, Elizabeth Barnes, Yoshua Bengio, Mark Chen, Rohin Shah, et al., 2025
https://scholar.google.com/scholar?q=Chain+of+Thought+Monitorability%3A+A+New+and+Fragile+Opportunity+for+AI+Safety
5. When Chain of Thought is Necessary, Language Models Struggle to Evade Monitors — Scott Emmons, Erik Jenner, David K. Elson, Rif A. Saurous, Senthooran Rajamanoharan, Heng Chen, Irhum Shafkat, Rohin Shah, 2025
https://scholar.google.com/scholar?q=When+Chain+of+Thought+is+Necessary%2C+Language+Models+Struggle+to+Evade+Monitors
6. Saturated Transformers are Constant-Depth Threshold Circuits — William Merrill, Ashish Sabharwal, Noah A. Smith, 2022
https://scholar.google.com/scholar?q=Saturated+Transformers+are+Constant-Depth+Threshold+Circuits
7. The Parallelism Tradeoff: Limitations of Log-Precision Transformers — William Merrill, Ashish Sabharwal, 2023
https://scholar.google.com/scholar?q=The+Parallelism+Tradeoff%3A+Limitations+of+Log-Precision+Transformers
8. The Expressive Power of Transformers with Chain of Thought — William Merrill, Ashish Sabharwal, 2024
https://scholar.google.com/scholar?q=The+Expressive+Power+of+Transformers+with+Chain+of+Thought
9. Theoretical Limitations of Self-Attention in Neural Sequence Models — Michael Hahn, 2020
https://scholar.google.com/scholar?q=Theoretical+Limitations+of+Self-Attention+in+Neural+Sequence+Models
10. Continuous Chain of Thought Enables Parallel Exploration and Reasoning — Halil Alperen Gozeten, M. Emrullah Ildiz, Xuechen Zhang, Hrayr Harutyunyan, Ankit Singh Rawat, Samet Oymak, 2025
https://scholar.google.com/scholar?q=Continuous+Chain+of+Thought+Enables+Parallel+Exploration+and+Reasoning
11. Reasoning Theater: Disentangling Model Beliefs from Chain-of-Thought — Siddharth Boppana, Annabel Ma, Max Loeffler, Raphael Sarfati, Eric Bigelow, Atticus Geiger, Owen Lewis, Jack Merullo, 2026
https://scholar.google.com/scholar?q=Reasoning+Theater%3A+Disentangling+Model+Beliefs+from+Chain-of-Thought
12. Measuring Chain of Thought Faithfulness by Unlearning Reasoning Steps — Martin Tutek et al., 2025
https://scholar.google.com/scholar?q=Measuring+Chain+of+Thought+Faithfulness+by+Unlearning+Reasoning+Steps
13. Counterfactual Simulation Training for Chain-of-Thought Faithfulness — Peter Hase, Christopher Potts, 2026
https://scholar.google.com/scholar?q=Counterfactual+Simulation+Training+for+Chain-of-Thought+Faithfulness
14. Why Models Know But Don't Say: Chain-of-Thought Faithfulness Divergence Between Thinking Tokens and Answers in Open-Weight Reasoning Models — Richard J. Young, 2026
https://scholar.google.com/scholar?q=Why+Models+Know+But+Don%27t+Say%3A+Chain-of-Thought+Faithfulness+Divergence+Between+Thinking+Tokens+and+Answers+in+Open-Weight+Reasoning+Models
15. Reasoning with Latent Thoughts: On the Power of Looped Transformers — Nikunj Saunshi et al., 2025
https://scholar.google.com/scholar?q=Reasoning+with+Latent+Thoughts%3A+On+the+Power+of+Looped+Transformers
16. Language Models are Hidden Reasoners: Unlocking Latent Reasoning Capabilities via Self-Rewarding — Haolin Chen et al., 2024
https://scholar.google.com/scholar?q=Language+Models+are+Hidden+Reasoners%3A+Unlocking+Latent+Reasoning+Capabilities+via+Self-Rewarding
17. Efficient Post-Training Refinement of Latent Reasoning in Large Language Models — Xinyuan Wang et al., 2025
https://scholar.google.com/scholar?q=Efficient+Post-Training+Refinement+of+Latent+Reasoning+in+Large+Language+Models
18. AI Post Transformers: Reasoning Theater and Unfaithful Chain-of-Thought — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-reasoning-theater-and-unfaithful-chain-o-a4507e.mp3
19. AI Post Transformers: Generative Recursive Reasoning in Latent Space — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-21-generative-recursive-reasoning-in-latent-a9371d.mp3
20. AI Post Transformers: How Models Detect Hidden Activation Steering — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-08-how-models-detect-hidden-activation-stee-577f73.mp3
21. AI Post Transformers: Latent Space as a New Computational Paradigm — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-05-latent-space-as-a-new-computational-para-810f39.mp3
22. AI Post Transformers: Neural Computers as Learned Latent Runtimes — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-11-neural-computers-as-learned-latent-runti-9fa282.mp3
Interactive Visualization: Opaque Serial Depth and Chain-of-Thought Limits

This episode explores TRELLIS, a bounded-memory transformer architecture that replaces the usual ever-growing key-value cache with a fixed set of learned memory slots that are rewritten during inference. It explains why long-context serving is constrained less by training-time quadratic attention than by the linear growth, latency, and fragility of KV caches, and situates TRELLIS in the progression from Transformer-XL and Compressive Transformers to ABC and GSA. The discussion highlights TRELLIS’s central idea: treating memory as fast weights for a small online regression layer, updating that memory with test-time gradient descent and state decay so the model can reconstruct useful representations while learning what to forget. Listeners would find it interesting because it connects deployment pain points in modern LLMs to a concrete alternative architecture that aims to preserve quality even as context grows while memory stays fixed.

Sources:
1. TRELLIS and Bounded-Memory Transformer KV Compression
https://arxiv.org/pdf/2512.23852
2. Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context — Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V. Le, Ruslan Salakhutdinov, 2019
https://scholar.google.com/scholar?q=Transformer-XL%3A+Attentive+Language+Models+Beyond+a+Fixed-Length+Context
3. Compressive Transformers for Long-Range Sequence Modelling — Jack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, Timothy P. Lillicrap, 2020
https://scholar.google.com/scholar?q=Compressive+Transformers+for+Long-Range+Sequence+Modelling
4. Recurrent Memory Transformer — Aydar Bulatov, Yury Kuratov, Mikhail Burtsev, 2022
https://scholar.google.com/scholar?q=Recurrent+Memory+Transformer
5. Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention — Tsendsuren Munkhdalai, Manaal Faruqui, Siddharth Gopal, 2024
https://scholar.google.com/scholar?q=Leave+No+Context+Behind%3A+Efficient+Infinite+Context+Transformers+with+Infini-attention
6. ABC: Attention with Bounded-Memory Control — Hao Peng et al., 2021
https://scholar.google.com/scholar?q=ABC%3A+Attention+with+Bounded-Memory+Control
7. Gated Slot Attention for Efficient Linear-Time Sequence Modeling — Yu Zhang et al., 2024
https://scholar.google.com/scholar?q=Gated+Slot+Attention+for+Efficient+Linear-Time+Sequence+Modeling
8. Learning to (Learn at Test Time): RNNs with Expressive Hidden States — Yu Sun et al., 2024
https://scholar.google.com/scholar?q=Learning+to+%28Learn+at+Test+Time%29%3A+RNNs+with+Expressive+Hidden+States
9. Lattice: Learning to Efficiently Compress the Memory — Mahdi Karami, Razvan Pascanu, Vahab Mirrokni, 2025
https://scholar.google.com/scholar?q=Lattice%3A+Learning+to+Efficiently+Compress+the+Memory
10. You Only Cache Once: Decoder-Decoder Architectures for Language Models — Yutao Sun et al., 2024
https://scholar.google.com/scholar?q=You+Only+Cache+Once%3A+Decoder-Decoder+Architectures+for+Language+Models
11. KVLink: Accelerating Large Language Models via Efficient KV Cache Reuse — Jingbo Yang et al., 2025
https://arxiv.org/abs/2502.16002
12. LMCache: An Efficient KV Cache Layer for Enterprise-Scale LLM Inference — Yihua Cheng et al., 2025
https://arxiv.org/abs/2510.09665
13. Selective KV-Cache Sharing to Mitigate Timing Side-Channels in LLM Inference — Kexin Chu et al., 2025
https://arxiv.org/abs/2508.08438
14. SparK: Query-Aware Unstructured Sparsity with Recoverable KV Cache Channel Pruning — Huanxuan Liao et al., 2025
https://arxiv.org/abs/2508.15212
15. Test-Time Training Provably Improves Transformers as In-context Learners — Halil Alperen Gozeten et al., 2025
https://arxiv.org/abs/2503.11842
16. Linearizing Vision Transformer with Test-Time Training — Yining Li et al., 2026
https://arxiv.org/abs/2605.02772
17. AI Post Transformers: Titans: Learning to Memorize at Test Time — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-20-titans-learning-to-memorize-at-test-time-054662.mp3
18. AI Post Transformers: Explicit Information Transmission for Context Compression — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-explicit-information-transmission-for-co-24e3c2.mp3
19. AI Post Transformers: Gated Linear Attention for Efficient Long Sequences — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-18-gated-linear-attention-for-efficient-lon-c858ab.mp3
20. AI Post Transformers: In-Place Test-Time Training for Transformers — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-09-in-place-test-time-training-for-transfor-d0b976.mp3
21. AI Post Transformers: Parallelizing DeltaNet Linear Transformers over Sequence Length — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-18-parallelizing-deltanet-linear-transforme-2d0377.mp3
22. AI Post Transformers: Long Context Pre-Training with Lighthouse Attention — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-13-long-context-pre-training-with-lighthous-e85bbe.mp3
23. AI Post Transformers: Compressed Convolutional Attention in Latent Space — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-25-compressed-convolutional-attention-in-la-61e1cf.mp3
Interactive Visualization: TRELLIS and Bounded-Memory Transformer KV Compression

This episode explores why batch-1 LLM decode for robots, edge copilots, and other single-session agents behaves very differently from high-throughput serving, and why next-token latency cannot be explained by memory bandwidth alone. It breaks down the paper’s main test: compare real decode time against an analytic memory floor based on model-weight and KV-cache traffic, then run that across Qwen-2.5-7B, Mistral-7B-v0.3, and Llama-3.1-8B on L4, L40S, A100, and H100 GPUs over contexts from 2048 to 16384. The discussion argues that because these models already use grouped-query attention to cut KV traffic, the remaining latency gap is driven by runtime details such as CUDA Graphs, launch overhead, kernel quality, and whether quantization actually helps in this tiny decode regime. Listeners would find it interesting because it challenges the simple idea that buying a faster-memory GPU automatically lowers token latency, especially for physical AI systems where one delayed token can stall the whole interaction.

Sources:
1. Memory-Bound, Not Bandwidth-Limited Batch-1 LLM Decode
https://arxiv.org/pdf/2605.30571
2. Orca: A Distributed Serving System for Transformer-Based Generative Models — Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim, Byung-Gon Chun, 2022
https://scholar.google.com/scholar?q=Orca%3A+A+Distributed+Serving+System+for+Transformer-Based+Generative+Models
3. Efficient Memory Management for Large Language Model Serving with PagedAttention — Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, Ion Stoica, 2023
https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention
4. Splitwise: Efficient Generative LLM Inference Using Phase Splitting — Pratyush Patel, Esha Choukse, Chaojie Zhang, Aashaka Shah, Inigo Goiri, Saeed Maleki, Ricardo Bianchini, 2024
https://scholar.google.com/scholar?q=Splitwise%3A+Efficient+Generative+LLM+Inference+Using+Phase+Splitting
5. Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving — Ruoyu Qin, Zheming Li, Weiran He, Mingxing Zhang, Yongwei Wu, Weimin Zheng, Xinran Xu, 2024
https://scholar.google.com/scholar?q=Mooncake%3A+A+KVCache-centric+Disaggregated+Architecture+for+LLM+Serving
6. Boosting Performance of Iterative Applications on GPUs: Kernel Batching with CUDA Graphs — Jonah Ekelund, Stefano Markidis, Ivy Peng, 2025
https://scholar.google.com/scholar?q=Boosting+Performance+of+Iterative+Applications+on+GPUs%3A+Kernel+Batching+with+CUDA+Graphs
7. PyGraph: Robust Compiler Support for CUDA Graphs in PyTorch — Abhishek Ghosh, Ajay Nayak, Ashish Panwar, Arkaprava Basu, 2025
https://scholar.google.com/scholar?q=PyGraph%3A+Robust+Compiler+Support+for+CUDA+Graphs+in+PyTorch
8. Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start — Xueshen Liu, Yongji Wu, Yuncheng Yao, Danyang Zhuo, Ion Stoica, Z. Morley Mao, 2026
https://scholar.google.com/scholar?q=Foundry%3A+Template-Based+CUDA+Graph+Context+Materialization+for+Fast+LLM+Serving+Cold+Start
9. Memory-Bound but Not Bandwidth-Limited: The Physical AI Inference Gap in Batch-1 LLM Decode — Josef Chen, 2026
https://scholar.google.com/scholar?q=Memory-Bound+but+Not+Bandwidth-Limited%3A+The+Physical+AI+Inference+Gap+in+Batch-1+LLM+Decode
10. GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints — Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebron, Sumit Sanghai, 2023
https://scholar.google.com/scholar?q=GQA%3A+Training+Generalized+Multi-Query+Transformer+Models+from+Multi-Head+Checkpoints
11. GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers — Elias Frantar, Saleh Ashkboos, Torsten Hoefler, Dan Alistarh, 2022
https://scholar.google.com/scholar?q=GPTQ%3A+Accurate+Post-Training+Quantization+for+Generative+Pre-trained+Transformers
12. AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration — Ji Lin et al., 2023
https://scholar.google.com/scholar?q=AWQ%3A+Activation-aware+Weight+Quantization+for+LLM+Compression+and+Acceleration
13. FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision — Jay Shah, Ganesh Bikshandi, Ying Zhang, Vijay Thakkar, Pradeep Ramani, Tri Dao, 2024
https://scholar.google.com/scholar?q=FlashAttention-3%3A+Fast+and+Accurate+Attention+with+Asynchrony+and+Low-precision
14. FlashDecoding++: Faster Large Language Model Inference on GPUs — Ke Hong et al., 2023
https://scholar.google.com/scholar?q=FlashDecoding%2B%2B%3A+Faster+Large+Language+Model+Inference+on+GPUs
15. Mind the Memory Gap: Unveiling GPU Bottlenecks in Large-Batch LLM Inference — Pol G. Recasens et al., 2025
https://scholar.google.com/scholar?q=Mind+the+Memory+Gap%3A+Unveiling+GPU+Bottlenecks+in+Large-Batch+LLM+Inference
16. Challenges and Research Directions for Large Language Model Inference Hardware — Xiaoyu Ma, David Patterson, 2026
https://scholar.google.com/scholar?q=Challenges+and+Research+Directions+for+Large+Language+Model+Inference+Hardware
17. Medusa: Accelerating Serverless LLM Inference with Materialization — Shaoxun Zeng et al., 2025
https://scholar.google.com/scholar?q=Medusa%3A+Accelerating+Serverless+LLM+Inference+with+Materialization
18. Hybrid JIT-CUDA Graph Optimization for Low-Latency Large Language Model Inference — Divakar Kumar Yadav and Tian Zhao, 2026
https://scholar.google.com/scholar?q=Hybrid+JIT-CUDA+Graph+Optimization+for+Low-Latency+Large+Language+Model+Inference
19. AWQ: Activation-aware Weight Quantization for On-Device LLM Compression and Acceleration — Ji Lin et al., 2024
https://scholar.google.com/scholar?q=AWQ%3A+Activation-aware+Weight+Quantization+for+On-Device+LLM+Compression+and+Acceleration
20. SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression — Tim Dettmers et al., 2024
https://scholar.google.com/scholar?q=SpQR%3A+A+Sparse-Quantized+Representation+for+Near-Lossless+LLM+Weight+Compression
21. Don't Waste Bits! Adaptive KV-Cache Quantization for Lightweight On-Device LLMs — Sayed Pedram Haeri Boroujeni et al., 2026
https://scholar.google.com/scholar?q=Don%27t+Waste+Bits%21+Adaptive+KV-Cache+Quantization+for+Lightweight+On-Device+LLMs
22. KVLink: Accelerating Large Language Models via Efficient KV Cache Reuse — Jingbo Yang et al., 2025
https://scholar.google.com/scholar?q=KVLink%3A+Accelerating+Large+Language+Models+via+Efficient+KV+Cache+Reuse
23. LMCache: An Efficient KV Cache Layer for Enterprise-Scale LLM Inference — Yuhan Liu, Yihua Cheng et al., 2025
https://scholar.google.com/scholar?q=LMCache%3A+An+Efficient+KV+Cache+Layer+for+Enterprise-Scale+LLM+Inference
24. Selective KV-Cache Sharing to Mitigate Timing Side-Channels in LLM Inference — Kexin Chu et al., 2025
https://scholar.google.com/scholar?q=Selective+KV-Cache+Sharing+to+Mitigate+Timing+Side-Channels+in+LLM+Inference
25. AI Post Transformers: Deep Kernel Fusion for Transformer Decoding — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-deep-kernel-fusion-for-transformer-decod-b1a703.mp3
26. AI Post Transformers: Speculative Decoding in Real vLLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-speculative-decoding-in-real-vllm-servin-6f4e2b.mp3
27. AI Post Transformers: LPU Chip for Low-Latency LLM Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-20-lpu-chip-for-low-latency-llm-inference-be13c3.mp3
28. AI Post Transformers: CXL Computational Memory Offloading for Lower Runtime — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-cxl-computational-memory-offloading-for-3b2124.mp3
29. AI Post Transformers: Computation-Bandwidth-Memory Trade-offs for AI Infrastructure — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-09-computation-bandwidth-memory-trade-offs-a83f2b.mp3
30. AI Post Transformers: Serving MoE Models with Disaggregated Expert Parallelism — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-19-serving-moe-models-with-disaggregated-ex-6979d2.mp3

This episode explores the 2008 Dragonfly network topology paper and why its ideas suddenly matter again for large-scale AI systems in 2026. It explains how Dragonfly uses high-radix routers and router groups to keep most traffic to a local hop, a single global hop, and another local hop, reducing the number of expensive long-distance optical links compared with flattened butterfly and folded Clos designs. The discussion highlights the paper’s core argument that topology and routing must be co-designed around pin bandwidth, cable cost, power, and congestion, with the authors claiming roughly 20 percent lower cost than flattened butterfly and 52 percent lower cost than folded Clos beyond 16K nodes under their assumptions. Listeners would find it interesting because it connects an old supercomputing interconnect idea to modern TPU fabrics, mixture-of-experts traffic, all-to-all communication, and the growing reality that network design now directly shapes AI system performance.

Sources:
1. Dragonfly Topology for Scalable AI Networks
https://static.googleusercontent.com/media/research.google.com/en//pubs/archive/34926.pdf
2. Technology-Driven, Highly-Scalable Dragonfly Topology — John Kim, William J. Dally, Steve Scott, Dennis Abts, 2008
https://scholar.google.com/scholar?q=Technology-Driven%2C+Highly-Scalable+Dragonfly+Topology
3. Flattened Butterfly: A Cost-Efficient Topology for High-Radix Networks — John Kim, William J. Dally, Dennis Abts, 2007
https://scholar.google.com/scholar?q=Flattened+Butterfly%3A+A+Cost-Efficient+Topology+for+High-Radix+Networks
4. Topological Characterization of Hamming and Dragonfly Networks and Its Implications on Routing — Cristobal Camarero, Enrique Vallejo, Ramon Beivide, 2014
https://scholar.google.com/scholar?q=Topological+Characterization+of+Hamming+and+Dragonfly+Networks+and+Its+Implications+on+Routing
5. Slim Fly: A Cost Effective Low-Diameter Network Topology — Maciej Besta, Torsten Hoefler, 2014
https://scholar.google.com/scholar?q=Slim+Fly%3A+A+Cost+Effective+Low-Diameter+Network+Topology
6. Microarchitecture of a High-Radix Router — John Kim, William J. Dally, Brian Towles, Amit K. Gupta, 2005
https://scholar.google.com/scholar?q=Microarchitecture+of+a+High-Radix+Router
7. The BlackWidow High-Radix Clos Network — Steve Scott, Dennis Abts, John Kim, William J. Dally, 2006
https://scholar.google.com/scholar?q=The+BlackWidow+High-Radix+Clos+Network
8. Scalable High-Radix Router Microarchitecture Using a Network Switch Organization — Jung Ho Ahn, Young Hoon Son, John Kim, 2013
https://scholar.google.com/scholar?q=Scalable+High-Radix+Router+Microarchitecture+Using+a+Network+Switch+Organization
9. A Scheme for Fast Parallel Communication — L. G. Valiant, 1982
https://scholar.google.com/scholar?q=A+Scheme+for+Fast+Parallel+Communication
10. Indirect Adaptive Routing on Large Scale Interconnection Networks — Nan Jiang, John Kim, William J. Dally, 2009
https://scholar.google.com/scholar?q=Indirect+Adaptive+Routing+on+Large+Scale+Interconnection+Networks
11. Rationale and Challenges for Optical Interconnects to Electronic Chips — David A. B. Miller, 2000
https://scholar.google.com/scholar?q=Rationale+and+Challenges+for+Optical+Interconnects+to+Electronic+Chips
12. Optical Interconnects for High-Performance Computing — Marc A. Taubenblatt, 2012
https://scholar.google.com/scholar?q=Optical+Interconnects+for+High-Performance+Computing
13. Optical Interconnects for Extreme Scale Computing Systems — Sebastien Rumley, Meisam Bahadori, Robert Polster, Simon D. Hammond, David M. Calhoun, Ke Wen, Arun Rodrigues, Keren Bergman, 2017
https://scholar.google.com/scholar?q=Optical+Interconnects+for+Extreme+Scale+Computing+Systems
14. Mission Apollo: Landing Optical Circuit Switching at Datacenter Scale — Ryohei Urata, Hong Liu, Kevin Yasumura, Erji Mao, Jill Berger, Xiang Zhou, Cedric Lam, Roy Bannon, Darren Hutchinson, Daniel Nelson, Leon Poutievski, Arjun Singh, Joon Ong, Amin Vahdat, 2022
https://scholar.google.com/scholar?q=Mission+Apollo%3A+Landing+Optical+Circuit+Switching+at+Datacenter+Scale
15. Adaptive Routing in High-Radix Clos Network — John Kim, William J. Dally, Dennis Abts, 2006
https://doi.org/10.1145/1188455.1188552
16. Efficient All-to-All Collective Communication Schedules for Direct-Connect Topologies — Prithwish Basu, Liangyu Zhao, Jason Fantl, Siddharth Pal, Arvind Krishnamurthy, Joud Khoury, 2024
https://doi.org/10.1145/3625549.3658656
17. Toward lower-diameter large-scale HPC and data center networks with co-packaged optics — Pavlos Maniotis, Laurent Schares, Benjamin G. Lee, Marc A. Taubenblatt, Daniel M. Kuchta, 2021
https://scholar.google.com/scholar?q=Toward+lower-diameter+large-scale+HPC+and+data+center+networks+with+co-packaged+optics
18. Toward higher-radix switches with co-packaged optics for improved network locality in data center and HPC networks [Invited] — Pavlos Maniotis, Laurent Schares, Daniel M. Kuchta, Bengi Karacali, 2022
https://scholar.google.com/scholar?q=Toward+higher-radix+switches+with+co-packaged+optics+for+improved+network+locality+in+data+center+and+HPC+networks+%5BInvited%5D
19. Exploring the benefits of using co-packaged optics in data center and AI supercomputer networks: a simulation-based analysis [Invited] — Pavlos Maniotis, Daniel M. Kuchta, 2024
https://scholar.google.com/scholar?q=Exploring+the+benefits+of+using+co-packaged+optics+in+data+center+and+AI+supercomputer+networks%3A+a+simulation-based+analysis+%5BInvited%5D
20. Enhanced UGAL Routing Schemes for Dragonfly Networks — Ram Sharan Chaulagain, Xin Yuan, 2024
https://scholar.google.com/scholar?q=Enhanced+UGAL+Routing+Schemes+for+Dragonfly+Networks
21. On Selection Functions in Adaptive Routing — Alejandro Cano, Cristobal Camarero, Carmen Martinez, 2025
https://scholar.google.com/scholar?q=On+Selection+Functions+in+Adaptive+Routing
22. Co-packaged optics (CPO): status, challenges, and solutions — Min Tan and coauthors, 2023
https://scholar.google.com/scholar?q=Co-packaged+optics+%28CPO%29%3A+status%2C+challenges%2C+and+solutions
23. AI Post Transformers: Computation-Bandwidth-Memory Trade-offs for AI Infrastructure — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-09-computation-bandwidth-memory-trade-offs-a83f2b.mp3
24. AI Post Transformers: FengHuang for Rack-Scale LLM Inference Memory — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-12-fenghuang-for-rack-scale-llm-inference-m-62708e.mp3
25. AI Post Transformers: Serving MoE Models with Disaggregated Expert Parallelism — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-19-serving-moe-models-with-disaggregated-ex-6979d2.mp3
26. AI Post Transformers: Lossless Sparse Deltas for RL Networks — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-lossless-sparse-deltas-for-rl-networks-84d676.mp3

This episode explores a paper proposing that language models could handle long-context reasoning by periodically pausing, replaying soon-to-be-evicted context offline, and consolidating it into fixed-size fast-weight memory instead of carrying an ever-growing KV cache. It explains the core machinery behind the idea, including state space models and Gated Delta Networks, and clarifies why this is more than prompt summarization or retrieval: the model is rewriting its internal bounded memory during inference. The discussion highlights the paper’s central argument that extra compute may be better spent during these offline “sleep” passes, so later token prediction stays cheap while older information is metabolized into usable latent state. Listeners would find it interesting because it frames long-context scaling as a memory-systems problem, raises concrete questions about whether this consolidation actually improves reasoning, and connects the proposal to broader debates about how future LLMs should trade off memory, compute, and exact recall.

Sources:
1. Do Language Models Need Sleep? Offline Recurrence for Improved Online Inference — Sangyun Lee, Sean McLeish, Tom Goldstein, Giulia Fanti, 2026
http://arxiv.org/abs/2605.26099
2. Replay in Deep Learning: Current Approaches and Missing Biological Elements — Tyler L. Hayes, Giri P. Krishnan, Maxim Bazhenov, Hava T. Siegelmann, Terrence J. Sejnowski, Christopher Kanan, 2021
https://scholar.google.com/scholar?q=Replay+in+Deep+Learning%3A+Current+Approaches+and+Missing+Biological+Elements
3. Can sleep protect memories from catastrophic forgetting? — Oscar C. Gonzalez, Yury Sokolov, Giri P. Krishnan, Jean Erik Delanois, Maxim Bazhenov, 2020
https://scholar.google.com/scholar?q=Can+sleep+protect+memories+from+catastrophic+forgetting%3F
4. Sleep-like unsupervised replay reduces catastrophic forgetting in artificial neural networks — Timothy Tadros, Giri P. Krishnan, Ramyaa Ramyaa, Maxim Bazhenov, 2022
https://scholar.google.com/scholar?q=Sleep-like+unsupervised+replay+reduces+catastrophic+forgetting+in+artificial+neural+networks
5. Do Language Models Need Sleep? Offline Recurrence for Improved Online Inference — Sangyun Lee, Sean McLeish, Tom Goldstein, Giulia Fanti, 2026
https://scholar.google.com/scholar?q=Do+Language+Models+Need+Sleep%3F+Offline+Recurrence+for+Improved+Online+Inference
6. Using Fast Weights to Attend to the Recent Past — Jimmy Ba, Geoffrey Hinton, Volodymyr Mnih, Joel Z. Leibo, Catalin Ionescu, 2016
https://scholar.google.com/scholar?q=Using+Fast+Weights+to+Attend+to+the+Recent+Past
7. Linear Transformers Are Secretly Fast Weight Programmers — Imanol Schlag, Kazuki Irie, Jürgen Schmidhuber, 2021
https://scholar.google.com/scholar?q=Linear+Transformers+Are+Secretly+Fast+Weight+Programmers
8. Fast weight programming and linear transformers: from machine learning to neurobiology — Kazuki Irie, Samuel J. Gershman, 2026
https://scholar.google.com/scholar?q=Fast+weight+programming+and+linear+transformers%3A+from+machine+learning+to+neurobiology
9. TRELLIS: Learning to Compress Key-Value Memory in Attention Models — Mahdi Karami, Ali Behrouz, Praneeth Kacham, Vahab Mirrokni, 2025
https://scholar.google.com/scholar?q=TRELLIS%3A+Learning+to+Compress+Key-Value+Memory+in+Attention+Models
10. Gated Delta Networks: Improving Mamba2 with Delta Rule — Songlin Yang, Jan Kautz, Ali Hatamizadeh, 2024
https://scholar.google.com/scholar?q=Gated+Delta+Networks%3A+Improving+Mamba2+with+Delta+Rule
11. Titans: Learning to Memorize at Test Time — Ali Behrouz, Peilin Zhong, Vahab Mirrokni, 2025
https://scholar.google.com/scholar?q=Titans%3A+Learning+to+Memorize+at+Test+Time
12. Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach — Jonas Geiping, Sean McLeish, Neel Jain, et al., 2025
https://scholar.google.com/scholar?q=Scaling+up+Test-Time+Compute+with+Latent+Reasoning%3A+A+Recurrent+Depth+Approach
13. In-context Autoencoder for Context Compression in a Large Language Model — Tao Ge, Jing Hu, Lei Wang, Xun Wang, Si-Qing Chen, Furu Wei, 2023
https://scholar.google.com/scholar?q=In-context+Autoencoder+for+Context+Compression+in+a+Large+Language+Model
14. Cartridges: Lightweight and general-purpose long context representations via self-study — Sabri Eyuboglu, Ryan Ehrlich, Simran Arora, et al., 2025
https://scholar.google.com/scholar?q=Cartridges%3A+Lightweight+and+general-purpose+long+context+representations+via+self-study
15. Repeat After Me: Transformers are Better than State Space Models at Copying — Samy Jelassi, David Brandfonbrener, Sham M. Kakade, Eran Malach, 2024
https://scholar.google.com/scholar?q=Repeat+After+Me%3A+Transformers+are+Better+than+State+Space+Models+at+Copying
16. End-to-End Test-Time Training for Long Context — Arnuv Tandon et al., 2025
https://scholar.google.com/scholar?q=End-to-End+Test-Time+Training+for+Long+Context
17. Let's (not) just put things in Context: Test-Time Training for Long-Context LLMs — Rachit Bansal et al., 2025
https://scholar.google.com/scholar?q=Let%27s+%28not%29+just+put+things+in+Context%3A+Test-Time+Training+for+Long-Context+LLMs
18. Test-Time Training Done Right — Tianyuan Zhang et al., 2025
https://scholar.google.com/scholar?q=Test-Time+Training+Done+Right
19. Not All Heads Matter: A Head-Level KV Cache Compression Method with Integrated Retrieval and Reasoning — Yu Fu et al., 2024
https://scholar.google.com/scholar?q=Not+All+Heads+Matter%3A+A+Head-Level+KV+Cache+Compression+Method+with+Integrated+Retrieval+and+Reasoning
20. Beyond RAG: Task-Aware KV Cache Compression for Comprehensive Knowledge Reasoning — Giulio Corallo et al., 2025
https://scholar.google.com/scholar?q=Beyond+RAG%3A+Task-Aware+KV+Cache+Compression+for+Comprehensive+Knowledge+Reasoning
21. SideQuest: Model-Driven KV Cache Management for Long-Horizon Agentic Reasoning — Sanjay Kariyappa and G. Edward Suh, 2026
https://scholar.google.com/scholar?q=SideQuest%3A+Model-Driven+KV+Cache+Management+for+Long-Horizon+Agentic+Reasoning
22. Loop, Think, & Generalize: Implicit Reasoning in Recurrent-Depth Transformers — Harsh Kohli et al., 2026
https://scholar.google.com/scholar?q=Loop%2C+Think%2C+%26+Generalize%3A+Implicit+Reasoning+in+Recurrent-Depth+Transformers
23. AI Post Transformers: Titans: Learning to Memorize at Test Time — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-20-titans-learning-to-memorize-at-test-time-054662.mp3
24. AI Post Transformers: In-Place Test-Time Training for Transformers — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-09-in-place-test-time-training-for-transfor-d0b976.mp3
25. AI Post Transformers: Recursive Language Models for Arbitrarily Long Prompts — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-recursive-language-models-for-arbitraril-fbcd1c.mp3
26. AI Post Transformers: Explicit Information Transmission for Context Compression — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-explicit-information-transmission-for-co-24e3c2.mp3
27. AI Post Transformers: KVzip for Query-Agnostic KV Cache Compression — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-29-kvzip-for-query-agnostic-kv-cache-compre-72afe5.mp3
28. AI Post Transformers: Gated Linear Attention for Efficient Long Sequences — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-18-gated-linear-attention-for-efficient-lon-c858ab.mp3
29. AI Post Transformers: MiA-Signature and Global Activation for Long Context — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-13-mia-signature-and-global-activation-for-5ad62f.mp3

This episode explores Google’s Snap system, which moves major host-networking functions out of the kernel and into isolated userspace services while trying to keep the performance benefits usually associated with kernel bypass. It examines why that shift mattered operationally at fleet scale: kernel networking changes could take one to two months to deploy, while Snap enabled roughly weekly releases and had already been adopted across more than half of Google’s machines. The discussion breaks down Snap’s architecture, including centralized host services, microkernel-style isolation, lock-free engine communication, the MicroQuanta scheduler design, latency-sensitive congestion control, and Pony Express as a flagship transport for reliable, asynchronous messaging. Listeners would find it interesting because it frames host networking as a platform-design problem, not just a packet-speed problem, and argues that upgradeability, policy control, and performance can be engineered together rather than traded off.

Sources:
1. Snap's Microkernel Approach to Host Networking
https://storage.googleapis.com/gweb-research2023-media/pubtools/5281.pdf
2. L4 Microkernels: The Lessons from 20 Years of Research and Deployment — Gernot Heiser, Kevin Elphinstone, 2016
https://trustworthy.systems/publications/nicta_full_text/8988.pdf
3. Arrakis: The Operating System is the Control Plane — Simon Peter, Jialin Li, Irene Zhang, Dan R. K. Ports, Doug Woos, Arvind Krishnamurthy, Thomas Anderson, Timothy Roscoe, 2014
https://www.usenix.org/conference/osdi14/technical-sessions/presentation/peter
4. Snap: a Microkernel Approach to Host Networking — Michael Marty, Marc de Kruijf, Jacob Adriaens, Nandita Dukkipati, Amin Vahdat, et al., 2019
https://research.google/pubs/snap-a-microkernel-approach-to-host-networking/
5. netmap: A Novel Framework for Fast Packet I/O — Luigi Rizzo, 2012
https://www.usenix.org/conference/atc12/technical-sessions/presentation/rizzo
6. mTCP: A Highly Scalable User-level TCP Stack for Multicore Systems — EunYoung Jeong, Shinae Woo, Muhammad Jamshed, Haewon Jeong, Sunghwan Ihm, Dongsu Han, KyoungSoo Park, 2014
https://www.usenix.org/system/files/conference/nsdi14/nsdi14-paper-jeong.pdf
7. IX: A Protected Dataplane Operating System for High Throughput and Low Latency — Adam Belay, George Prekas, Ana Klimovic, Samuel Grossman, Christos Kozyrakis, Edouard Bugnion, 2014
https://csl.stanford.edu/~christos/publications/2014.ix.osdi.pdf
8. VL2: A Scalable and Flexible Data Center Network — Albert Greenberg, James R. Hamilton, Navendu Jain, Srikanth Kandula, Changhoon Kim, Parantap Lahiri, Dave Maltz, Parveen Patel, Sudipta Sengupta, 2009
https://www.microsoft.com/en-us/research/publication/vl2-a-scalable-and-flexible-data-center-network/
9. Andromeda: Performance, Isolation, and Velocity at Scale in Cloud Network Virtualization — Michael Dalton, David Schultz, Jacob Adriaens, Ahsan Arefin, Anshuman Gupta, Amin Vahdat, et al., 2018
https://www.usenix.org/conference/nsdi18/presentation/dalton
10. Carousel: Scalable Traffic Shaping at End-Hosts — Ahmed Saeed, Nandita Dukkipati, Valas Valancius, Terry Lam, Carlo Contavalli, Amin Vahdat, 2017
https://research.google/pubs/carousel-scalable-traffic-shaping-at-end-hosts/
11. FaRM: Fast Remote Memory — Aleksandar Dragojevic, Dushyanth Narayanan, Orion Hodson, Miguel Castro, 2014
https://www.usenix.org/conference/nsdi14/technical-sessions/dragojevi%C4%87
12. Using RDMA Efficiently for Key-Value Services — Anuj Kalia, Michael Kaminsky, David G. Andersen, 2014
https://www.pdl.cmu.edu/PDL-FTP/Storage/herd-sigcomm2014.pdf
13. Datacenter RPCs can be General and Fast — Anuj Kalia, Michael Kaminsky, David Andersen, 2019
https://www.usenix.org/conference/nsdi19/presentation/kalia
14. Shenango: Achieving High CPU Efficiency for Latency-sensitive Datacenter Workloads — Amy Ousterhout, Joshua Fried, Jonathan Behrens, Adam Belay, Hari Balakrishnan, 2019
https://www.usenix.org/conference/nsdi19/presentation/ousterhout
15. Caladan: Mitigating Interference at Microsecond Timescales — Joshua Fried, Zhenyuan Ruan, Amy Ousterhout, Adam Belay, 2020
https://www.usenix.org/conference/osdi20/presentation/fried
16. TAS: TCP Acceleration as an OS Service — Antoine Kaufmann, Tim Stamler, Simon Peter, Naveen Kr. Sharma, Arvind Krishnamurthy, and Thomas Anderson, 2019
https://scholar.google.com/scholar?q=TAS%3A+TCP+Acceleration+as+an+OS+Service
17. FaSST: Fast, Scalable and Simple Distributed Transactions with Two-Sided (RDMA) Datagram RPCs — Anuj Kalia, Michael Kaminsky, and David G. Andersen, 2016
https://scholar.google.com/scholar?q=FaSST%3A+Fast%2C+Scalable+and+Simple+Distributed+Transactions+with+Two-Sided+%28RDMA%29+Datagram+RPCs
18. Implementing Network Protocols at User Level — C. A. Thekkath, T. D. Nguyen, E. Moy, and E. D. Lazowska, 1993
https://scholar.google.com/scholar?q=Implementing+Network+Protocols+at+User+Level
19. NetEdit: An Orchestration Platform for eBPF Network Functions at Scale — Theophilus A. Benson et al., 2024
https://doi.org/10.1145/3651890.3672227
20. Demystifying Performance of eBPF Network Applications — Farbod Shahinfar, Sebastiano Miano, Aurojit Panda, Gianni Antichi, 2025
https://cs.nyu.edu/~apanda/assets/papers/conext25.pdf
21. Unleashing Unprivileged eBPF Potential with Dynamic Sandboxing — Soo Yee Lim, Xueyuan Han, Thomas Pasquier, 2023
https://arxiv.org/abs/2308.01983
22. Efficient Scheduler Live Update for Linux Kernel with Modularization — Teng Ma et al., 2023
https://doi.org/10.1145/3582016.3582054
23. Communication Offloading on SmartNIC DPUs: A Quantitative Approach — Jacob Wahlgren et al., 2026
https://arxiv.org/abs/2605.04842

This episode explores SmolLM2, a 1.7 billion parameter language model from Hugging Face that tries to compete with stronger small models not by changing the transformer architecture, but by radically improving the training data mix and sequencing across roughly 11 trillion tokens. It explains the distinction between pretraining and instruction tuning, then argues that for compact models, dataset quality and curriculum can function almost like part of the architecture itself. The discussion connects SmolLM2 to earlier work such as Chinchilla, TinyStories, Textbooks Are All You Need, FineWeb-Edu, and DataComp-LM to show why educational web text, curated math and code data, and staged rebalancing matter so much when model capacity is tight. Listeners would find it interesting because it frames a practical question with real deployment stakes: whether careful data design can make smaller, cheaper, lower-latency models genuinely useful without relying on giant-scale compute.

Sources:
1. SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model — Loubna Ben Allal, Anton Lozhkov, Elie Bakouch, Gabriel Martín Blázquez, Guilherme Penedo, Lewis Tunstall, Andrés Marafioti, Hynek Kydlíček, Agustín Piqueres Lajarín, Vaibhav Srivastav, Joshua Lochner, Caleb Fahlgren, Xuan-Son Nguyen, Clémentine Fourrier, Ben Burtenshaw, Hugo Larcher, Haojun Zhao, Cyril Zakka, Mathieu Morlon, Colin Raffel, Leandro von Werra, Thomas Wolf, 2025
http://arxiv.org/abs/2502.02737
2. Training Compute-Optimal Large Language Models — Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Nalisnick, Daniel Yamins, Timothy Lillicrap, Oriol Vinyals, Jeff Dean, et al., 2022
https://scholar.google.com/scholar?q=Training+Compute-Optimal+Large+Language+Models
3. TinyStories: How Small Can Language Models Be and Still Speak Coherent English? — Ronen Eldan, Yuanzhi Li, 2023
https://scholar.google.com/scholar?q=TinyStories%3A+How+Small+Can+Language+Models+Be+and+Still+Speak+Coherent+English%3F
4. Textbooks Are All You Need — Suriya Gunasekar, Yi Zhang, Jyoti Aneja, Caio C. T. Mendes, Allie Del Giorno, Sivakanth Gopi, Mojan Javaheripi, Piero Kauffmann, Gustavo de Rosa, Olli Saarikivi, Adil Salim, Shital Shah, Harkirat Singh Behl, Xin Wang, Sebastien Bubeck, Ronen Eldan, Adam Tauman Kalai, Yin Tat Lee, Yuanzhi Li, 2023
https://scholar.google.com/scholar?q=Textbooks+Are+All+You+Need
5. MobileLLM: Optimizing Sub-billion Parameter Language Models for On-Device Use Cases — Zechun Liu, Changsheng Zhao, Forrest Iandola, Chen Lai, Yuandong Tian, Igor Fedorov, Yunyang Xiong, Ernie Chang, Yangyang Shi, Raghuraman Krishnamoorthi, Liangzhen Lai, Vikas Chandra, 2024
https://scholar.google.com/scholar?q=MobileLLM%3A+Optimizing+Sub-billion+Parameter+Language+Models+for+On-Device+Use+Cases
6. Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research — Luca Soldaini, Rodney Kinney, Dustin Schwenk, Siddharth Goyal, Alessandro Sordoni, Kyle Lo, Noah A. Smith, and collaborators, 2024
https://scholar.google.com/scholar?q=Dolma%3A+an+Open+Corpus+of+Three+Trillion+Tokens+for+Language+Model+Pretraining+Research
7. The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale — Guilherme Penedo, Hynek Kydlíček, Loubna Ben Allal, Anton Lozhkov, Margaret Mitchell, Colin Raffel, Leandro von Werra, Thomas Wolf, 2024
https://scholar.google.com/scholar?q=The+FineWeb+Datasets%3A+Decanting+the+Web+for+the+Finest+Text+Data+at+Scale
8. Data-Centric AI in the Age of Large Language Models — Xinyi Xu, Zhaoxuan Wu, Rui Qiao, Arun Verma, Yao Shu, Jingtan Wang, Xinyuan Niu, Zhenfeng He, Jiangwei Chen, Zijian Zhou, Gregory Kang Ruey Lau, Hieu Dao, Lucas Agussurja, Rachael Hwee Ling Sim, Xiaoqiang Lin, Wenyang Hu, Zhongxiang Dai, Pang Wei Koh, Bryan Kian Hsiang Low, 2024
https://scholar.google.com/scholar?q=Data-Centric+AI+in+the+Age+of+Large+Language+Models
9. The Stack: 3 TB of permissively licensed source code — Denis Kocetkov, Raymond Li, Loubna Ben Allal, Jia Li, Chenghao Mou, Carlos Muñoz Ferrandis, Yacine Jernite, Margaret Mitchell, Sean Hughes, Thomas Wolf, Dzmitry Bahdanau, Leandro von Werra, Harm de Vries, 2022
https://scholar.google.com/scholar?q=The+Stack%3A+3+TB+of+permissively+licensed+source+code
10. Enhancing Chat Language Models by Scaling High-quality Instructional Conversations — Ning Ding, Yulin Chen, Bokai Xu, et al., 2023
https://scholar.google.com/scholar?q=Enhancing+Chat+Language+Models+by+Scaling+High-quality+Instructional+Conversations
11. OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction Data — Shubham Toshniwal, Wei Du, Ivan Moshkov, Branislav Kisacanin, Alexan Ayrapetyan, Igor Gitman, 2024
https://scholar.google.com/scholar?q=OpenMathInstruct-2%3A+Accelerating+AI+for+Math+with+Massive+Open-Source+Instruction+Data
12. SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model — Loubna Ben Allal, Anton Lozhkov, Elie Bakouch, Gabriel Martín Blázquez, Guilherme Penedo, Lewis Tunstall, Andrés Marafioti, Hynek Kydlíček, Agustín Piqueres Lajarín, Vaibhav Srivastav, Joshua Lochner, Caleb Fahlgren, Xuan-Son Nguyen, Clémentine Fourrier, Ben Burtenshaw, Hugo Larcher, Haojun Zhao, Cyril Zakka, Mathieu Morlon, Colin Raffel, Leandro von Werra, Thomas Wolf, 2025
https://scholar.google.com/scholar?q=SmolLM2%3A+When+Smol+Goes+Big+--+Data-Centric+Training+of+a+Small+Language+Model
13. DataComp-LM: In search of the next generation of training sets for language models — Jeffrey Li, Alex Fang, Georgios Smyrnis, Maor Ivgi, Matt Jordan, Samir Gadre, Hritik Bansal, Etash Guha, Sedrick Keh, Kushal Arora, Saurabh Garg, Rui Xin, Niklas Muennighoff, and many others, 2024
https://scholar.google.com/scholar?q=DataComp-LM%3A+In+search+of+the+next+generation+of+training+sets+for+language+models
14. OpenWebMath: An Open Dataset of High-Quality Mathematical Web Text — Keiran Paster, Marco Dos Santos, Zhangir Azerbayev, Jimmy Ba, 2023
https://scholar.google.com/scholar?q=OpenWebMath%3A+An+Open+Dataset+of+High-Quality+Mathematical+Web+Text
15. InfiMM-WebMath-40B: Advancing Multimodal Pre-Training for Enhanced Mathematical Reasoning — Xiaotian Han, Yiren Jian, Xuefeng Hu, Haogeng Liu, Yiqi Wang, Qihang Fan, Yuang Ai, Huaibo Huang, Ran He, Zhenheng Yang, Quanzeng You, 2024
https://scholar.google.com/scholar?q=InfiMM-WebMath-40B%3A+Advancing+Multimodal+Pre-Training+for+Enhanced+Mathematical+Reasoning
16. DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models — Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, Y.K. Li, Y. Wu, Daya Guo, 2024
https://scholar.google.com/scholar?q=DeepSeekMath%3A+Pushing+the+Limits+of+Mathematical+Reasoning+in+Open+Language+Models
17. 2 OLMo 2 Furious — Kyle Lo and the OLMo team, 2025
https://scholar.google.com/scholar?q=2+OLMo+2+Furious
18. Revisiting Scaling Laws for Language Models: The Role of Data Quality and Training Strategies — Zhengyu Chen, Siqi Wang, Teng Xiao, Yudong Wang, Shiqi Chen, Xunliang Cai, Junxian He, Jingang Wang, 2025
https://scholar.google.com/scholar?q=Revisiting+Scaling+Laws+for+Language+Models%3A+The+Role+of+Data+Quality+and+Training+Strategies
19. GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining — Simin Fan, Maria Ios Glarou, Martin Jaggi, 2025
https://scholar.google.com/scholar?q=GRAPE%3A+Optimize+Data+Mixture+for+Group+Robust+Multi-target+Adaptive+Pretraining
20. Optimizing Pre-Training Data Mixtures with Mixtures of Data Expert Models — Lior Belenki, Alekh Agarwal, Tianze Shi, Kristina Toutanova, 2025
https://scholar.google.com/scholar?q=Optimizing+Pre-Training+Data+Mixtures+with+Mixtures+of+Data+Expert+Models
21. Scaling Laws with Vocabulary: Larger Models Deserve Larger Vocabularies — Chaofan Tao, Qian Liu, Longxu Dou, Niklas Muennighoff, Zhongwei Wan, Ping Luo, Min Lin, Ngai Wong, 2024
https://scholar.google.com/scholar?q=Scaling+Laws+with+Vocabulary%3A+Larger+Models+Deserve+Larger+Vocabularies
22. Distilling Reasoning Capabilities into Smaller Language Models — Kumar Shridhar, Alessandro Stolfo, Mrinmaya Sachan, 2023
https://scholar.google.com/scholar?q=Distilling+Reasoning+Capabilities+into+Smaller+Language+Models
23. Teaching Small Language Models Reasoning through Counterfactual Distillation — Tao Feng, Yicheng Li, Chenglin Li, Hao Chen, Fei Yu, Yin Zhang, 2024
https://scholar.google.com/scholar?q=Teaching+Small+Language+Models+Reasoning+through+Counterfactual+Distillation
24. AI Post Transformers: Self-Improving Pretraining With Post-Trained Models — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-02-self-improving-pretraining-with-post-tra-e37460.mp3
25. AI Post Transformers: Scaling Laws for Multilingual Code Pretraining — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-scaling-laws-for-multilingual-code-pretr-7d220e.mp3
26. AI Post Transformers: Can Models Learn from Long Context? — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-can-models-learn-from-long-context-77533e.mp3
27. AI Post Transformers: ASI-Evolve for Data, Architectures, and RL — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-05-asi-evolve-for-data-architectures-and-rl-197b2b.mp3
28. AI Post Transformers: Muon Is Scalable for LLM Training — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-25-muon-is-scalable-for-llm-training-587ed8.mp3

This episode explores a post-training method for making mixture-of-experts language models cheaper at inference time without retraining them from scratch. It explains how the paper converts a fully trained static MoE into a dynamic one by adding parameter-free zero experts, allowing some tokens to skip normal experts, and then uses self-distillation to preserve the original model’s behavior under this lower-compute routing scheme. The discussion highlights why this deployment-focused approach matters for real production systems, especially when pretraining, fine-tuning, and alignment are already complete and inference cost is the main bottleneck. Listeners would find it interesting for its clear breakdown of dynamic versus static MoE compute, its practical framing around latency and serving costs, and its focus on whether large post-trained models can cut expert FLOPs substantially without losing capability.

Sources:
1. Post-Trained MoE Can Skip Half Experts via Self-Distillation — Xingtai Lv, Li Sheng, Kaiyan Zhang, Yichen You, Siyan Gao, Xueheng Luo, Yuxin Zuo, Yuchen Fan, Junlin Yang, Ganqu Cui, Bingning Wang, Fan Yang, Youbang Sun, Ning Ding, Bowen Zhou, 2026
http://arxiv.org/abs/2605.18643
2. MoE++: Accelerating Mixture-of-Experts Methods with Zero-Computation Experts — Peng Jin, Bo Zhu, Li Yuan, Shuicheng Yan, 2024
https://scholar.google.com/scholar?q=MoE%2B%2B%3A+Accelerating+Mixture-of-Experts+Methods+with+Zero-Computation+Experts
3. Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models — Xudong Lu, Qi Liu, Yuhui Xu, Aojun Zhou, Siyuan Huang, Bo Zhang, Junchi Yan, Hongsheng Li, 2024
https://scholar.google.com/scholar?q=Not+All+Experts+are+Equal%3A+Efficient+Expert+Pruning+and+Skipping+for+Mixture-of-Experts+Large+Language+Models
4. Task-Specific Expert Pruning for Sparse Mixture-of-Experts — Tianyu Chen, Shaohan Huang, Yuan Xie, Binxing Jiao, Daxin Jiang, Haoyi Zhou, Jianxin Li, Furu Wei, 2022
https://scholar.google.com/scholar?q=Task-Specific+Expert+Pruning+for+Sparse+Mixture-of-Experts
5. Auxiliary-Loss-Free Load Balancing Strategy for Mixture-of-Experts — DeepSeek-AI et al., 2024
https://scholar.google.com/scholar?q=Auxiliary-Loss-Free+Load+Balancing+Strategy+for+Mixture-of-Experts
6. ST-MoE: Designing Stable and Transferable Sparse Expert Models — Barret Zoph, Noam Shazeer, William Fedus, et al., 2022
https://scholar.google.com/scholar?q=ST-MoE%3A+Designing+Stable+and+Transferable+Sparse+Expert+Models
7. AdaMoE: Token-Adaptive Routing with Null Experts for Mixture-of-Experts Language Models — Zihao Zeng, Yibo Miao, Hongcheng Gao, Hao Zhang, Zhijie Deng, 2024
https://scholar.google.com/scholar?q=AdaMoE%3A+Token-Adaptive+Routing+with+Null+Experts+for+Mixture-of-Experts+Language+Models
8. Harder Task Needs More Experts: Dynamic Routing in MoE Models — Quzhe Huang, Zhenwei An, Nan Zhuang, Mingxu Tao, Chen Zhang, Yang Jin, Kun Xu, Liwei Chen, Songfang Huang, Yansong Feng, 2024
https://scholar.google.com/scholar?q=Harder+Task+Needs+More+Experts%3A+Dynamic+Routing+in+MoE+Models
9. MoE Pathfinder: Trajectory-driven Expert Pruning — Xican Yang, Yuanhe Tian, Yan Song, 2025
https://scholar.google.com/scholar?q=MoE+Pathfinder%3A+Trajectory-driven+Expert+Pruning
10. Discovering Important Experts for Mixture-of-Experts Models Pruning Through a Theoretical Perspective — approximate only; title verified, authors not confidently recovered, 2025/2026
https://scholar.google.com/scholar?q=Discovering+Important+Experts+for+Mixture-of-Experts+Models+Pruning+Through+a+Theoretical+Perspective
11. MoEEdit: Efficient and Routing-Stable Knowledge Editing for Mixture-of-Experts LLMs — Yupu Gu, Rongzhe Wei, Andy Zhu, Pan Li, 2026
https://scholar.google.com/scholar?q=MoEEdit%3A+Efficient+and+Routing-Stable+Knowledge+Editing+for+Mixture-of-Experts+LLMs
12. ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning — Chao Jin, Xinming Wei, Yinmin Zhong, Chengxu Yang, Bingyang Wu, Ruidong Zhu, Zili Zhang, Yuliang Liu, Xin Jin, 2026
https://scholar.google.com/scholar?q=ReLibra%3A+Routing-Replay-Guided+Load+Balancing+for+MoE+Training+in+Reinforcement+Learning
13. Sparse MoE Students for Efficient Knowledge Distillation — approximate only; exact author list not confidently recovered, 2025
https://scholar.google.com/scholar?q=Sparse+MoE+Students+for+Efficient+Knowledge+Distillation
14. AI Post Transformers: Batch-Aware Expert Routing for Faster MoE Decoding — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-batch-aware-expert-routing-for-faster-mo-683ab6.mp3
15. AI Post Transformers: Serving MoE Models with Disaggregated Expert Parallelism — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-19-serving-moe-models-with-disaggregated-ex-6979d2.mp3
16. AI Post Transformers: Ministral 3: Cascade Distillation for Long-Context Multimodal Models — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-cascade-distillation-for-long-context-mu-0ebd1a.mp3
17. AI Post Transformers: Nemotron 3 Super Hybrid Mamba-Transformer MoE — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-19-nemotron-3-super-hybrid-mamba-transforme-31ac75.mp3
18. AI Post Transformers: LPU Chip for Low-Latency LLM Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-20-lpu-chip-for-low-latency-llm-inference-be13c3.mp3

Browse by month

Sep 2026 Aug 2026 Jul 2026 Jun 2026 May 2026 Apr 2026 Mar 2026 Feb 2026 Jan 2026 Dec 2025 Nov 2025 Oct 2025 Sep 2025 Aug 2025