This episode explores the AI+HW 2035 roadmap, arguing that the next decade of AI progress will depend less on raw compute growth and more on coordinated design across models, compilers, runtimes, memory systems, and chips. It breaks down the memory wall in concrete terms, showing how moving weights, activations, and KV caches can cost more time and energy than the math itself, especially for inference, autoregressive serving, and state-heavy workloads like video world models. The discussion examines quantization, mixed precision, sparsity, pruning, distillation, tiered memory, IO-aware attention, and hardware-aware scheduling, with the key claim that these methods only matter when the full stack preserves locality and avoids wasteful data movement. Listeners would find it interesting because it treats AI efficiency as a practical systems problem and policy agenda, not just a matter of inventing better model architectures.
Sources:
1. AI+HW 2035: Shaping the Next Decade — Deming Chen, Jason Cong, Azalia Mirhoseini, Christos Kozyrakis, Subhasish Mitra, Jinjun Xiong, Cliff Young, Anima Anandkumar, Michael Littman, Aron Kirschen, Sophia Shao, Serge Leef, Naresh Shanbhag, Dejan Milojicic, Michael Schulte, Gert Cauwenberghs, Jerry M. Chow, Tri Dao, Kailash Gopalakrishnan, Richard Ho, Hoshik Kim, Kunle Olukotun, David Z. Pan, Mark Ren, Dan Roth, Aarti Singh, Yizhou Sun, Yusu Wang, Yann LeCun, Ruchir Puri, 2026
http://arxiv.org/abs/2603.052252. Mixed Precision Training — Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, and others, 2017
https://scholar.google.com/scholar?q=Mixed+Precision+Training3. Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference — Benoit Jacob, Skirmantas Kligys, Bo Chen, Menglong Zhu, Matthew Tang, Andrew Howard, and others, 2018
https://scholar.google.com/scholar?q=Quantization+and+Training+of+Neural+Networks+for+Efficient+Integer-Arithmetic-Only+Inference4. FP8 Formats for Deep Learning — Paulius Micikevicius, Dusan Stosic, Neil Burgess, Marius Cornea, Pradeep Dubey, and others, 2022
https://scholar.google.com/scholar?q=FP8+Formats+for+Deep+Learning5. AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration — Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Wei-Ming Chen, Wei-Chen Wang, Guangxuan Xiao, Chuang Gan, Song Han, and others, 2023
https://scholar.google.com/scholar?q=AWQ%3A+Activation-aware+Weight+Quantization+for+LLM+Compression+and+Acceleration6. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness — Tri Dao et al., 2022
https://scholar.google.com/scholar?q=FlashAttention%3A+Fast+and+Memory-Efficient+Exact+Attention+with+IO-Awareness7. A Full-Stack Search Technique for Domain Optimized Deep Learning Accelerators — Dan Zhang et al., 2022
https://scholar.google.com/scholar?q=A+Full-Stack+Search+Technique+for+Domain+Optimized+Deep+Learning+Accelerators8. A Compute-in-Memory Chip Based on Resistive Random-Access Memory — Weier Wan et al., 2022
https://scholar.google.com/scholar?q=A+Compute-in-Memory+Chip+Based+on+Resistive+Random-Access+Memory9. Intelligence per Watt: Measuring Intelligence Efficiency of Local AI — Jon Saad-Falcon et al., 2025
https://scholar.google.com/scholar?q=Intelligence+per+Watt%3A+Measuring+Intelligence+Efficiency+of+Local+AI10. PinDrop: Breaking the Silence on SDCs in a Large-Scale Fleet — Peter W. Deutsch et al., 2026
https://scholar.google.com/scholar?q=PinDrop%3A+Breaking+the+Silence+on+SDCs+in+a+Large-Scale+Fleet11. LMCache: An Efficient KV Cache Layer for Enterprise-Scale LLM Inference — Yuhan Liu et al., 2025
https://scholar.google.com/scholar?q=LMCache%3A+An+Efficient+KV+Cache+Layer+for+Enterprise-Scale+LLM+Inference12. TraCT: Disaggregated LLM Serving with CXL Shared Memory KV Cache at Rack-Scale — Dongha Yoon et al., 2025
https://scholar.google.com/scholar?q=TraCT%3A+Disaggregated+LLM+Serving+with+CXL+Shared+Memory+KV+Cache+at+Rack-Scale13. Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads — Cunchen Hu et al., 2024
https://scholar.google.com/scholar?q=Inference+without+Interference%3A+Disaggregate+LLM+Inference+for+Mixed+Downstream+Workloads14. CXL-SpecKV: A Disaggregated FPGA Speculative KV-Cache for Datacenter LLM Serving — Dong Liu and Yanxuan Yu, 2025
https://scholar.google.com/scholar?q=CXL-SpecKV%3A+A+Disaggregated+FPGA+Speculative+KV-Cache+for+Datacenter+LLM+Serving15. KVLink: Accelerating Large Language Models via Efficient KV Cache Reuse — Jingbo Yang et al., 2025
https://scholar.google.com/scholar?q=KVLink%3A+Accelerating+Large+Language+Models+via+Efficient+KV+Cache+Reuse16. Selective KV-Cache Sharing to Mitigate Timing Side-Channels in LLM Inference — Kexin Chu et al., 2025
https://scholar.google.com/scholar?q=Selective+KV-Cache+Sharing+to+Mitigate+Timing+Side-Channels+in+LLM+Inference17. Test-Time Training on Nearest Neighbors for Large Language Models — Moritz Hardt and Yu Sun, 2023
https://scholar.google.com/scholar?q=Test-Time+Training+on+Nearest+Neighbors+for+Large+Language+Models18. Test-Time Learning for Large Language Models — Jinwu Hu et al., 2025
https://scholar.google.com/scholar?q=Test-Time+Learning+for+Large+Language+Models19. AI Post Transformers: FlatAttention for Tile-Based Accelerator Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-flatattention-for-tile-based-accelerator-56e6ca.mp320. AI Post Transformers: LPU Chip for Low-Latency LLM Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-20-lpu-chip-for-low-latency-llm-inference-be13c3.mp321. AI Post Transformers: Mamba-3 for Efficient Sequence Modeling — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-16-mamba-3-for-efficient-sequence-modeling-97a22a.mp322. AI Post Transformers: Mistral 7B: Superior Performance in a Smaller Package — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/mistral-7b-superior-performance-in-a-smaller-package/23. AI Post Transformers: PALOMA: Benchmarking Language Model Fit Across Domains — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-23-paloma-benchmarking-language-model-fit-a-360060.mp324. AI Post Transformers: Automating CNN Mapping on Embedded FPGAs — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-06-automating-cnn-mapping-on-embedded-fpgas-4c1dc3.mp325. AI Post Transformers: FPGA Neural Network Accelerators for Space — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-26-fpga-neural-network-accelerators-for-spa-3087ae.mp326. AI Post Transformers: ASI-Evolve for Data, Architectures, and RL — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-05-asi-evolve-for-data-architectures-and-rl-197b2b.mp3