This episode explores Meta-Harness, a paper arguing that a large share of LLM system performance comes from the surrounding harness code that manages memory, retrieval, tool use, context formatting, and control flow rather than from model weights alone. It explains how the method uses an outer-loop coding agent to rewrite harness code, inspect raw traces and logs stored on disk, and search for better system designs across tasks like text classification, retrieval-based math reasoning, and agentic coding. The discussion highlights why this matters: in multi-step systems, the same fixed model can perform very differently depending on what information it sees, when it sees it, and how the wrapper code structures the interaction. Listeners would find it interesting because it reframes progress in AI systems as a systems-engineering problem, raising the possibility that better scaffolding around existing models may unlock major gains without retraining the models themselves.
Sources:
1. Meta-Harness: End-to-End Optimization of Model Harnesses — Yoonho Lee, Roshen Nair, Qizheng Zhang, Kangwook Lee, Omar Khattab, Chelsea Finn, 2026
http://arxiv.org/abs/2603.280522.
https://yoonholee.com/meta-harness/ https://yoonholee.com/meta-harness/3. DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines — Omar Khattab, Arnav Singhvi, Paridhi Maheshwari, Zhiyuan Zhang, Keshav Santhanam, Matei Zaharia, Christopher Potts, and others, 2023
https://scholar.google.com/scholar?q=DSPy%3A+Compiling+Declarative+Language+Model+Calls+into+Self-Improving+Pipelines4. Reflexion: Language Agents with Verbal Reinforcement Learning — Noah Shinn, Federico Cassano, Edward Berman, Ashwin Gopinath, Karthik Narasimhan, Shunyu Yao, 2023
https://scholar.google.com/scholar?q=Reflexion%3A+Language+Agents+with+Verbal+Reinforcement+Learning5. MemGPT: Towards LLMs as Operating Systems — Charles Packer, Sarah Wooders, Kevin Lin, Vivian Fang, Shishir G. Patil, Ion Stoica, Joseph E. Gonzalez, 2023
https://scholar.google.com/scholar?q=MemGPT%3A+Towards+LLMs+as+Operating+Systems6. Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models — Qizheng Zhang, Changran Hu, Shubhangi Upasani, Boyuan Ma, Fenglu Hong, James Zou, Kunle Olukotun, and others, 2025
https://scholar.google.com/scholar?q=Agentic+Context+Engineering%3A+Evolving+Contexts+for+Self-Improving+Language+Models7. GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning — Lakshya A. Agrawal, Shangyin Tan, Dilara Soylu, Noah Ziems, Rishi Khare, Krista Opsahl-Ong, Arnav Singhvi, Herumb Shandilya, Michael J. Ryan, Meng Jiang, Christopher Potts, Koushik Sen, Alexandros G. Dimakis, Ion Stoica, Dan Klein, Matei Zaharia, Omar Khattab, 2025
https://scholar.google.com/scholar?q=GEPA%3A+Reflective+Prompt+Evolution+Can+Outperform+Reinforcement+Learning8. TextGrad: Automatic "Differentiation" via Text — Mert Yuksekgonul, Federico Bianchi, Joseph Boen, Sheng Liu, Zhi Huang, Carlos Guestrin, James Zou, 2024
https://scholar.google.com/scholar?q=TextGrad%3A+Automatic+%22Differentiation%22+via+Text9. Large Language Models as Optimizers — Chengrun Yang, Xuezhi Wang, Yifeng Lu, Hanxiao Liu, Quoc V. Le, Denny Zhou, Xinyun Chen, 2023
https://scholar.google.com/scholar?q=Large+Language+Models+as+Optimizers10. AlphaEvolve: A Coding Agent for Scientific and Algorithmic Discovery — Alexander Novikov, Ngan Vu, Marvin Eisenberger, Emilien Dupont, Po-Sen Huang, Adam Zsolt Wagner, Sergey Shirobokov, Borislav Kozlovskii, Francisco J. R. Ruiz, Abbas Mehrabian, M. Pawan Kumar, Abigail See, Swarat Chaudhuri, George Holland, Alex Davies, Sebastian Nowozin, Pushmeet Kohli, Matej Balog, 2025
https://scholar.google.com/scholar?q=AlphaEvolve%3A+A+Coding+Agent+for+Scientific+and+Algorithmic+Discovery11. Learning to Discover at Test Time — Mert Yuksekgonul, Daniel Koceja, Xinhao Li, Federico Bianchi, Jed McCaleb, Xiaolong Wang, Jan Kautz, Yejin Choi, James Zou, Carlos Guestrin, Yu Sun, 2026
https://scholar.google.com/scholar?q=Learning+to+Discover+at+Test+Time12. Grounded Test-Time Adaptation for LLM Agents — Arthur Chen et al., 2025
https://scholar.google.com/scholar?q=Grounded+Test-Time+Adaptation+for+LLM+Agents13. Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory — Tianxin Wei et al., 2025
https://scholar.google.com/scholar?q=Evo-Memory%3A+Benchmarking+LLM+Agent+Test-time+Learning+with+Self-Evolving+Memory14. M^2: Dual-Memory Augmentation for Long-Horizon Web Agents via Trajectory Summarization and Insight Retrieval — Dawei Yan et al., 2026
https://scholar.google.com/scholar?q=M%5E2%3A+Dual-Memory+Augmentation+for+Long-Horizon+Web+Agents+via+Trajectory+Summarization+and+Insight+Retrieval15. Reinforcement Fine-Tuning for History-Aware Dense Retriever in RAG — Yicheng Zhang et al., 2026
https://scholar.google.com/scholar?q=Reinforcement+Fine-Tuning+for+History-Aware+Dense+Retriever+in+RAG16. MAIN-RAG: Multi-Agent Filtering Retrieval-Augmented Generation — Chia-Yuan Chang et al., 2024
https://scholar.google.com/scholar?q=MAIN-RAG%3A+Multi-Agent+Filtering+Retrieval-Augmented+Generation17. Fine-tuning with RAG for Improving LLM Learning of New Skills — Humaid Ibrahim, Nikolai Rozanov, Marek Rei, 2025
https://scholar.google.com/scholar?q=Fine-tuning+with+RAG+for+Improving+LLM+Learning+of+New+Skills18. AI Post Transformers: Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/agentic-context-engineering-evolving-contexts-for-self-improving-language-models/19. AI Post Transformers: Mem0: Scalable Long-Term Memory for AI Agents — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/mem0-scalable-long-term-memory-for-ai-agents/20. AI Post Transformers: Agentic AI and the Next Intelligence Explosion — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-28-agentic-ai-and-the-next-intelligence-exp-d06561.mp3