This episode explores MemPO, a self-memory policy optimization framework for long-horizon AI agents developed by researchers at Tsinghua University and Alibaba's Tongyi Lab. The discussion contrasts MemPO's approach against the dominant ReAct pattern, which accumulates full interaction history and suffers from both ballooning token costs and the "lost in the middle" degradation documented in prior research, as well as against passive retrieval-based memory systems like MemGPT and Mem0 that rely on embedding similarity rather than task outcomes. The hosts unpack how MemPO trains an agent to write compressed memory notes as a learned, RL-optimized action — discarding raw tool outputs and reasoning traces at each step in favor of a single distilled note — using Group Relative Policy Optimization to solve the credit-assignment problem of rewarding intermediate memory decisions from a single end-of-trajectory success signal. Listeners interested in agent architecture, RL training objectives, or the tradeoffs between context-window scaling and structured memory will find the episode's walk-through of the mem/think/tool_call decomposition particularly useful. The conversation also traces the intellectual lineage of the ideas, from Minsky's original framing of credit assignment to retrieval-augmented generation's origins at Facebook AI Research.
Sources:
1. MemPO: Self-Memory Policy Optimization for Long-Horizon Agents — Ruoran Li, Xinghua Zhang, Haiyang Yu, Shitong Duan, Xiang Li, Wenxin Xiang, Chonghua Liao, Xudong Guo, Yongbin Li, Jinli Suo, 2026
http://arxiv.org/abs/2603.006802. Policy Gradient Methods for Reinforcement Learning with Function Approximation — Richard S. Sutton, David McAllester, Satinder Singh, Yishay Mansour, 1999/2000
https://scholar.google.com/scholar?q=Policy+Gradient+Methods+for+Reinforcement+Learning+with+Function+Approximation3. High-Dimensional Continuous Control Using Generalized Advantage Estimation — John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, Pieter Abbeel, 2016
https://scholar.google.com/scholar?q=High-Dimensional+Continuous+Control+Using+Generalized+Advantage+Estimation4. RUDDER: Return Decomposition for Delayed Rewards — Jose A. Arjona-Medina, Michael Gillhofer, Michael Widrich, Thomas Adler, Johannes Brandstetter, Sepp Hochreiter, 2019
https://scholar.google.com/scholar?q=RUDDER%3A+Return+Decomposition+for+Delayed+Rewards5. Let's Verify Step by Step — Hunter Lightman, Vineet Kosaraju, Yura Burda, Harrison Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, Karl Cobbe, 2023
https://scholar.google.com/scholar?q=Let%27s+Verify+Step+by+Step6. Information Gain-based Policy Optimization: A Simple and Effective Approach for Multi-Turn LLM Agents — Guoqing Wang, Sunhao Dai, Guangze Ye, Zeyu Gan, Wei Yao, Yong Deng, Xiaofeng Wu, Zhenzhe Ying, 2025
https://scholar.google.com/scholar?q=Information+Gain-based+Policy+Optimization%3A+A+Simple+and+Effective+Approach+for+Multi-Turn+LLM+Agents7. MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon Agents — Zijian Zhou, Ao Qu, Zhaoxuan Wu, Sunghwan Kim, Alok Prakash, Daniela Rus, Jinhua Zhao, Bryan Kian Hsiang Low, Paul Pu Liang, 2025
https://scholar.google.com/scholar?q=MEM1%3A+Learning+to+Synergize+Memory+and+Reasoning+for+Efficient+Long-Horizon+Agents8. Reflexion: Language Agents with Verbal Reinforcement Learning — Noah Shinn, Federico Cassano, Edward Berman, Ashwin Gopinath, Karthik Narasimhan, Shunyu Yao, 2023
https://scholar.google.com/scholar?q=Reflexion%3A+Language+Agents+with+Verbal+Reinforcement+Learning9. Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning — Bowen Jin, Hansi Zeng, Zhenrui Yue, Jinsung Yoon, Sercan Arik, Dong Wang, Hamed Zamani, Jiawei Han, 2025
https://scholar.google.com/scholar?q=Search-R1%3A+Training+LLMs+to+Reason+and+Leverage+Search+Engines+with+Reinforcement+Learning10. Attnpo: Attention-guided Process Supervision for Efficient Reasoning — Shuaiyi Nie, Siyu Ding, Wenyuan Zhang, Linhao Yu, Tianmeng Yang, Yao Chen, Tingwen Liu, Weichong Yin, Yu Sun, Hua Wu, 2026
https://scholar.google.com/scholar?q=Attnpo%3A+Attention-guided+Process+Supervision+for+Efficient+Reasoning11. MemGPT: Towards LLMs as Operating Systems — Charles Packer, Sarah Wooders, Kevin Lin, Vivian Fang, Shishir G. Patil, Ion Stoica, Joseph E. Gonzalez, 2024
https://scholar.google.com/scholar?q=MemGPT%3A+Towards+LLMs+as+Operating+SystemsInteractive Visualization: MemPO: Teaching Agents to Write Their Own Memory