AI Post Transformers — Episode Companion

MemPO: Teaching Agents to Write Their Own Memory

A self-memory policy optimization framework that trains a long-horizon agent to write its own compressed memory note, as a learned RL action, instead of bolting on retrieval after the fact.

arXiv:2603.00680 Ruoran Li, Xinghua Zhang, Haiyang Yu, et al. — 10 authors Tsinghua University · Alibaba Tongyi Lab 2026

ReAct vs. MemPO: what actually travels between steps

ReAct (Yao et al., 2022) appends every tool result and reasoning trace to the prompt, so context grows every round. MemPO discards all of that and carries forward only a single compressed memory note.

Context size over a 10-step trajectory

Hover any point for the exact token count at that step. ReAct's context climbs steadily toward the window limit; MemPO's stays flat because only the note carries forward.

The mem / think / tool_call decomposition

Every step writes a memory note, reasons about what to do next, then acts. Only the note from step t-1 becomes step t's input — the reasoning trace and raw tool output are discarded.

Compression ratio per step

A single step can involve thousands of tokens of tool output and reasoning. Almost all of it is thrown away the moment the note is written.

Trajectory-level advantage (AT) from GRPO

Sixteen rollouts answer the same question. Reward is 1 for a clean, correct trajectory and 0 otherwise. Advantage is each rollout's reward normalized against the group's mean and standard deviation. Hover a cell for the exact values.

Memory-level advantage (AM): grading a note by its own confidence

Show only the note to the model and have it try to produce the ground-truth answer. A confident, correct guess means the note captured what mattered; subtracting the same probability conditioned on full history isolates the note's marginal contribution.

mem tokens get A_T + A_M · every other token gets A_T only

Averaged over 2–10 chained objectives on local Wikipedia search: MemPO 37.63 F1, GRPO with no memory reward 30.53, MEM1 26.34.

Context-truncation ablation

Full history wins on short-horizon tasks; as objective count climbs, one-step truncation overtakes it — the same lost-in-the-middle effect, this time self-inflicted by hoarding old memory.

Isolating the memory reward (Figure 3, left panel)

GRPO with no dedicated memory reward already beats MEM1 by several points using the same warm start and truncated context. The reward mechanism — the one thing MemPO adds on top of that — is real but smaller than the headline gap over MEM1 suggests.

The abstract's "7.1 points over previous SOTA" bundles the GPT-4.1 warm-start distillation and the truncated-context recipe together with the memory reward. The isolated ablation — reward on vs. reward off, everything else held fixed — is the real marginal effect of what the paper claims to have invented.

References

  1. MemPO: Self-Memory Policy Optimization for Long-Horizon Agents — Ruoran Li, Xinghua Zhang, Haiyang Yu, Shitong Duan, Xiang Li, Wenxin Xiang, Chonghua Liao, Xudong Guo, Yongbin Li, Jinli Suo, 2026. arxiv.org/abs/2603.00680
  2. Policy Gradient Methods for Reinforcement Learning with Function Approximation — Richard S. Sutton, David McAllester, Satinder Singh, Yishay Mansour, 1999/2000. Google Scholar
  3. High-Dimensional Continuous Control Using Generalized Advantage Estimation — John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, Pieter Abbeel, 2016. Google Scholar
  4. RUDDER: Return Decomposition for Delayed Rewards — Jose A. Arjona-Medina, Michael Gillhofer, Michael Widrich, Thomas Adler, Johannes Brandstetter, Sepp Hochreiter, 2019. Google Scholar
  5. Let's Verify Step by Step — Hunter Lightman, Vineet Kosaraju, Yura Burda, Harrison Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, Karl Cobbe, 2023. Google Scholar
  6. Information Gain-based Policy Optimization: A Simple and Effective Approach for Multi-Turn LLM Agents — Guoqing Wang, Sunhao Dai, Guangze Ye, Zeyu Gan, Wei Yao, Yong Deng, Xiaofeng Wu, Zhenzhe Ying, 2025. Google Scholar
  7. MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon Agents — Zijian Zhou, Ao Qu, Zhaoxuan Wu, Sunghwan Kim, Alok Prakash, Daniela Rus, Jinhua Zhao, Bryan Kian Hsiang Low, Paul Pu Liang, 2025. Google Scholar
  8. Reflexion: Language Agents with Verbal Reinforcement Learning — Noah Shinn, Federico Cassano, Edward Berman, Ashwin Gopinath, Karthik Narasimhan, Shunyu Yao, 2023. Google Scholar
  9. Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning — Bowen Jin, Hansi Zeng, Zhenrui Yue, Jinsung Yoon, Sercan Arik, Dong Wang, Hamed Zamani, Jiawei Han, 2025. Google Scholar
  10. Attnpo: Attention-guided Process Supervision for Efficient Reasoning — Shuaiyi Nie, Siyu Ding, Wenyuan Zhang, Linhao Yu, Tianmeng Yang, Yao Chen, Tingwen Liu, Weichong Yin, Yu Sun, Hua Wu, 2026. Google Scholar
  11. MemGPT: Towards LLMs as Operating Systems — Charles Packer, Sarah Wooders, Kevin Lin, Vivian Fang, Shishir G. Patil, Ion Stoica, Joseph E. Gonzalez, 2024. Google Scholar