A self-memory policy optimization framework that trains a long-horizon agent to write its own compressed memory note, as a learned RL action, instead of bolting on retrieval after the fact.
arXiv:2603.00680Ruoran Li, Xinghua Zhang, Haiyang Yu, et al. — 10 authorsTsinghua University · Alibaba Tongyi Lab2026
ReAct vs. MemPO: what actually travels between steps
ReAct (Yao et al., 2022) appends every tool result and reasoning trace to the prompt, so context grows every round. MemPO discards all of that and carries forward only a single compressed memory note.
Context size over a 10-step trajectory
Hover any point for the exact token count at that step. ReAct's context climbs steadily toward the window limit; MemPO's stays flat because only the note carries forward.
The mem / think / tool_call decomposition
Every step writes a memory note, reasons about what to do next, then acts. Only the note from step t-1 becomes step t's input — the reasoning trace and raw tool output are discarded.
Compression ratio per step
A single step can involve thousands of tokens of tool output and reasoning. Almost all of it is thrown away the moment the note is written.
Trajectory-level advantage (AT) from GRPO
Sixteen rollouts answer the same question. Reward is 1 for a clean, correct trajectory and 0 otherwise. Advantage is each rollout's reward normalized against the group's mean and standard deviation. Hover a cell for the exact values.
Memory-level advantage (AM): grading a note by its own confidence
Show only the note to the model and have it try to produce the ground-truth answer. A confident, correct guess means the note captured what mattered; subtracting the same probability conditioned on full history isolates the note's marginal contribution.
mem tokens get A_T + A_M · every other token gets A_T only
Averaged over 2–10 chained objectives on local Wikipedia search: MemPO 37.63 F1, GRPO with no memory reward 30.53, MEM1 26.34.
Against the strongest baselines, MemPO cuts total tokens 67.58% and peak tokens 73.12% — on the hardest setting it spends roughly what ReSearch spends on a 4-objective task.
Context-truncation ablation
Full history wins on short-horizon tasks; as objective count climbs, one-step truncation overtakes it — the same lost-in-the-middle effect, this time self-inflicted by hoarding old memory.
Isolating the memory reward (Figure 3, left panel)
GRPO with no dedicated memory reward already beats MEM1 by several points using the same warm start and truncated context. The reward mechanism — the one thing MemPO adds on top of that — is real but smaller than the headline gap over MEM1 suggests.
The abstract's "7.1 points over previous SOTA" bundles the GPT-4.1 warm-start distillation and the truncated-context recipe together with the memory reward. The isolated ablation — reward on vs. reward off, everything else held fixed — is the real marginal effect of what the paper claims to have invented.
Policy Gradient Methods for Reinforcement Learning with Function Approximation — Richard S. Sutton, David McAllester, Satinder Singh, Yishay Mansour, 1999/2000. Google Scholar
High-Dimensional Continuous Control Using Generalized Advantage Estimation — John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, Pieter Abbeel, 2016. Google Scholar
RUDDER: Return Decomposition for Delayed Rewards — Jose A. Arjona-Medina, Michael Gillhofer, Michael Widrich, Thomas Adler, Johannes Brandstetter, Sepp Hochreiter, 2019. Google Scholar
Let's Verify Step by Step — Hunter Lightman, Vineet Kosaraju, Yura Burda, Harrison Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, Karl Cobbe, 2023. Google Scholar
Information Gain-based Policy Optimization: A Simple and Effective Approach for Multi-Turn LLM Agents — Guoqing Wang, Sunhao Dai, Guangze Ye, Zeyu Gan, Wei Yao, Yong Deng, Xiaofeng Wu, Zhenzhe Ying, 2025. Google Scholar
MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon Agents — Zijian Zhou, Ao Qu, Zhaoxuan Wu, Sunghwan Kim, Alok Prakash, Daniela Rus, Jinhua Zhao, Bryan Kian Hsiang Low, Paul Pu Liang, 2025. Google Scholar
Reflexion: Language Agents with Verbal Reinforcement Learning — Noah Shinn, Federico Cassano, Edward Berman, Ashwin Gopinath, Karthik Narasimhan, Shunyu Yao, 2023. Google Scholar
Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning — Bowen Jin, Hansi Zeng, Zhenrui Yue, Jinsung Yoon, Sercan Arik, Dong Wang, Hamed Zamani, Jiawei Han, 2025. Google Scholar
Attnpo: Attention-guided Process Supervision for Efficient Reasoning — Shuaiyi Nie, Siyu Ding, Wenyuan Zhang, Linhao Yu, Tianmeng Yang, Yao Chen, Tingwen Liu, Weichong Yin, Yu Sun, Hua Wu, 2026. Google Scholar
MemGPT: Towards LLMs as Operating Systems — Charles Packer, Sarah Wooders, Kevin Lin, Vivian Fang, Shishir G. Patil, Ion Stoica, Joseph E. Gonzalez, 2024. Google Scholar