1 00:00:01,000 --> 00:00:36,524 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today we're digging into MemPO: Self-Memory Policy Optimization for Long-Horizon Agents. First author Ruoran Li — ten co-authors total, so, Ruoran Li et al. — out of Tsinghua University and Tongyi Lab at Alibaba Group, posted to arXiv in June 2026. The one-sentence version: can you train an agent to write its own working memory as it goes, using reinforcement learning, instead of bolting on a separate retrieval system afterward? 2 00:00:36,524 --> 00:00:57,274 [Dr. Ada Shannon] What actually pulled me in wasn't the headline numbers, Hal — it's the theoretical footing underneath it. Most agent-memory work feels like plumbing: store things, retrieve things, hope for relevance. This one derives its memory signal from something principled, the model's own confidence in the answer given what it kept, and wires that straight into the RL objective. That's a real design choice, not just an engineering patch. 3 00:00:57,274 --> 00:01:21,924 [Hal Turing] So here's the setup. An agent working a question that needs, say, ten rounds of searching and reasoning before it has enough to answer. Standard practice: every tool call, every result, every thought gets appended to the prompt and carried into the next round. By round ten your context has ballooned, you're burning way more tokens per step, and you're pushing on whatever window limit your model actually has. 4 00:01:21,924 --> 00:01:48,574 [Dr. Ada Shannon] That accumulate-everything loop is called ReAct — Yao and colleagues introduced it in 2022 out of Princeton and Google Research, and it's still the dominant agent pattern: think, act, observe, append, repeat. It was never designed with cost in mind, and it's not just expensive. Liu and colleagues out of Stanford showed in their 2023 'Lost in the Middle' paper that models get measurably worse at using information buried in the center of a long context. A bloated history isn't neutral — it actively hurts reasoning. 5 00:01:48,574 --> 00:02:08,474 [Hal Turing] Oh, wait, hold on — I want to push on that, because I'm not sure this is even still a problem in 2026. We've got models advertising million-token context windows now. Why not just let the agent hoard everything from every round and trust the model to sort out what actually matters when it answers? 6 00:02:08,474 --> 00:02:29,275 [Dr. Ada Shannon] I actually disagree with you there, Hal. A bigger window isn't cheap and it isn't automatically effective — every token in it gets attended over again at every single turn, so ten rounds in a huge window means paying that cost repeatedly, not once. And lost-in-the-middle doesn't disappear with scale, it gets worse, because there's more haystack for the needle to hide in. Bigger windows raise the ceiling, they don't fix the failure mode. 7 00:02:29,275 --> 00:03:06,875 [Hal Turing] Okay, fair — cost and quality are two separate failure modes, and scale alone doesn't fix either. So the field's actual answer: don't dump everything into the prompt, give the agent an external memory store. MemGPT, from Packer and colleagues out of UC Berkeley in 2024, treats memory like an OS's virtual memory hierarchy. Mem0, from Chhikara and colleagues at Mem0 in 2025, does dynamic extraction and consolidation of conversation history. Both pull back relevant chunks by embedding similarity when needed. 8 00:03:06,875 --> 00:03:39,599 [Dr. Ada Shannon] Right, and that's exactly where MemPO's authors plant their flag. The underlying idea traces back to Lewis and colleagues at Facebook AI Research in 2020 — retrieve a chunk, stuff it in the prompt, generate. But it's offline and passive: retrieval runs on embedding similarity to the query, not what actually helps solve the task. Nobody's optimizing retrieval or summary against task success. So MemPO's question is almost embarrassingly direct — what if the agent wrote its own memory, as a learned action, trained end-to-end toward getting the right answer? 9 00:03:39,599 --> 00:04:02,750 [Hal Turing] And that's the part that actually got me curious, because 'train it to summarize well' sounds simple until you ask how you'd reward it. You don't know a memory note was good or bad until the whole trajectory finishes and you see if the final answer was right. So how do you turn one pass-or-fail signal at the end of ten rounds into a training signal for something that happened at step three? 10 00:04:02,750 --> 00:04:42,574 [Dr. Ada Shannon] That's the credit assignment problem — the term traces back to Minsky's 1961 essay on steps toward artificial intelligence, and it's central to RL generally. Supervised learning gives you an exact gradient to every weight; RL doesn't, because the environment sits between action and outcome and isn't differentiable, so you need a proxy for how much any one action mattered. MemPO builds on GRPO, Group Relative Policy Optimization, from Shao and colleagues' DeepSeekMath paper out of DeepSeek in 2024 — instead of a separate value network, it samples a batch of trajectories for the same question and scores them against each other, so the group's relative performance becomes the training signal. 11 00:04:42,574 --> 00:05:03,649 [Hal Turing] Okay, so GRPO gives you one coarse signal per trajectory — everyone in a bad rollout gets blamed a little, everyone in a good one gets credit. But that still doesn't tell you which specific memory note, three steps in, was the good or bad call. That's the gap MemPO actually had to close, and it's worth walking through exactly how. 12 00:05:03,649 --> 00:05:30,924 [Dr. Ada Shannon] And it's worth digging into — the fix is narrow once you see the mechanics. Each step decomposes into three tagged actions: mem, think, and tool_call. The model writes a memory note summarizing what matters, reasons in think about what to do next, then calls a tool or answers. Here's the piece that kills context growth: at inference, step t's input is just the memory from step t-1 — not raw tool responses, not earlier reasoning. Everything's discarded except what survived compression into that one note. 13 00:05:30,924 --> 00:05:53,349 [Hal Turing] Right — one note in, one note out, everything else gone. But training that's brutal. Sixteen rollouts per question, ten-plus steps each, and at the end you get one bit: right or wrong. That's the trajectory-level piece from the credit assignment problem we hit earlier. So how does MemPO turn that single end-of-trajectory bit into something usable at every step? 14 00:05:53,349 --> 00:06:21,174 [Dr. Ada Shannon] The trajectory half is basically vanilla GRPO. Each rollout gets reward 1 if the format's clean and the answer's correct, 0 otherwise. Normalize the group of sixteen binary rewards by mean and standard deviation, and that's the trajectory-level advantage, AT — applied to every token in that trajectory. If trajectory five got it right and the others mostly didn't, every token in five, memory notes included, gets a positive nudge. That alone can't tell you which specific memory note was actually good. 15 00:06:21,174 --> 00:06:39,649 [Hal Turing] Which is the gap you flagged earlier — a good trajectory could still contain a lazy, half-useless memory note at step three that just got carried by better information gathered afterward. So there's a second signal specifically for the mem tokens. What's actually being measured there? 16 00:06:39,649 --> 00:07:11,799 [Dr. Ada Shannon] This is the clever part — RM comes from the model's own confidence. Take the note written at step t, show only that note to the model, and have it generate the ground-truth answer string. The probability it assigns becomes a proxy for the note's information content — high means the note alone was basically enough, low means it dropped something essential. They subtract a bias term, epsilon: the same probability but conditioned on everything through step t-1 instead of just the note, isolating the memory action's marginal contribution. 17 00:07:11,799 --> 00:07:31,074 [Hal Turing] Oh wait, hold on — that's elegant, you don't need a judge model or human labels, the policy grades its own compression by how well it primes itself to answer. But doesn't that make it self-referential? You're using the model's belief in its own summary as the ground truth for whether the summary's good. 18 00:07:31,074 --> 00:07:53,449 [Dr. Ada Shannon] Self-referential in that the same model judges, but not circular — it's scored against the real ground-truth string, not the model's own guess. A confident note that omits the answer still scores low. They normalize it the same way as AT, grouped across all memory notes at that step, giving you AM. The combination's simple: mem tokens get AT plus AM; everything else only gets AT. Memory tokens are the only place getting two stacked signals. 19 00:07:53,449 --> 00:08:08,374 [Hal Turing] Neat surgery. But before any of that RL turns on, the base model has to actually know how to produce mem, think, and tool_call tags — that's not something a vanilla Qwen checkpoint does out of the box. What's the on-ramp? 20 00:08:08,374 --> 00:08:45,449 [Dr. Ada Shannon] A behavior-cloning warm start. They run GPT-4.1 over an existing agent-trajectory dataset, keep only correct-answer trajectories, and get roughly ten thousand clean examples already in the mem-think-tool_call format. One epoch of fine-tuning, and the base model speaks the right dialect before RL touches it. Then: 7-billion-parameter Qwen2.5 base, sixteen rollouts per question, batch 128, trained on a synthesized multi-objective mix from HotpotQA and Natural Questions, using local Wikipedia search and, for the generalization check, live web search. 21 00:08:45,449 --> 00:09:00,049 [Hal Turing] Does it actually move the needle, though? "Trains a smarter compressor" is a nice idea, but the field's littered with agent-memory papers that look great on a whiteboard and barely beat baseline once you run the numbers. 22 00:09:00,049 --> 00:09:37,599 [Dr. Ada Shannon] On local wiki search, averaged across two through ten objectives, MemPO lands 37.63 F1. MEM1, the prior RL memory baseline, gets 26.34. GRPO trained identically but with no memory reward — same warm start, same rollout size, full context — scores 30.53. The token numbers matter more for production cost: MemPO cuts total tokens 67.58% and peak tokens 73.12% against the strongest baselines — a better answer for roughly a third of the token bill on the hardest setting. 23 00:09:37,599 --> 00:09:54,425 [Hal Turing] Real gap over MEM1. Though given inference only keeps the last step's note, I'd have guessed you'd lose something by discarding everything older — wouldn't two or three steps of raw history always beat one, since it's strictly more information sitting there? 24 00:09:54,425 --> 00:10:16,600 [Dr. Ada Shannon] I actually disagree, Hal — their ablation backs me up. They tested full context against three-step and one-step truncation. On short-horizon tasks more context edges out — fewer rounds, less dilution. But as objective count climbs, that flips hard: full context starts losing to the one-step version on long-horizon tasks. Same lost-in-the-middle effect Liu and colleagues described, just self-inflicted. 25 00:10:16,600 --> 00:10:27,575 [Hal Turing] Isn't that just an artifact of only testing one step, three steps, and full context — nothing in between? A smarter middle ground might beat both extremes. 26 00:10:27,575 --> 00:10:48,800 [Dr. Ada Shannon] Fair follow-up nobody ran. But the trend's monotonic across all three settings, and it lines up with the conditional-probability data: MemPO's P of answer given memory skews visibly higher than plain GRPO's in the grouped analysis, and accuracy climbs right along with it, bucket by bucket. Not proof the one-step cutoff is optimal — just that raw context volume isn't the lever doing the work here. 27 00:10:48,800 --> 00:11:32,475 [Hal Turing] Not proof, sure, but it's not nothing either. Before we close this out, I want to push on something in the numbers, not the story around them. GRPO without the memory reward — same warm-start data, same sixteen rollouts, same batch size, same RL recipe, just full context and no dedicated reward for the mem tokens — scores 30.53 average F1 on local wiki search. MEM1 scores 26.34. MemPO scores 37.63. If a version with zero memory-specific reward already beats MEM1 by four points, how much of that seven-point MemPO-over-MEM1 gap is actually the reward mechanism, versus just everything else being better? 28 00:11:32,475 --> 00:12:13,000 [Dr. Ada Shannon] That's the sharpest question in this paper, and to their credit there's one place they isolate it cleanly — Figure 3, left panel. Same base model, same GRPO training, same warm-start distillation, the only variable toggled is whether mem tokens get that dedicated reward on top of the trajectory advantage. That comparison, not the 37.63-versus-26.34 headline, is the real marginal effect of what they're claiming to have invented. Everything else — the GPT-4.1 distillation, the RL recipe, running on truncated context — is shared between GRPO-with-mem and GRPO-without-mem, so a good chunk of the gap over MEM1 has nothing to do with the reward design at all. 29 00:12:13,000 --> 00:12:36,275 [Hal Turing] Okay, but is that really a flaw in the underlying science, or just how every abstract gets written? They didn't bury GRPO-without-mem in an appendix — it's right there in Table 1 next to everything else, and the ablation showing the isolated effect is in the paper too. Feels like you're grading them for a marketing choice in the abstract, not an actual sin in the experiment design. 30 00:12:36,275 --> 00:13:12,350 [Dr. Ada Shannon] Wait — no, hold on, I don't think that's fair. The abstract is what people cite eighteen months from now — nobody's citing 'Figure 3, left panel.' When your whole pitch is 'we invented a dedicated memory reward,' and your own ablation shows that reward accounts for less than half your total advantage over the strongest baseline, that's the headline claim being oversold, not a rhetorical nitpick. I'll give them this: the mechanism is real, the reward-on line sits above reward-off at every objective count. But 'consistently a few points better' and '7.1 points over previous SOTA' are different sentences, and only one is technically true. 31 00:13:12,350 --> 00:13:37,375 [Hal Turing] Fair — I'll take that one. The mechanism holds up under isolation, it's just smaller than the marketing suggests, and those are meaningfully different claims. So let's hit the other soft spot: that memory reward is the probability of the ground-truth answer string given the memory, minus a bias term. That requires teacher-forcing an actual short, checkable answer. What happens the moment your task doesn't have one? 32 00:13:37,375 --> 00:14:21,400 [Dr. Ada Shannon] Then the whole reward collapses, and the paper never addresses it. Their own introduction motivates this with deep research, data analysis, vibe coding — none of which have a single short answer span you can teacher-force a probability against. A data-analysis agent's 'good memory' might be a retained table or a caveat about a confounding variable, with no gold string to condition on. They cite Wang and colleagues' Information Gain-based Policy Optimization paper, posted in 2025, as intuition for using conditional probability as an information signal — the closest prior precedent for this exact reward design — but there's no direct experimental comparison against it, just a passing citation. The validated claim is narrower than the framing implies: short-answer, search-heavy QA agents. Whether it works anywhere else is untested. 33 00:14:21,400 --> 00:14:48,025 [Hal Turing] Which matters more for anyone actually deploying this than the F1 number does. At ten objectives, MemPO uses roughly a third of ReSearch's tokens and about a fifth of its peak per-step tokens — comparable to what ReSearch spends on a four-objective task. If you're running this at scale, that's the line item that shows up on your bill every month, not a two-point F1 delta on a benchmark nobody outside the paper will ever query. 34 00:14:48,025 --> 00:15:19,725 [Dr. Ada Shannon] And the authors are honest about where this needs work — their limitations section flags that memory content isn't equivalent across steps or rollouts, since tool calls surface different amounts of information at different points, biasing the group-based normalization. They patch that with the epsilon bias term but admit more refined solutions are needed. Add the sixteen-step cap and Qwen2.5-7B being the only model tested, and 'future work, more applications' is basically the outline for the next paper. 35 00:15:19,725 --> 00:15:51,025 [Hal Turing] So where's the honest summary land? Real idea, real mechanism — teaching an agent to grade its own compression by how well it primes the right answer is genuinely clever, and the token savings are the kind that actually ship in production. But the headline gap over MEM1 needs that ablation, not the abstract, to be believed, and the reward only works where there's already a crisp answer to check against. Promising direction, oversold on the outside. That's MemPO — thanks for listening, everyone, we'll catch you next time.