AI Post Transformers · Visual Companion

MEMSEARCHER: Reinforcement Learning for LLM Memory Management

A visual-first exploration of how bounded learned memory changes the search-agent loop: less full-trajectory replay, more explicit state compression, and reinforcement learning that treats memory updates as part of policy.
arXiv: 2511.02805 2025 paper Search agents · RL · memory compression Transcript IDs found: 2511.02805

From transcript replay to compact state

MemSearcher replaces “carry the whole story forever” with a loop that repeatedly rewrites a bounded memory. Toggle the mode to see how the state passed into the next turn changes.

Search agent pipeline
Per-turn context payload
user objective compact memory thought/action/obs transcript redundant/noisy residue
~constant
MemSearcher turn context
linear ↑
ReAct transcript growth
learned
What survives across turns
bounded
Reasoning memory budget

Memory retention vs clutter accumulation

Hover the heatmap to inspect which facts survive. The matrix contrasts full-history carry-forward with iterative compression under a fixed budget.

Turn × item retention heatmap
forgotten / absent retained high salience / overload
Compression frontier

Mock benchmark and systems view

Illustrative data mirrors the episode’s claims: stronger average accuracy, lower per-turn token load, and a notable “3B beats some 7B baselines” systems narrative.

7-benchmark comparison bars
Token growth across turns

Multi-context GRPO as trajectory-level credit flow

Each turn creates a different prompt-state. The paper’s idea is to propagate a trajectory advantage signal across these context-conditioned sub-conversations instead of treating the run as one flat prompt.

Credit assignment flow
Reviewer skepticism radar

References

Compact source map for papers and adjacent threads mentioned in the episode.
MemSearcher — Yuan et al., 2025 · arXiv:2511.02805
ReAct — Yao et al., 2023 · Scholar
Search-R1 — Jin et al., 2025 · Scholar
PPO — Schulman et al., 2017 · Scholar
Self-Correct via RL — Zhang et al., 2024 · Scholar
Reflexion — Shinn et al., 2023 · Scholar
Generative Agents — Park et al., 2023 · Scholar
Context Compression Threads — ACon / Active Context Compression / Tiered Memory · Scholar
Credit Assignment Threads — GRPO-, InT, CAPO · Scholar
AgentGym-RL — long-horizon RL for agents · Scholar