Memory Intelligence Agents for Deep Research

A visual explainer of MIA’s core claim: combine non-parametric memory (compressed external search experience) with parametric memory (planner skill in weights) so a research agent reuses prior investigations without drowning in raw trajectories.

arXiv: 2604.04503 2026 paper Manager · Planner · Executor Test-time self-evolution Alternating RL
11
benchmarks discussed
9%
reported max gain on LiveVQA
6%
reported gain on HotpotQA
31%
avg gain with lightweight executor
Visuals below use mock-but-plausible data to illustrate the paper’s scaling and architecture ideas.

MIA Architecture Flow

Three specialized roles split the workload: the Manager curates compressed experience, the Planner turns question + memory into strategy, and the Executor interacts with tools and evidence.

parametric memory / learned planning non-parametric memory / external bank tool interaction / execution feedback / judgment / updates

Role Separation

Hover boxes and links for state load, update frequency, and auditability tradeoffs.

Compression vs Raw Trajectory Memory

Existing agent memory often accumulates long episodes verbatim. MIA’s pitch is to retain planning-relevant structure while reducing retrieval clutter and storage burden.

Trajectory Density Heatmap

low signal medium high / clutter

What Survives Compression

A conceptual schema: not every thought survives, but task decomposition, productive tools, evidence paths, and failure recovery do.

Retrieval Signal Mix

Performance Gains vs Efficiency Evidence

The episode highlights a tension: the paper seems strongest on benchmark performance and weaker on hard operational efficiency curves such as store growth, latency, and context cost.

Toggle between benchmark improvements and notional memory economics over long deployments.

Claim Strength Radar

strongly supported moderate under-evidenced

Test-Time Self-Evolution

The core loop is retrieval → planning → execution → reflection → judgment → update. This tab shows the progressive migration between external episodes and planner skill.

Parametric ↔ Non-Parametric Balance

Conceptual trend: explicit memory dominates early, then planner internalization rises while external memory remains a searchable safety net.

Reference Constellation

Compact map of the cited ideas around memory, reflection, RAG, long-horizon agents, and test-time adaptation.

Memory Intelligence Agent

Qiao et al., 2026 · 2604.04503

RAG

Lewis et al., 2020 · external non-parametric retrieval

Generative Agents

Park et al., 2023 · observation → reflection → planning

Voyager

Wang et al., 2023 · reusable skill accumulation

Reflexion / Self-Refine

self-critique and iterative improvement

MemGPT / LongMem / MemoryBank-style

long-term agent memory systems

Parametric vs Non-Parametric Memory Studies

2024 evaluation & interpretability papers

Related Podcast Episodes

MemSearcher, DeepVerifier, Mem0, DeepSeek Engram, MetaClaw