Hal: Memory-Augmented LLM Agents Still Hit Continual Learning's Wall

arXiv:2604.27003 Qisheng Hu, Quanyu Long, Wenya Wang · NTU ALFWorld · BabyAI · ReMe · BM25

Memory-augmented LLM agents were pitched as a way to dodge catastrophic forgetting entirely — don't retrain, just remember. This paper runs the classic continual-learning protocol on that setup and finds the same bottleneck resurfaces, relocated from parameter capacity to retrieval capacity: retrieval pollution, context competition, and memory dilution.

Two Continual-Learning Pipelines, Same Shaped Failure

Parametric CL overwrites weights; Memory CL overwrites nothing — it just runs out of room to retrieve. Hover the bottleneck boxes.

parametric path memory path bottleneck

Relocation Table

The paper's core claim, as data: the carrier and interference mechanism change, the shape of the problem does not.

Retrieval Pollution

A BM25 query pulls back top-k memory entries — some relevant to the current task, most not. Hover a slot.

Context Competition

Toggle memory-store size. As more entries compete for the same fixed context window, useful memories get displaced.

Memory Dilution Over Time

Heatmap: rows are memory slots, columns are time steps as the pool grows. Color = probability a query for old content still surfaces this slot. Hover any cell.

easily retrieved competing buried / unreachable

Study 1 — Representation: Raw Trajectory vs Insight (Forward Transfer, A → B)

Detail that helps within one task actively misleads a different task. Toggle environment.

Raw Insight

ALFWorld: Easy vs Hard Subset (Raw, A→B)

Raw's detail barely hurts easy cases — it wrecks the hard ones.

BabyAI Backward Transfer (B→A)

Insight cuts forgetting too, not just forward transfer.

Study 2 — Organization: Cond-Agg vs Cond-Ind vs Cond-Step (BabyAI, Insight fixed)

Toggle direction. The condition that adapts best to the new task can damage the old one most, in the same run.

Forward Transfer Backward Transfer

Retrieval Diversity Collapse

Cond-Ind's fine-grained index only helps if the underlying content is actually diverse. On repetitive content, near-identical entries funnel every query into the same handful of items.

The (k, v) Decomposition

Every memory design is a value (how experience is written) crossed with a key (how it's indexed). Hover a cell for its definition; highlighted cells are the ones actually tested.

References