AI Post Transformers · Episode Companion

Adapting Without Forgetting

A Lifelong Learning Roadmap for LLM Agents — how agents that never update their weights still find ways to forget, and what a 2025 survey does and doesn't prove about fixing it.

arXiv:2501.07278 Zheng, Shi, Cai, Li, Zhang, Li, Yu, Ma · 2025 South China University of Technology · MBZUAI · Tencent AI Lab

The survey welds two lineages together — internal-knowledge LLM surveys and agent-architecture surveys — into one loop: perception, memory, and action, formalized as a goal-conditioned POMDP. Click a memory block to see how each type is actually used.

Perception → Memory → Action loop

Click a memory block in the diagram above — Working, Episodic, Semantic, or Parametric — to see how it's used in practice.

Mock retrieval-activation intensity across a 10-task episode, by memory type. Episodic memory climbs as an agent's skill library grows (the Voyager pattern); parametric memory stays cold except when a knowledge-editing pass fires. Hover a cell for the story behind the number.

Memory activation by task step

Freezing the backbone doesn't remove the stability-plasticity dilemma — it relocates it. Toggle between an EWC-style protected update and an open update to see which weights move, then compare what that tradeoff does to performance across a task sequence.

Weight update pattern — Stability Mode

Average Performance (AP) across a 10-task sequence

The survey inherits five formal metrics from classical continual learning — AP, AIP, FGT, BWT, FWT — but the systems it cites report free-form success rates instead. A taxonomy checkmark is not evidence a system was ever tested for forgetting.

0
of ~360 cited works have a table where the survey computes FGT or BWT for a named system. Confucius, ART, and GITM get the same taxonomy checkmark as papers actually designed to measure forgetting.

Taxonomy membership vs. metrics actually reported

References