The survey welds two lineages together — internal-knowledge LLM surveys and agent-architecture surveys — into one loop: perception, memory, and action, formalized as a goal-conditioned POMDP. Click a memory block to see how each type is actually used.
Perception → Memory → Action loop
Mock retrieval-activation intensity across a 10-task episode, by memory type. Episodic memory climbs as an agent's skill library grows (the Voyager pattern); parametric memory stays cold except when a knowledge-editing pass fires. Hover a cell for the story behind the number.
Memory activation by task step
Freezing the backbone doesn't remove the stability-plasticity dilemma — it relocates it. Toggle between an EWC-style protected update and an open update to see which weights move, then compare what that tradeoff does to performance across a task sequence.
Weight update pattern — Stability Mode
Average Performance (AP) across a 10-task sequence
The survey inherits five formal metrics from classical continual learning — AP, AIP, FGT, BWT, FWT — but the systems it cites report free-form success rates instead. A taxonomy checkmark is not evidence a system was ever tested for forgetting.