A frozen transformer gets a tiny mutable state matrix, updates it with a delta rule, and feeds the readout back as a low-rank attention correction. The claim is not “more context,” but “actual online memory.”
Instead of replaying a growing transcript, the model keeps a compact evolving matrix and injects its readout into attention. The SVG below contrasts prompt hauling with latent memory steering.
Hover any cell. The matrix shows a mock latent state as it is read, corrected by residual error, and selectively decayed by a forget gate.
This ordering is the paper’s sharpest systems point. The model uses old memory to shape current attention, then updates the state after seeing new evidence.
The chart groups memory-centric benchmarks apart from memory-adjacent capability checks. The largest lifts appear where persistence and incremental updating are directly stressed.
The scatterplot maps methods by how tightly they are integrated into model computation and how mutable they are at inference time. Bubble size approximates state or storage footprint.
Compact latent memory is elegant, but explicit stores are easier to inspect, delete, and audit. That operational tradeoff is the main caution flag for real assistants.