AI Post Transformers • Interactive Visualization arXiv: 2605.12357 Posted: 2026-05-12

δ-mem and Online Memory for LLMs

A frozen transformer gets a tiny mutable state matrix, updates it with a delta rule, and feeds the readout back as a low-rank attention correction. The claim is not “more context,” but “actual online memory.”

Online state
8×8
Microscopic fast memory
Core ordering
read → steer → write
Prior memory shapes reasoning first
Reported lift
1.10×
vs frozen backbone average
TTL subtask
26.14 → 50.50
Memory-heavy jump
Architecture

Long context is a buffer. δ-mem tries to be a stateful pathway.

Instead of replaying a growing transcript, the model keeps a compact evolving matrix and injects its readout into attention. The SVG below contrasts prompt hauling with latent memory steering.

tiny mutable state attention correction token/context cost

What this page emphasizes

Persistence
Can a small state keep useful facts alive across many turns?
Conflict
Can later evidence overwrite stale state instead of appending noise?
Compression
The state stores residual error, not a miniature transcript.
Coupling
Memory enters generation as a low-rank routing nudge, not plain text.
Benchmarks like LoCoMo and MemoryAgentBench matter most here because they test persistence, incremental updating, and conflict resolution across turns.
Technical Deep Dive

Interactive 8×8 memory state

Hover any cell. The matrix shows a mock latent state as it is read, corrected by residual error, and selectively decayed by a forget gate.

low / cold mid / active high / hot

Read, steer, then write

This ordering is the paper’s sharpest systems point. The model uses old memory to shape current attention, then updates the state after seeing new evidence.

Results

Memory-heavy tasks separate memory claims from generic capability claims

The chart groups memory-centric benchmarks apart from memory-adjacent capability checks. The largest lifts appear where persistence and incremental updating are directly stressed.

Signal strength by benchmark type

Strong evidence: LoCoMo, MemoryAgentBench.
Secondary evidence: HotpotQA, IFEval, GPQA-Diamond.
Design Space

Where δ-mem sits among memory strategies

The scatterplot maps methods by how tightly they are integrated into model computation and how mutable they are at inference time. Bubble size approximates state or storage footprint.

Practical tension

Compact latent memory is elegant, but explicit stores are easier to inspect, delete, and audit. That operational tradeoff is the main caution flag for real assistants.

Pros
Tiny state, low replay cost, direct steering of attention.
Risks
Opaque edits, overwrite limits, unclear scaling law beyond reported setups.
References

Sources and adjacent work