AI Post Transformers Visual Companion

Titans: Learning to Memorize at Test Time

This page treats Titans as a systems-design argument in pictures: exact short-range attention, a writable long-term memory, and a test-time update rule driven by surprise. The main question is not whether memory is desirable, but whether mutable neural memory beats simpler alternatives like bigger windows, retrieval, or recurrent cache reuse once deployment constraints show up.

Theme
Inference-time memory writes as small online learning steps
Core split
Attention = precise near field; memory = persistent fuzzy far field
Transcript scan
Additional arXiv IDs mentioned explicitly in transcript: none detected
Surprise-driven writes 340M / 400M / 760M 4K train context Needle + BABILong

Short-Term Precision vs Long-Term Persistence

The architecture below shows the paper’s pitch: keep local attention for crisp recall, but route surprising events into a separate learned memory that survives beyond the active window. Hover nodes and links to see what changes when the system swaps from pure attention to a hybrid memory path.

Interpretation

The learned memory is not a passive cache. In the Titans framing, writes happen through a loss-based update rule, so memory behaves more like a tiny optimizer state than like ordinary recurrent hidden state.

Why it matters

This moves the long-context question from “how big is the window?” to “what deserves to survive?” The operational tension is that writable parameters are harder to batch, inspect, reset, and isolate than ordinary KV state.

Mode contrast

Toggle the diagram to compare a plain attention stack against the hybrid path. The baseline keeps exact local recall but drops old information once it falls outside the active segment.

Local recall
0.96
Far-context retention
0.78
Serving simplicity
0.42
active attention path
learned memory path
surprise-triggered write
deployment friction

Surprise-Driven Memory Writes

Titans’ distinct claim is that memory updates are scaled by surprise. This section visualizes a toy token stream, a write heatmap over memory slots, and the evolving optimizer-like state over four steps.

Update intuition

Large prediction error increases gradient magnitude, which makes a write stronger. Forgetting and momentum keep the memory from chasing every token equally hard.

Why this is not a cache

A cache stores exact entries. This mechanism mutates parameters, so multiple events interfere, blend, and decay over time.

Current step

A routine local phrase passes through attention with only a light memory write.

Surprise
0.18
Write strength
0.22
Forgetting gate
0.86
low to high write energy
token surprise curve
memory slots

Tradeoff Surface: Accuracy, Length, and Operational Cost

These are illustrative mock numbers shaped by the episode’s claims. The chart compares families of long-context solutions across sparse retrieval, reasoning retention, and infrastructure friction. Switch views to see how the ranking changes when the objective changes.

Families shown

Attention-only long windows, retrieval-augmented systems, KV-cache recurrence, state-space or linear recurrent models, and Titans-style writable memory. The point is not exact leaderboard numbers but the shape of the design space.

Episode takeaway

Titans sits in an appealing middle zone: stronger far-context behavior than pure recurrence, more internal integration than retrieval, but more mutable-state burden than either.

The largest uncertainty is scaling evidence. Medium-scale wins do not settle whether writable memory beats stronger baselines at frontier budgets.
Best sparse retrieval
Titans
Cheapest serving
KV-Fold
Best ops hygiene
Retrieval

Genealogy of Long-Context Memory Ideas

This map places Titans inside a broader family tree: explicit external memory, segment recurrence, compressed memory, retrieval stores, state-space recurrence, and test-time learners. Hover any node to see the paper’s role in the memory tradeoff.

2014 → 2026 Memory vs cost vs controllability

References

Compact source list for the episode and nearby comparison points.

Titans: Learning to Memorize at Test TimearXiv:2501.00663
Neural Turing MachinesarXiv:1410.5401
Transformer-XLarXiv:1901.02860
Compressive TransformersOpenReview
Memorizing TransformersarXiv:2203.08913
Learning to (learn at test time): RNNs with expressive hidden statesScholar link
Gated Delta NetworksScholar link
BABILongScholar link