This page treats Titans as a systems-design argument in pictures: exact short-range attention, a writable long-term memory, and a test-time update rule driven by surprise. The main question is not whether memory is desirable, but whether mutable neural memory beats simpler alternatives like bigger windows, retrieval, or recurrent cache reuse once deployment constraints show up.
The architecture below shows the paper’s pitch: keep local attention for crisp recall, but route surprising events into a separate learned memory that survives beyond the active window. Hover nodes and links to see what changes when the system swaps from pure attention to a hybrid memory path.
The learned memory is not a passive cache. In the Titans framing, writes happen through a loss-based update rule, so memory behaves more like a tiny optimizer state than like ordinary recurrent hidden state.
This moves the long-context question from “how big is the window?” to “what deserves to survive?” The operational tension is that writable parameters are harder to batch, inspect, reset, and isolate than ordinary KV state.
Toggle the diagram to compare a plain attention stack against the hybrid path. The baseline keeps exact local recall but drops old information once it falls outside the active segment.
Titans’ distinct claim is that memory updates are scaled by surprise. This section visualizes a toy token stream, a write heatmap over memory slots, and the evolving optimizer-like state over four steps.
Large prediction error increases gradient magnitude, which makes a write stronger. Forgetting and momentum keep the memory from chasing every token equally hard.
A cache stores exact entries. This mechanism mutates parameters, so multiple events interfere, blend, and decay over time.
A routine local phrase passes through attention with only a light memory write.
These are illustrative mock numbers shaped by the episode’s claims. The chart compares families of long-context solutions across sparse retrieval, reasoning retention, and infrastructure friction. Switch views to see how the ranking changes when the objective changes.
Attention-only long windows, retrieval-augmented systems, KV-cache recurrence, state-space or linear recurrent models, and Titans-style writable memory. The point is not exact leaderboard numbers but the shape of the design space.
Titans sits in an appealing middle zone: stronger far-context behavior than pure recurrence, more internal integration than retrieval, but more mutable-state burden than either.
This map places Titans inside a broader family tree: explicit external memory, segment recurrence, compressed memory, retrieval stores, state-space recurrence, and test-time learners. Hover any node to see the paper’s role in the memory tradeoff.
Compact source list for the episode and nearby comparison points.