After Titans: Behrouz on Nested Learning and Hope

A visual map of the Titans → Nested Learning → Hope arc: nested optimizers, self-modifying memory, and a continuum of update rates. The page emphasizes mechanism, compression flow, and what the paper’s evidence appears to support.

Arc

Titans reframed long-term memory as learnable at test time. Nested Learning broadens that bet into a hierarchy of learners, each compressing a different stream and updating on a different clock.
3 core contributions tracked below
4 timescales shown in Hope continuum
2 paper IDs extracted from episode context

Source Footprint

Primary paper: Behrouz, Razaviyayn, Zhong, Mirrokni, NeurIPS 2025
nested optimization optimizer-as-memory self-modification continuum memory Hope

From Depth to Nested Clocks

Hover any module to inspect what stream it compresses, how quickly it changes, and where the paper claims novelty.

Compression Flow

The visual follows the episode’s framing: activations compress tokens, optimizer state compresses gradients, and memory modules compress experience. The research bet is that these are not side effects but designable learners nested inside one another.
fast local state learned update rule slow persistent memory

Episode Tension

The hosts’ split shows up here too: the mechanism becomes more concrete as you move from lens → module → system, but the theory still outruns the diagnostics. The page keeps both: architecture on the left, skepticism encoded in the evidence tab.

Optimizer State as Associative Memory

Switch optimizers to compare state capacity, horizon, and the shape of compressed gradient traces. Hover cells to inspect recoverability.

mode

Interpretation vs Measurement

The transcript presses on a key gap: “state stores information” is obvious, but “state functions as usable associative memory” needs direct recall diagnostics. The heatmap shows a mock retrieval profile; its purpose is to make that missing measurement concrete.

Why Richer State Matters

The paper’s intuition is that tiny first- and second-moment summaries bottleneck adaptation. More expressive state can preserve longer, more structured traces of gradient history, potentially making the update path itself part of the learned architecture.

Hope: Self-Modifying Memory Across a Continuum

Move step-by-step through the writable module, then compare memory-band activity under different workloads.

step
workload

What Changes

This is not unrestricted whole-model self-rewriting. The writable part is a bounded module whose update rule is itself learned end-to-end. The stepper exposes that distinction visually.

Why the Continuum Matters

The episode rejects a single short-term vs long-term split. Hope’s story is a spectrum: some state changes nearly every step, some drifts, some persists. The band chart turns that qualitative claim into a visible rate profile.

Evidence Lab: Wins, Controls, and Missing Diagnostics

Toggle the comparison view to see how the same result can read as mechanism, extra state, or extra compute depending on the control.

compare

Results Framing

The mock bar chart mirrors the episode’s stance: Hope looks promising on adaptation-heavy tasks, but the headline chart is not the same thing as clean attribution. When budgets are equalized, some of the margin narrows.

Tests the Hosts Want

Corrupt memory bands, fix writable-state budget, hold recurrence and wall-clock compute fixed, and decode what facts survive inside state across time. The matrix on the right visualizes those stress axes as an experimental dashboard the paper still owes.

References

Nested Learning: The Illusion of Deep Learning Architectures
Ali Behrouz, Meisam Razaviyayn, Peilin Zhong, Vahab Mirrokni. NeurIPS 2025. arXiv:2512.24695
Titans: Learning to Memorize at Test Time
Ali Behrouz, Peilin Zhong, Vahab Mirrokni. January 2025. arXiv:2501.00663
Episode Link
AI Post Transformers: “Titans: Learning to Memorize at Test Time” podcast audio