← All episodes Learning at Test Time with Expressive RNN States

Learning at Test Time with Expressive RNN States

Jun 11, 2026
This episode explores the paper Learning to (Learn at Test Time): RNNs with Expressive Hidden States and its attempt to give recurrent models transformer-like long-context behavior without the quadratic cost of attention. It explains why standard RNN hidden states are a bottleneck, compares that limitation to transformers’ growing KV cache, and highlights a key empirical motivation: in the paper’s setup, Mamba’s token-level perplexity improvements flatten around 16k tokens while transformers keep improving deeper into a 32k context. The discussion focuses on the paper’s core idea of test-time training, where the hidden state is treated as a small inner model whose parameters are updated online with a self-supervised learning rule, rather than as a fixed vector summary. Listeners would find it interesting because it connects old fast-weights and dynamic-evaluation ideas to a new systems-level proposal for long-context efficiency, while also noting the open question of whether better perplexity truly translates into stronger retrieval and reasoning.
Sources:
1. Learning to (Learn at Test Time): RNNs with Expressive Hidden States — Yu Sun, Xinhao Li, Karan Dalal, Jiarui Xu, Arjun Vikram, Genghan Zhang, Yann Dubois, Xinlei Chen, Xiaolong Wang, Sanmi Koyejo, Tatsunori Hashimoto, Carlos Guestrin, 2024
http://arxiv.org/abs/2407.04620
2. Using Fast Weights to Attend to the Recent Past — Jimmy Ba, Geoffrey Hinton, Volodymyr Mnih, Joel Z. Leibo, Catalin Ionescu, 2016
https://arxiv.org/abs/1610.06258
3. Linear Transformers Are Secretly Fast Weight Programmers — Imanol Schlag, Kazuki Irie, Jürgen Schmidhuber, 2021
https://arxiv.org/abs/2102.11174
4. Mamba: Linear-Time Sequence Modeling with Selective State Spaces — Albert Gu, Tri Dao, 2023
https://arxiv.org/abs/2312.00752
5. Learning to (Learn at Test Time): RNNs with Expressive Hidden States — Yu Sun, Xinhao Li, Karan Dalal, Xiaolong Wang, Tatsunori Hashimoto, Carlos Guestrin, 2024
https://arxiv.org/abs/2407.04620
6. Dynamic Evaluation of Transformer Language Models — Ben Krause, Emmanuel Kahembwe, Iain Murray, Steve Renals, 2019
https://arxiv.org/abs/1904.08378
7. Effective Long-Context Scaling of Foundation Models — Wenhan Xiong et al., 2023
https://arxiv.org/abs/2309.16039
8. Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models — Soham De et al., 2024
https://arxiv.org/abs/2402.19427
9. Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality — Tri Dao, Albert Gu, 2024
https://arxiv.org/abs/2405.21060
10. An Empirical Study of Mamba-based Language Models — Roger Waleffe et al., 2024
https://arxiv.org/abs/2406.07887
11. KVzip: Query-Agnostic KV Cache Compression with Context Reconstruction — Jang-Hyun Kim et al., 2025
https://arxiv.org/abs/2505.23416
12. KVLink: Accelerating Large Language Models via Efficient KV Cache Reuse — Jingbo Yang et al., 2025
https://arxiv.org/abs/2502.16002
13. ReMamba: Equip Mamba with Effective Long-Sequence Modeling — Danlong Yuan et al., 2024
https://arxiv.org/abs/2408.15496
14. LongMamba: Enhancing Mamba's Long Context Capabilities via Training-Free Receptive Field Enlargement — Zhifan Ye et al., 2025
https://arxiv.org/abs/2504.16053
15. Fast-weight Product Key Memory — Tianyu Zhao, Llion Jones, 2026
https://arxiv.org/abs/2601.00671
16. Test-Time Learning for Large Language Models — Jinwu Hu et al., 2025
https://arxiv.org/abs/2505.20633
17. Test-Time Training on Nearest Neighbors for Large Language Models — Moritz Hardt, Yu Sun, 2023
https://arxiv.org/abs/2305.18466
18. AI Post Transformers: Titans: Learning to Memorize at Test Time — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-20-titans-learning-to-memorize-at-test-time-054662.mp3
19. AI Post Transformers: Gated Linear Attention for Efficient Long Sequences — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-18-gated-linear-attention-for-efficient-lon-c858ab.mp3
20. AI Post Transformers: Kimi Linear: Efficient Expressive Attention Architecture — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/kimi-linear-efficient-expressive-attention-architecture/
21. AI Post Transformers: How Induction Heads Emerge in Transformers — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-03-how-induction-heads-emerge-in-transforme-a7bfcb.mp3
22. AI Post Transformers: Explicit Information Transmission for Context Compression — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-explicit-information-transmission-for-co-24e3c2.mp3
23. AI Post Transformers: Memory-Bound, Not Bandwidth-Limited Batch-1 LLM Decode — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-02-memory-bound-not-bandwidth-limited-batch-114799.mp3
Interactive Visualization: Learning at Test Time with Expressive RNN States