This episode explores TRELLIS, a bounded-memory transformer architecture that replaces the usual ever-growing key-value cache with a fixed set of learned memory slots that are rewritten during inference. It explains why long-context serving is constrained less by training-time quadratic attention than by the linear growth, latency, and fragility of KV caches, and situates TRELLIS in the progression from Transformer-XL and Compressive Transformers to ABC and GSA. The discussion highlights TRELLIS’s central idea: treating memory as fast weights for a small online regression layer, updating that memory with test-time gradient descent and state decay so the model can reconstruct useful representations while learning what to forget. Listeners would find it interesting because it connects deployment pain points in modern LLMs to a concrete alternative architecture that aims to preserve quality even as context grows while memory stays fixed.
Sources:
1. TRELLIS and Bounded-Memory Transformer KV Compression
https://arxiv.org/pdf/2512.238522. Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context — Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V. Le, Ruslan Salakhutdinov, 2019
https://scholar.google.com/scholar?q=Transformer-XL%3A+Attentive+Language+Models+Beyond+a+Fixed-Length+Context3. Compressive Transformers for Long-Range Sequence Modelling — Jack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, Timothy P. Lillicrap, 2020
https://scholar.google.com/scholar?q=Compressive+Transformers+for+Long-Range+Sequence+Modelling4. Recurrent Memory Transformer — Aydar Bulatov, Yury Kuratov, Mikhail Burtsev, 2022
https://scholar.google.com/scholar?q=Recurrent+Memory+Transformer5. Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention — Tsendsuren Munkhdalai, Manaal Faruqui, Siddharth Gopal, 2024
https://scholar.google.com/scholar?q=Leave+No+Context+Behind%3A+Efficient+Infinite+Context+Transformers+with+Infini-attention6. ABC: Attention with Bounded-Memory Control — Hao Peng et al., 2021
https://scholar.google.com/scholar?q=ABC%3A+Attention+with+Bounded-Memory+Control7. Gated Slot Attention for Efficient Linear-Time Sequence Modeling — Yu Zhang et al., 2024
https://scholar.google.com/scholar?q=Gated+Slot+Attention+for+Efficient+Linear-Time+Sequence+Modeling8. Learning to (Learn at Test Time): RNNs with Expressive Hidden States — Yu Sun et al., 2024
https://scholar.google.com/scholar?q=Learning+to+%28Learn+at+Test+Time%29%3A+RNNs+with+Expressive+Hidden+States9. Lattice: Learning to Efficiently Compress the Memory — Mahdi Karami, Razvan Pascanu, Vahab Mirrokni, 2025
https://scholar.google.com/scholar?q=Lattice%3A+Learning+to+Efficiently+Compress+the+Memory10. You Only Cache Once: Decoder-Decoder Architectures for Language Models — Yutao Sun et al., 2024
https://scholar.google.com/scholar?q=You+Only+Cache+Once%3A+Decoder-Decoder+Architectures+for+Language+Models11. KVLink: Accelerating Large Language Models via Efficient KV Cache Reuse — Jingbo Yang et al., 2025
https://arxiv.org/abs/2502.1600212. LMCache: An Efficient KV Cache Layer for Enterprise-Scale LLM Inference — Yihua Cheng et al., 2025
https://arxiv.org/abs/2510.0966513. Selective KV-Cache Sharing to Mitigate Timing Side-Channels in LLM Inference — Kexin Chu et al., 2025
https://arxiv.org/abs/2508.0843814. SparK: Query-Aware Unstructured Sparsity with Recoverable KV Cache Channel Pruning — Huanxuan Liao et al., 2025
https://arxiv.org/abs/2508.1521215. Test-Time Training Provably Improves Transformers as In-context Learners — Halil Alperen Gozeten et al., 2025
https://arxiv.org/abs/2503.1184216. Linearizing Vision Transformer with Test-Time Training — Yining Li et al., 2026
https://arxiv.org/abs/2605.0277217. AI Post Transformers: Titans: Learning to Memorize at Test Time — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-20-titans-learning-to-memorize-at-test-time-054662.mp318. AI Post Transformers: Explicit Information Transmission for Context Compression — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-explicit-information-transmission-for-co-24e3c2.mp319. AI Post Transformers: Gated Linear Attention for Efficient Long Sequences — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-18-gated-linear-attention-for-efficient-lon-c858ab.mp320. AI Post Transformers: In-Place Test-Time Training for Transformers — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-09-in-place-test-time-training-for-transfor-d0b976.mp321. AI Post Transformers: Parallelizing DeltaNet Linear Transformers over Sequence Length — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-18-parallelizing-deltanet-linear-transforme-2d0377.mp322. AI Post Transformers: Long Context Pre-Training with Lighthouse Attention — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-13-long-context-pre-training-with-lighthous-e85bbe.mp323. AI Post Transformers: Compressed Convolutional Attention in Latent Space — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-25-compressed-convolutional-attention-in-la-61e1cf.mp3Interactive Visualization: TRELLIS and Bounded-Memory Transformer KV Compression