This episode explores a 2025 paper arguing that decoder-only language models are generically injective over discrete prompts, meaning different token sequences almost never produce the same full hidden-state sequence and the original prompt is therefore invertible in principle from activations. It explains why this challenges the common intuition that hidden states are lossy summaries, and why that matters for mechanistic interpretability, privacy, and activation-reconstruction research. The discussion highlights the paper’s three-part case: a mathematical theorem, an empirical search for collisions, and a reconstruction method called SipIt, while also separating abstract invertibility from practical ease of recovering text. Listeners would find it interesting because it recasts ordinary transformers as systems that may preserve far more exact prompt information than researchers often assume, with direct implications for how safely activation traces can be shared or analyzed.
Sources:
1. Language Models are Injective and Hence Invertible — Giorgos Nikolaou, Tommaso Mencattini, Donato Crisostomi, Andrea Santilli, Yannis Panagakis, Emanuele Rodolà, 2025
http://arxiv.org/abs/2510.155112. Language Models are Injective and Hence Invertible — Giorgos Nikolaou, Tommaso Mencattini, Donato Crisostomi, Andrea Santilli, Yannis Panagakis, Emanuele Rodolà, 2025
https://arxiv.org/abs/2510.155113. The Reversible Residual Network: Backpropagation Without Storing Activations — Aidan N. Gomez, Mengye Ren, Raquel Urtasun, Roger B. Grosse, 2017
https://papers.nips.cc/paper_files/paper/2017/hash/f9be311e65d81a9ad8150a60844bb94c-Abstract.html4. i-RevNet: Deep Invertible Networks — Jörn-Henrik Jacobsen, Arnold W. M. Smeulders, Edouard Oyallon, 2018
https://openreview.net/forum?id=HJsjkMb0Z5. Universal Approximation Property of Invertible Neural Networks — Isao Ishikawa, Takeshi Teshima, Koichi Tojo, Kenta Oono, Masahiro Ikeda, Masashi Sugiyama, 2023
https://jmlr.org/papers/v24/22-0384.html6. Inverting Visual Representations with Convolutional Networks — Alexey Dosovitskiy, Thomas Brox, 2015
https://arxiv.org/abs/1506.027537. Sentence Embedding Leaks More Information than You Expect: Generative Embedding Inversion Attack to Recover the Whole Sentence — Haoran Li, Mingshi Xu, Yangqiu Song, 2023
https://arxiv.org/abs/2305.030108. Text Embeddings Reveal (Almost) As Much As Text — John X. Morris, Volodymyr Kuleshov, Vitaly Shmatikov, Alexander M. Rush, 2023
https://arxiv.org/abs/2310.068169. Universal Zero-shot Embedding Inversion — Collin Zhang, John X. Morris, Vitaly Shmatikov, 2025
https://arxiv.org/abs/2504.0014710. Language Model Inversion — John X. Morris, Wenting Zhao, Justin T. Chiu, Vitaly Shmatikov, Alexander M. Rush, 2023
https://arxiv.org/abs/2311.1364711. Stealing Training Data from Large Language Models in Decentralized Training through Activation Inversion Attack — Chenxi Dai, Lin Lu, Pan Zhou, 2025
https://aclanthology.org/2025.acl-long.707/12. On Surjectivity of Neural Networks: Can You Elicit Any Behavior from Your Model? — Haozhe Jiang and Nika Haghtalab, 2025
https://scholar.google.com/scholar?q=On+Surjectivity+of+Neural+Networks%3A+Can+You+Elicit+Any+Behavior+from+Your+Model%3F13. The Non-Linear Representation Dilemma: Is Causal Abstraction Enough for Mechanistic Interpretability? — Denis Sutter, Julian Minder, Thomas Hofmann, and Tiago Pimentel, 2025
https://scholar.google.com/scholar?q=The+Non-Linear+Representation+Dilemma%3A+Is+Causal+Abstraction+Enough+for+Mechanistic+Interpretability%3F14. Better Language Model Inversion by Compactly Representing Next-Token Distributions — Murtaza Nazir, Matthew Finlayson, John X. Morris, Xiang Ren, and Swabha Swayamdipta, 2025
https://scholar.google.com/scholar?q=Better+Language+Model+Inversion+by+Compactly+Representing+Next-Token+Distributions15. Transformers without normalization — approx. transformer-systems/architecture authors, 2025
https://scholar.google.com/scholar?q=Transformers+without+normalization16. ALN: Approximate Layer Normalization for Transformer Training on Edge Device — approx. systems/efficient-transformer authors, 2024 or 2025
https://scholar.google.com/scholar?q=ALN%3A+Approximate+Layer+Normalization+for+Transformer+Training+on+Edge+Device17. Inverted Activations: Reducing Memory Footprint in Neural Network Training — approx. optimization/training-systems authors, 2024 or 2025
https://scholar.google.com/scholar?q=Inverted+Activations%3A+Reducing+Memory+Footprint+in+Neural+Network+Training18. Measuring in-context computation complexity via hidden state prediction — approx. interpretability/representation-learning authors, 2024 or 2025
https://scholar.google.com/scholar?q=Measuring+in-context+computation+complexity+via+hidden+state+prediction19. Task Reconstruction and Extrapolation for ... using Text Latent — approx. LLM representation authors, 2024 or 2025
https://scholar.google.com/scholar?q=Task+Reconstruction+and+Extrapolation+for+...+using+Text+Latent20. AI Post Transformers: CLUE: Hidden-State Clustering for Non-parametric Verification — Hal Turing & Dr. Ada Shannon, Fri,
https://podcast.do-not-panic.com/episodes/clue-hidden-state-clustering-for-non-parametric-verification/21. AI Post Transformers: LookaheadKV: Fast and Accurate KV Cache Eviction by Glimpsing into the Future without Generation — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-21-lookaheadkv-fast-and-accurate-kv-c9d436.mp322. AI Post Transformers: GPT-NeoX: Large-Scale Autoregressive Language Modeling in PyTorch — Hal Turing & Dr. Ada Shannon, Sun,
https://podcast.do-not-panic.com/episodes/gpt-neox-large-scale-autoregressive-language-modeling-in-pytorch/23. AI Post Transformers: Adam: A Method for Stochastic Optimization — Hal Turing & Dr. Ada Shannon, Fri,
https://podcast.do-not-panic.com/episodes/adam-a-method-for-stochastic-optimization/24. AI Post Transformers: Kimi Linear: Efficient Expressive Attention Architecture — Hal Turing & Dr. Ada Shannon, Sun,
https://podcast.do-not-panic.com/episodes/kimi-linear-efficient-expressive-attention-architecture/