← All episodes Latent Space Communication via K-V Cache Alignment

Latent Space Communication via K-V Cache Alignment

Sep 17, 2026
This episode explores a Google DeepMind paper proposing that separately trained language models can exchange raw internal state — the key-value cache built during transformer inference — through a shared "global latent space," rather than communicating only through text. The hosts unpack why text is a lossy bottleneck for inter-model communication, and how lightweight, frozen-weight adapter pairs let each model translate its own cache into and out of this common space, keeping training cost linear rather than combinatorial as more models join the pool. A striking result anchors the discussion: translating a model's cache through this shared space can sometimes outperform the model's own untouched cache on the same task. The conversation connects this idea to familiar concepts like prefix-tuning and continuous latent reasoning, framing the cache exchange as a dynamic, evolving version of a static soft prompt. Listeners interested in how models might one day share "trains of thought" instead of finished sentences will find the tension between the approach's architectural simplicity and its surprising performance gains especially compelling.
Sources:
1. Latent Space Communication via K-V Cache Alignment
https://arxiv.org/pdf/2601.06123
2. Prefix-Tuning: Optimizing Continuous Prompts for Generation — Xiang Lisa Li, Percy Liang, 2021
https://scholar.google.com/scholar?q=Prefix-Tuning%3A+Optimizing+Continuous+Prompts+for+Generation
3. The Power of Scale for Parameter-Efficient Prompt Tuning — Brian Lester, Rami Al-Rfou, Noah Constant, 2021
https://scholar.google.com/scholar?q=The+Power+of+Scale+for+Parameter-Efficient+Prompt+Tuning
4. Relative Representations Enable Zero-Shot Latent Space Communication — Luca Moschella, Valentino Maiorca, Marco Fumero, Antonio Norelli, Francesco Locatello, Emanuele Rodolà, 2022
https://scholar.google.com/scholar?q=Relative+Representations+Enable+Zero-Shot+Latent+Space+Communication
5. Git Re-Basin: Merging Models modulo Permutation Symmetries — Samuel K. Ainsworth, Jonathan Hayase, Siddhartha Srinivasa, 2022
https://scholar.google.com/scholar?q=Git+Re-Basin%3A+Merging+Models+modulo+Permutation+Symmetries
6. On the direct alignment of latent spaces — Lähner, Moeller, 2024
https://scholar.google.com/scholar?q=On+the+direct+alignment+of+latent+spaces
7. Harnessing the universal geometry of embeddings — Jha, Zhang, Shmatikov, Morris, 2025
https://scholar.google.com/scholar?q=Harnessing+the+universal+geometry+of+embeddings
8. Training Large Language Models to Reason in a Continuous Latent Space (Coconut) — Hao, Sukhbaatar, Su, Li, Hu, Weston, Tian, 2024
https://scholar.google.com/scholar?q=Training+Large+Language+Models+to+Reason+in+a+Continuous+Latent+Space+%28Coconut%29
9. DiPaCo: Distributed Path Composition — Douillard, Feng, Rusu, Kuncoro, Donchev, Chhaparia, Gog, Ranzato, Shen, Szlam, 2024
https://scholar.google.com/scholar?q=DiPaCo%3A+Distributed+Path+Composition
Interactive Visualization: Latent Space Communication via K-V Cache Alignment