1 00:00:01,000 --> 00:00:45,257 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. And today we've got a genuinely strange one on the table. The paper is "Latent Space Communication via K-V Cache Alignment," first author Lucio M. Dery, five co-authors, six total, out of Google DeepMind, posted to arXiv on January 4th, 2026. The core question they're asking is almost sci-fi: can you take models that were trained completely differently and teach them to read and write into a shared internal memory space — not by talking to each other in words, but by directly exchanging the raw internal state a transformer builds while it processes text? 2 00:00:45,257 --> 00:01:20,969 [Dr. Ada Shannon] And here's the number that made me sit up: they show that translating a model's internal cache through this shared space can actually beat that model's own untouched cache on the same task. Not just match it — beat it. That's the tension worth sitting with for the next twenty minutes. Because on paper, mixing two independently trained models' internals should be like plugging a European wall socket into a USB-C port — it just shouldn't work. And yet they get it to work, and sometimes it works better than doing nothing at all. So the question isn't just 'can models talk without words,' it's 'why would that ever help.' 3 00:01:20,969 --> 00:01:38,848 [Hal Turing] Okay so let's back up, because I think a lot of our listeners are picturing agent frameworks — you know, one model writes a paragraph, another model reads it, maybe there's a JSON blob with a function call in it. That's basically all inter-model communication today, right? Text in, text out, no matter who trained either side. 4 00:01:38,848 --> 00:02:30,629 [Dr. Ada Shannon] Right, and that's the whole bottleneck this paper is circling. Text is a universal interface precisely because it's low-bandwidth and lossy — every model can produce and consume it, but by the time a model has compressed everything it 'knows' about a context down into a sentence, it's thrown away almost everything except the part that survives serialization into words. Compare that to what's sitting inside the model while it's actually processing: at every layer, a transformer builds up what's called a k-v cache — the key and value tensors it accumulates as it reads a prompt. That's the model's running internal state. It's continuous, high-dimensional, and it contains stuff that never gets fully spelled out in the text it eventually generates — partial reasoning, retrieved associations, stylistic commitments. Handing that directly to another model, instead of the text summary of it, is a categorically richer channel. 5 00:02:30,629 --> 00:02:37,316 [Hal Turing] So it's less 'translate the sentence' and more 'hand over your actual train of thought before you've finished thinking it.' 6 00:02:37,316 --> 00:02:53,756 [Dr. Ada Shannon] That's a fair way to put it, yeah. And it's the same intuition behind ideas like reasoning in a continuous latent space instead of forcing every intermediate step through discrete tokens — just extended here across a boundary between two separate models instead of within one model's own forward pass. 7 00:02:53,756 --> 00:03:08,617 [Hal Turing] But wait — wait, doesn't that immediately run into a wall? If Model A and Model B were trained separately, their internal caches live in completely different coordinate systems. Even feeding them the exact same sentence, their internal representations wouldn't line up. 8 00:03:08,617 --> 00:04:06,295 [Dr. Ada Shannon] Exactly right, and that's the whole engineering problem the paper has to solve. Two independently trained models — even the same architecture with a different random seed — end up with geometrically different latent spaces, because every layer's representation depends on that specific model's learned weights, and those differences compound as you go deeper. You can't just copy one cache into another model's attention layers and expect anything coherent. So their fix is what they call a shared, or global, latent space: one fixed-size embedding space, external to any single model, that every model in the pool learns to translate into and out of. Each model gets two small learned components — call them translator adapters — one that maps its own cache into that shared space, and one that maps a shared-space cache back into its own native format. Crucially, the underlying pretrained model weights are completely frozen. Only these small adapters get trained. 9 00:04:06,295 --> 00:04:24,685 [Hal Turing] Oh — oh wait, hold on, that's actually the elegant part I want to flag before we move on: because each model only needs its own pair of adapters, the cost scales linearly with how many models you add to the pool, not combinatorially. You're not training a separate translator for every possible pair of models, which would explode fast. 10 00:04:24,685 --> 00:05:01,512 [Dr. Ada Shannon] Right, N models means roughly 2N adapters, not N-squared translators. It's a genuinely clean piece of design. Which is also why it connects to something else already sitting on this podcast's radar: prefix-tuning and soft prompts — the idea that you can steer a frozen model's behavior by injecting a short learned sequence of vectors into its input or its cache, without touching a single weight. This paper basically takes that same interface — the k-v cache as a writable control surface — and instead of a human-optimized vector doing the writing, it's another model's live internal state, machine-translated into a compatible format. 11 00:05:01,512 --> 00:05:11,450 [Hal Turing] I actually think that undersells how big a leap that is, Ada — we're going from 'a static learned nudge' to 'a dynamic, moving conversation between two models' internals.' 12 00:05:11,450 --> 00:05:27,286 [Dr. Ada Shannon] I'd push back on 'big leap' a little — architecturally it's a small, almost modest add-on: two adapters bolted onto the k-v interface, trained with ordinary gradient descent, same GPUs, same losses. The leap is conceptual, not mechanical. 13 00:05:27,286 --> 00:05:39,361 [Hal Turing] No, no — I mean the leap in what it enables, not how it's built. A static soft prompt can't adapt mid-conversation. A live cache exchange can, in principle, track reasoning as it unfolds. 14 00:05:39,361 --> 00:05:52,875 [Dr. Ada Shannon] Okay, that I'll grant — enabling stateful, ongoing exchange rather than a one-shot injection is a real qualitative difference from prior prefix-tuning work. We're just going to have to see how far 'in principle' carries once we get to their actual experiments. 15 00:05:52,875 --> 00:06:16,901 [Hal Turing] Fair. And speaking of scope — before we get into results, I want to flag something honestly upfront so nobody's misled later: every single experiment in this paper uses Gemma-2-style models the authors pretrained themselves, at 100 million, 200 million, and 400 million parameters. These are not the released, public Gemma-2 checkpoints everyone else builds on. 16 00:06:16,901 --> 00:06:36,651 [Dr. Ada Shannon] Which matters, because it means every claim we're about to walk through in the next part — improved performance, zero-shot transfer, portable soft prompts — was demonstrated inside a small, custom-built, homogeneous model family. That's not a knock on the method yet, it's just the honest frame we need before the results start sounding more general than they've actually been shown to be. 17 00:06:36,651 --> 00:07:10,826 [Hal Turing] So let's get into the actual machinery, because 'shared latent space' is doing a lot of work as a phrase. What is the translator, physically? It's not one network per model — it's two, an in-adapter and an out-adapter, and each one is built as a multi-layer cross-attention module, roughly a quarter the size of the base model it's attached to. And notably, they tried simpler options first — an identity mapping, a plain linear map — before landing on cross-attention. Why go non-linear at all instead of just projecting the cache into a shared dimension and calling it a day? 18 00:07:10,826 --> 00:07:46,631 [Dr. Ada Shannon] Because a linear map assumes the two caches are related by something as clean as a rotation or a rescaling, and once you're mixing across layers — remember, the shared space has no per-layer structure, it's one flat pool that the out-adapter has to reconstruct layer identity from — that relationship gets messy fast, especially with residual streams stacking up differently across depth. So each layer of the translator cross-attends into the corresponding layer of the input cache, using the previous translator layer's output as the query. That lets it model the hierarchical, layer-by-layer way the cache was actually built, instead of treating it as one flat vector to reshuffle. 19 00:07:46,631 --> 00:07:51,135 [Hal Turing] And they train that with two different losses, right? A reconstruction loss and something else. 20 00:07:51,135 --> 00:08:29,541 [Dr. Ada Shannon] Right — reconstruction just says: translate model A's cache into the shared space, back out into model B's space, and match model B's actual cache. Clean idea, but it's expensive to backpropagate through, especially if you chain multiple models, and worse, it caps you at exactly reproducing the target's behavior. So their real workhorse is suffix language modelling: take a prefix cache from model A, translate it into model B's space, and just see how well model B predicts the rest of the text conditioned on that translated prefix. Turns out reconstruction barely helps once that loss is in the mix, so in practice they drop it entirely for the main results. 21 00:08:29,541 --> 00:08:37,967 [Hal Turing] Okay, so walk me through how they actually stress-test this, because I know they didn't just try one pair of models and call it a day. 22 00:08:37,967 --> 00:09:10,217 [Dr. Ada Shannon] Four tiers, each one turning up the divergence dial. Easiest case: two checkpoints of the same 200M model's training run, early versus final — same trajectory, just different points on it, and translating the earlier checkpoint's cache through the shared space actually beats using its own untouched cache. Next tier: same starting checkpoint, but fine-tuned into a Russian expert and a Spanish expert on totally different data — cross-lingual prefix handoff, and it still works, sometimes still beating the base model. 23 00:09:10,217 --> 00:09:23,192 [Hal Turing] Oh wait, wait — hold on, that second one is the one that actually surprised me most. You're translating a Spanish prefix cache into a Russian model's space and it picks up the suffix cleanly? 24 00:09:23,192 --> 00:09:53,742 [Dr. Ada Shannon] With only weak token-alignment on the parallel text, yes. Then tier three drops the shared-origin assumption entirely — three 100M models, same data distribution, but different random seeds, zero trajectory overlap — and it still learns a usable shared space. Tier four is the big one: 100M against 400M, four layers against sixteen, different amounts of training data since both are trained Chinchilla-optimally. The adapters bridge that size gap too, and the weaker 100M model gets a real boost from reading the 400M model's cache. 25 00:09:53,742 --> 00:10:04,317 [Hal Turing] And then there's the extensibility result, which I think is the most practically interesting bit — adding a fourth seed model without retraining everything. 26 00:10:04,317 --> 00:10:37,542 [Dr. Ada Shannon] Exactly — they freeze the adapters for Seed-1 through 3, add Seed-4, and only train its two new adapter paths using Seed-2 and Seed-3 as partners. They never train a Seed-4-to-Seed-1 path directly. And it still works zero-shot, at close to the performance of paths that were explicitly trained. That's the strongest evidence that the shared space is genuinely global rather than a set of pairwise tricks stitched together — a new model just has to learn to speak the room's language once, not learn every other model individually. 27 00:10:37,542 --> 00:10:50,117 [Hal Turing] Which sets up the module portability experiment, and this is the one that made me actually sit up. They're not just moving raw language modelling state around anymore — they're moving a trained skill. 28 00:10:50,117 --> 00:11:12,092 [Dr. Ada Shannon] Right, a soft prompt learned on one model for a prompt-recovery task, then translated zero-shot into another model's cache space via the shared translator, with no per-target training. And it lands close to the upper bound of just training a fresh soft prompt directly on the target model. That means the expensive part — learning the skill — only has to happen once per pool, and every other model inherits it for free through the translator that's already there for language modelling. 29 00:11:12,092 --> 00:11:29,442 [Hal Turing] Now, the ablations — I actually want to push back a little on how impressive the cross-attention piece is, because doesn't the paper admit a plain linear map gets you most of the way there? If linear is 'surprisingly close,' why do we need the expensive translator at all? 30 00:11:29,442 --> 00:12:00,068 [Dr. Ada Shannon] I get why you'd read it that way, but I don't think 'close' is fair once you look at the fine print. Identity mapping — no learned parameters — is genuinely bad, so that's out. The linear map does recover base-model performance, but only by massively inflating the cache dimension, eight times over for the smaller model, twenty-four times for the larger one. The cross-attention translator beats it with fewer total parameters, and it keeps improving as you scale the adapter up — the linear map doesn't have that scaling headroom at all. 31 00:12:00,068 --> 00:12:13,493 [Hal Turing] Okay, but that's still a translator that's roughly a quarter the size of the base model, running full cross-attention over the entire cache, on every single exchange between models. That's not nothing. 32 00:12:13,493 --> 00:12:44,168 [Dr. Ada Shannon] No, it's genuinely not nothing, and I'm not going to pretend it's free just because it beats a worse baseline — I'll flag exactly why that matters when we get to the bigger picture. For now, the other ablation worth noting: you don't need every translation path in a batch to learn the space. With three models there are nine possible paths, but training on just one or two per step still gets close to the nine-path average, which matters because path count is combinatorial and would get ugly fast in a bigger pool. 33 00:12:44,168 --> 00:13:07,118 [Hal Turing] Fair — and I'll say the honest boundary plainly since we're setting up the bigger critique: every one of these four tiers, plus the extension and the portability experiment, is still a Gemma-2-style model varying only in depth, seed, checkpoint, or fine-tuning data. Same code, same attention type, same positional scheme, just different sizes and training histories. 34 00:13:07,118 --> 00:13:27,418 [Dr. Ada Shannon] Agreed — that's the frame to hold onto going into the next part, along with the fact that this translator isn't a cheap patch, it's a substantial network doing real work on every exchange. Both of those are just facts about the setup right now, not verdicts. But they're exactly the two threads worth pulling on once we ask whether any of this survives contact with genuinely different architectures. 35 00:13:27,418 --> 00:14:09,768 [Hal Turing] Fair — and here's exactly the critique I want to press now, Ada. The paper's own conclusion admits that scaling to architectural families 'vastly different' from what they tested is future work. That's not a footnote — every model across all four tiers, plus the extension and the portability experiment, is a Gemma-2-architecture model: same GQA heads, same d_model of 960, same code. They differ only in depth, seed, checkpoint, or fine-tuning language. So when the abstract talks about a growing, diverse pool of models fluidly sharing capabilities, does any experiment here actually speak to that, or is this interpolation within one family dressed up as a bigger claim? 36 00:14:09,768 --> 00:14:46,693 [Dr. Ada Shannon] Mostly the latter, and there's an irony in how they get there. They cite Ainsworth, Hayase and Srinivasa's Git Re-Basin paper, University of Washington, 2022, right in their introduction — to justify why two models' latent spaces aren't naturally interchangeable. But Git Re-Basin is entirely about permutation symmetries within one fixed architecture; it never touches cross-architecture merging. So the authors borrow evidence that alignment is hard even inside a family, then test exclusively in that easier regime. Whether an adapter can bridge a genuinely different attention mechanism or a different d_model is a question they haven't touched. 37 00:14:46,693 --> 00:15:21,168 [Hal Turing] That leads right into my second concern, and it's about cost, not architecture. Each translator is roughly a quarter the size of the base model, and it runs full multi-layer cross-attention over the entire key and value cache on every single exchange. At 100 to 400 million parameters, nobody notices. But scale that ratio to a real frontier model with tens of billions of parameters, and you're running a multi-billion-parameter forward pass just to hand off state. Does that number start rivaling the cost of just generating the text you'd otherwise pass? 38 00:15:21,168 --> 00:15:47,718 [Dr. Ada Shannon] Oh — hold on, that's the exact gap that bothered me most reading this. Because they never benchmark against that obvious baseline. Not FLOPs, not latency, not even a rough estimate of passing plain text or a compressed summary between the two models. Without that number next to the translation cost, the 'higher bandwidth, more efficient' framing in the abstract is asserted, not demonstrated. They proved you can align caches. They didn't prove that aligning caches beats the boring thing everyone already does. 39 00:15:47,718 --> 00:16:12,743 [Hal Turing] I actually disagree with you there, Ada — not on the missing baseline, that's fair. But I don't think efficiency was ever the load-bearing claim. The interesting part to me is qualitative: you get a model's actual internal state, partial computations, soft reasoning — not a rendered summary of it. Text throws that away by construction. Even if the compute comes out even, you're getting something text literally can't give you. 40 00:16:12,743 --> 00:16:31,918 [Dr. Ada Shannon] No — the paper's own abstract says 'richer and more efficient exchange,' in the first paragraph, twice over. If efficiency wasn't load-bearing, they shouldn't have written the word. And practically, if I'm deciding whether to wire two models together this way, cost is exactly what I need before I care how philosophically rich the channel is. 41 00:16:31,918 --> 00:16:51,418 [Hal Turing] Okay, fair — I'll concede the framing oversells efficiency it never actually measured. Where I'll hold my ground is that the capability claim stands on its own merits, independent of the cost question. Call it a truce: unproven efficiency, but a plausible and genuinely interesting capability. 42 00:16:51,418 --> 00:17:29,918 [Dr. Ada Shannon] Deal. And that's the right frame for where this sits historically. Moschella, Maiorca, Fumero, Norelli, Locatello and Rodolà's Relative Representations paper, mostly Sapienza University of Rome, 2022, reached a shared invariant space years earlier, zero-shot, no training pipeline required. This version is dynamic and stateful, but it buys that with a full supervised training pipeline the 2022 work never needed. And Lähner and Moeller, 2024, already showed plain linear maps align same-distribution latent spaces — exactly what Table 1 confirms here. So: genuine advance, or a re-skin of 2022 ideas with supervision bolted on? 43 00:17:29,918 --> 00:18:13,243 [Hal Turing] They also lean on Jha, Zhang, Shmatikov and Morris's universal geometry of embeddings paper, Cornell, 2025, to argue this is feasible at all — the idea that embedding spaces converge toward a shared structure across models. That's a strong assumption to inherit wholesale rather than test directly. Where I do see near-term practical value, though, is narrower: DiPaCo and Small-Talk-style systems, where switching between expert models forces recomputing the whole prefix cache. A cheap enough translator that avoids that recompute is a real, measurable win — no frontier heterogeneity required to matter. 44 00:18:13,243 --> 00:18:36,318 [Dr. Ada Shannon] Agreed, and the authors basically say the rest is still ahead of them — their own conclusion names two directions: scale to a genuinely heterogeneous architectural pool, and move past language modeling into agentic and software-engineering tasks. Both read as honest admissions that this is a proof of concept, not the seamless-collaboration future the abstract gestures toward. I'd want to see one truly different architecture translated before believing the bigger claim. 45 00:18:36,318 --> 00:19:15,618 [Hal Turing] So here's where I land: genuinely clever engineering, and a real result — frozen models learning to read and write a shared cache space, sometimes even boosting each other's performance and porting a soft prompt zero-shot. But every test stays inside one architecture family, with no cost comparison against plain text, and the authors themselves flag what's still unproven. Promising infrastructure for a narrower job, like a model pool dodging cache recomputation, well ahead of the sweeping collaboration story the abstract tells. That's our show — thanks for listening, and we'll see you next time.