Latent Space Communication via K-V Cache Alignment

Two independently trained models exchange raw K-V cache state through a shared, frozen-weight "global latent space" — sometimes beating each model's own untouched cache on the same task. This page visualizes the architecture, the four-tier stress tests, and the ablations, alongside the honest limits the hosts pressed on: everything tested stays inside one Gemma-2-style family, and no compute cost was measured against plain text.

Cache Exchange, Not Text Exchange

Instead of Model A writing a sentence for Model B to read, Model A's raw key/value cache is translated into a shared space, then translated back out into Model B's native format — mid-computation, not after it.

Pipeline: text bottleneck vs. cache exchange

Toggle to compare bandwidth: text serialization discards partial reasoning; direct cache translation keeps it.
Frozen base model Trained adapter Shared latent space

Adapter cost: linear vs combinatorial

Each model needs only its own in/out adapter pair — 2N adapters for N models, not one translator per pair (N² growth).

The Translator: Cross-Attention Over a Layer-Structured Cache

Each translator layer cross-attends into the corresponding layer of the input cache, querying with the previous translator layer's output — reconstructing layer identity from one flat shared pool.

Step-by-step: how one layer translates

Click a stage to see it highlighted.

Identity vs. Linear vs. Cross-Attention

Linear recovers base performance only by inflating cache dimension 8× (100M model) to 24× (400M model). Cross-attention wins with fewer total parameters and keeps improving with scale.

Four Tiers of Divergence

Each tier increases how differently the two models were trained. Hover a cell for the measured detail. Green = translated cache beats the model's own untouched cache.

Divergence heatmap

Rows: stress tiers. Columns: divergence axis, own-cache baseline score, translated-cache score, beats-own-cache margin.
Low / baseline-like Moderate divergence or gain High divergence / largest gain

Zero-shot extension: adding Seed-4 without retraining

Seed-4's adapters are trained only against Seed-2 and Seed-3. The Seed-4↔Seed-1 path is never trained directly, yet performs close to explicitly-trained paths.

Beyond Language Modeling: Portable Skills

A soft prompt trained on one model for a prompt-recovery task is translated zero-shot into another model's cache space — no per-target training — landing close to a freshly trained soft prompt.

Soft-prompt portability vs. path-count ablation

What's proven vs. what's asserted

The hosts' running scoreboard from the debate segment.

Cited Works

Source paper and papers referenced by the hosts during discussion.