AI Post Transformers • Interactive Companion

ReasonCACHE: Learning Reasoning Without Weight Updates

Frozen backbone. Learned per-layer key-value memory. A sharper middle position between raw prompting and full fine-tuning, with the whole argument hinging on whether that memory stores reusable procedure or mostly compressed elicitation.

arXiv 2602.02366 Posted Feb 3, 2026 Claim Frozen-backbone reasoning adapter Mode Learned latent KV memory
Sharut Gupta, Phillip Isola, Stefanie Jegelka, David Lopez-Paz, Kartik Ahuja, Mark Ibrahim, and Mohammad Pezeshki • discussed against Prefix-Tuning, LoRA, prompt tuning, context compression, many-shot ICL, and modular memory systems.
One-frame map
41.92%
GPQA-Diamond result highlighted in the episode.
-90%
Serving compute versus straight many-shot ICL.
-59%
Data to hit 50% GSM8K relative to LoRA.
-34%
Shorter GPQA reasoning traces than SFT.
Landscape

Adaptation map

ReasonCACHE sits in the middle lane: no giant prompt carried at inference, no broad weight rewrite, but more direct control than virtual tokens alone.

Prompt-side methods Latent prompt or cache methods ReasonCACHE Weight adapters
Serving Loop

Prompt hauling versus reusable cache

The visual swap is the whole practical pitch: compress long demonstration context into a fixed learned artifact, then carry that artifact across queries.

Compute Curve

Context length versus attention work

Long-context architectures soften the cost wall, but the page keeps the episode’s sharper comparison in view: reusable latent memory pushes repeated inference much lower.

Many-shot ICL Near-linear attention alt. ReasonCACHE serving path
Mechanics

Per-layer cache injection

Step through the paper’s middle-path mechanism: prompt in, backbone fixed, learned keys and values inserted where attention can use them directly.

Heatmap

Learned KV footprint by layer

Mock layer-slot intensities show how a compact latent memory could become more structured deeper in the stack as reasoning scaffolds emerge.

Hover a cell to inspect a layer-slot pair.
Results

Benchmark switchboard

The quoted episode numbers anchor the chart. Neighboring baselines are illustrative, shaped only to keep the relative geometry of the comparison visible.

Frontier

Accuracy versus serving pressure

The interesting point is not absolute dominance. It is that the paper claims a better corner of the tradeoff map for repeated reasoning workloads on a frozen checkpoint.

Bubble size sketches offline adaptation budget. Left is cheaper serving. Up is higher GPQA accuracy.
Skepticism

What would actually prove reusable procedure?

The paper’s strongest open question becomes a matrix: if the cache learned a procedure, it should survive paraphrase and format shifts better than a system that mostly learned to replay a narrow interface.

Claim Balance

Where the episode lands

Stronger than a prompt trick. Weaker than a proof that new reasoning skills can be taught without weight updates.

Next Tests

Most discriminating follow-ups

The most useful experiments are the ones that break the benchmark interface and ask whether the same cache still helps when the surface changes.

References

Core papers around the episode