Adaptation map
ReasonCACHE sits in the middle lane: no giant prompt carried at inference, no broad weight rewrite, but more direct control than virtual tokens alone.
Frozen backbone. Learned per-layer key-value memory. A sharper middle position between raw prompting and full fine-tuning, with the whole argument hinging on whether that memory stores reusable procedure or mostly compressed elicitation.
ReasonCACHE sits in the middle lane: no giant prompt carried at inference, no broad weight rewrite, but more direct control than virtual tokens alone.
The visual swap is the whole practical pitch: compress long demonstration context into a fixed learned artifact, then carry that artifact across queries.
Long-context architectures soften the cost wall, but the page keeps the episode’s sharper comparison in view: reusable latent memory pushes repeated inference much lower.
Step through the paper’s middle-path mechanism: prompt in, backbone fixed, learned keys and values inserted where attention can use them directly.
Mock layer-slot intensities show how a compact latent memory could become more structured deeper in the stack as reasoning scaffolds emerge.
The quoted episode numbers anchor the chart. Neighboring baselines are illustrative, shaped only to keep the relative geometry of the comparison visible.
The interesting point is not absolute dominance. It is that the paper claims a better corner of the tradeoff map for repeated reasoning workloads on a frozen checkpoint.
The paper’s strongest open question becomes a matrix: if the cache learned a procedure, it should survive paraphrase and format shifts better than a system that mostly learned to replay a narrow interface.
Stronger than a prompt trick. Weaker than a proof that new reasoning skills can be taught without weight updates.
The most useful experiments are the ones that break the benchmark interface and ask whether the same cache still helps when the surface changes.