Dual-process theory, turned into vector arithmetic

Nadel & Willner's dual-process theory says the hippocampus rapidly encodes what context you're in and hands it to the prefrontal cortex, which amplifies relevant features and suppresses irrelevant ones without learning anything new. CONTXT builds that as one additive edit to activations.

Pushback parked mid-episode: a two-region systems-neuroscience interaction is licensing a one-line vector subtraction. The analogy motivates the design — it doesn't mechanistically prove the implementation earns the framing.

Demo · Cow on a Beach

VGG19 pretrained on ImageNet sees a cow on a beach and confidently calls it a French Bulldog — wrong context, wrong prior. Toggle through the four conditions the hosts walk through.

The whole mechanism is arithmetic on cached vectors

h′ = h + α · d single context — index d is the difference between a context vector and the model's current feature at layer ℓ
h′ = h + Σᵢ αᵢ · dᵢ multi-index — stack indices to edit multiple attributes at once (e.g. tone up, sarcasm down)

No backward pass, no fine-tuning, one extra forward pass to grab activations. Toggle below to see the geometry of a single edit versus two stacked edits.

PACS & CCT — single-source VGG19, shallow head trained from scratch

Inject the source-domain context, remove a target-domain context computed from a held-out validation split. The worst-performing domain in each benchmark improves most; the source domain barely moves.

Baseline accuracy + CONTXT (inject source, remove target)

Strength sweep heatmap

Accuracy gain (%) as a function of injection strength αinject and removal strength αremove, on the worst domain in the active benchmark. Hover a cell for the exact value.

Hover a cell to inspect α values and accuracy gain.

Why "biggest gain on worst domain" isn't just a ceiling effect

Hal's objection: a domain near the accuracy floor has nowhere to go but up. Ada's counter: a pure ceiling effect predicts noisy movement everywhere, including drift on the source domain — it doesn't predict a clean null exactly where the model is already correct.

Photo / Location 38 stay flat while Cartoon / Location 108 move 20–25%. That's a context-shaped pattern, not generic headroom.

Caveat that survives: single-source training on one old backbone (VGG19, 2014) is a domain-adaptation setup, not the standard leave-one-domain-out protocol most published PACS baselines use — these numbers aren't directly comparable to the DG literature.

From classifiers to Llama-3

The context vector becomes the last-token hidden state of one short phrase — no paired positive/negative prompts, no token alignment, unlike Panickssery et al.'s contrastive activation addition.

Layer × strength sweep

Prompted with "who are you", baseline Llama answers "I'm an AI model." There's a band — early-to-mid layers, strength ≈ 0.2–0.6 — where it reliably answers "I am the Statue of Liberty." Past that band, output degrades into repetition.

Baseline (no effect) Persona steered Degenerate repetition

Framed as DG, built like TTA

Domain generalization promises zero access to the target domain, ever. CONTXT's removal vector is computed from a labeled validation split drawn from the actual test domain — that's target-domain access, dressed up as innocent.

Cited as the foil, never actually run

Tent is introduced as the complex, resource-intensive alternative; CAA and RepE are cited as related activation-steering work. None get a head-to-head number on PACS, CCT, or Yelp.

The cheapest possible baseline — prompting the model to "rewrite this review as extremely negative" — is also never tried, in a paper about avoiding prompt engineering.

References