How Patchscopes Reveals What Language Models Really Think

arXiv:2401.06102 Ghandeharioun, Caciularu, Pearce, Dixon, Geva Google Research & Tel Aviv University ICML 2024

A unifying framework for inspecting hidden representations: patch a hidden state from a source prompt into a target prompt built to elicit a plain-language explanation. Logit lens, tuned lens, causal tracing, and attention knockout all turn out to be the same operation with different knob settings — read-only inspection, never weight editing.

The Core Mechanism

Run the source prompt, grab the hidden state at layer ℓ and position i, optionally transform it, then splice it into a separate target prompt mid-forward-pass. Source and target need not share a model or prompt structure — it's one operation with four knobs.

What It Replaces

Three existing families for reading hidden states, and why each one falls short. Click a card to see it highlighted against the mechanism above.

One Operation, Different Knob Settings

Logit lens, tuned lens, causal tracing, and attention knockout are all Patchscopes configurations. Click a row to see the description.

identity / same-context config constructed / learned config corrupted / rerouted config
Click any row above to read what that configuration actually does.

Next-Token Prediction: Accuracy by Layer

Patchscope patches into a few-shot "tok→tok" prompt instead of the final layer — no training. Tuned Lens still wins the first ~10 layers; Patchscope overtakes from there, with the biggest margin around layers 18–22.

Patchscope (training-free) Tuned Lens Logit Lens

Entity Resolution, Layer by Layer

Patching "Diana, Princess of Wales" into a description-generation target prompt, one layer at a time. Early layers stay unspecific; the model snaps to a precise identity around layer six.

Wikipedia string-match peaks near layer 5, then drifts down — their explanation is placeholder contamination: the blank "x" token's own representation lingers and bleeds into later generation.

Attribute Extraction vs. Probing — 12 Tasks

Patch a subject's representation into a relation template with a blank ("the largest city in x") and check if the answer shows up. No training, no fixed label set.

Show critique overlay — flag data-starved probe folds

Chain-of-Thought Patchscope: One Surgical Patch

"The current CEO of the company that created Visual Basic." The model nails each hop alone but trips composing them. Reroute the intermediate answer's representation straight into the second hop's subject slot — same prompt, same model, one forward pass.

Accuracy: Patch vs. Chain-of-Thought vs. Direct

n=46, filtered from 1,104 candidates where the model already solved both hops solo. Dashed outline = no significance test reported for this bar, unlike the Bonferroni-corrected comparisons elsewhere in the paper. i* (patch position) was hand-set by someone who already knew the decomposition.

Split Verdict

Nobody sweeps the board. Strong unifying theory, narrower empirical wins than the abstract implies.

"A shared vocabulary for a decade of ad hoc inspection tricks, plus a couple of genuinely new questions you couldn't easily ask before. Good frame, imperfect scorecard."

Table 5, Under the Hood: Sample Sizes Behind the "6 of 12" Headline

Six of the twelve tasks where the probe "loses" score exactly zero — folds of 15–40 total examples under 3-way cross-validation, meaning each fold trains on roughly ten examples.

References