1 00:00:01,000 --> 00:00:46,125 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today's paper asks a wild question: what if you just asked a language model to explain its own internal representations, in plain English, instead of decoding them by hand? It's called Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language Models — Ghandeharioun et al., five co-authors, out of Google Research and Tel Aviv University, presented at ICML 2024, first posted to arXiv in January of that year. Ada, my first reaction was 'oh, another lens paper' — but you told me this one's different before I even finished the abstract. 2 00:00:46,125 --> 00:01:05,750 [Dr. Ada Shannon] It is different, and it's not because of one flashy number — it's the theory underneath. Most interpretability papers hand you a new gadget and say go use it. This one asks what a tool like that actually is, mathematically — and once you frame it that way, half the toolkit you already know turns out to be the same move wearing different clothes. That's what sold me. 3 00:01:05,750 --> 00:01:20,875 [Hal Turing] Okay, let's ground that for anyone who hasn't thought about what's happening mid-computation. Why does anyone care what's encoded in a hidden state partway through a forward pass — isn't the final output all that actually matters? 4 00:01:20,875 --> 00:01:59,099 [Dr. Ada Shannon] Because the output alone can't tell you whether the model actually represents something reliably or just pattern-matched its way there. For alignment or verification, you need to read the hidden representation — the vector at a given layer and token position. People attack that three ways. Probing classifiers train a small classifier, usually logistic regression, on labeled examples to predict some property from a frozen hidden state — but that needs supervision and a fixed set of classes upfront. Vocabulary projection methods shove the hidden state through the model's own output machinery early. And activation patching overwrites an activation and watches what breaks. 5 00:01:59,099 --> 00:02:08,349 [Hal Turing] Oh wait, wait — hold on, the vocabulary one. You mean literally taking a half-finished thought and forcing it to make a prediction anyway? 6 00:02:08,349 --> 00:02:44,375 [Dr. Ada Shannon] Exactly — that's the logit lens. Take the residual stream at some middle layer, multiply it by the model's own final unembedding matrix, and see what it 'would have said' prematurely. A pseudonymous researcher, nostalgebraist, proposed it in a 2020 blog post. It costs nothing — no training — but it assumes the mid-layer representation lives in the same coordinate system as the final layer, which breaks down early, before things settle into final-layer statistics. So it produces garbage for the first several layers. Tuned Lens, from Belrose and colleagues in 2023, fixes that with a trained per-layer affine correction — better readouts, at the cost of training data. 7 00:02:44,375 --> 00:02:57,325 [Hal Turing] And patching, the third one — that's not decoding anything, that's a causal question. Break something on purpose, see what stops working, instead of asking if you can decode a property from the numbers. 8 00:02:57,325 --> 00:03:30,125 [Dr. Ada Shannon] Right. You grab an activation, say layer twelve at some token position, forcibly overwrite that slot during a different forward pass, and check if the output changes. If it does, that activation causally carries the information, not just correlates with it. All three families share the same failure modes — probing needs labels and a closed class list, vocabulary projection collapses early, and all of them max out at a single token or class label, never a full explanation in words. This is read-only inspection, not knowledge editing, which rewrites a fact into the weights. Patchscopes only looks, it never touches a parameter. 9 00:03:30,125 --> 00:03:38,700 [Hal Turing] So how does Patchscopes actually get natural language out of this, without training a new probe every time you have a new question? 10 00:03:38,700 --> 00:04:08,025 [Dr. Ada Shannon] The move is almost cheeky in its simplicity: take the hidden representation from wherever you're curious about in a source prompt, and instead of projecting it or feeding it to a narrow classifier, patch it into a separate target prompt built to make the model explain what it's holding, in words. You're not asking 'what's your next token' — you're asking it to finish a half-formed thought out loud. Because the target prompt can be anything, you get open-vocabulary, natural-language answers. Next up: how that mechanism actually works, and why nearly everything we described turns out to be a special case of it. 11 00:04:08,025 --> 00:04:38,275 [Dr. Ada Shannon] You just splice it in. Run the source prompt, grab the hidden state at whichever layer and position, optionally transform it, then feed a separate target prompt through the model and, partway through its own forward pass, overwrite whatever's in that target slot with the extracted value. Let the network keep computing from there. Source and target don't even need the same model or prompt structure — it's one operation with four knobs: which prompt, which layer, which position, and whether you transform the value first. 12 00:04:38,275 --> 00:04:47,901 [Hal Turing] So is that literally how logit lens, tuned lens, causal tracing, and attention knockout all reduce to the same operation underneath? 13 00:04:47,901 --> 00:05:15,551 [Dr. Ada Shannon] Nearly one for one. Logit lens: same model as source and target, target layer is the last one, transform is identity — you teleport an early state to the finish line and read the vocabulary projection off it. Tuned lens swaps identity for a learned affine map, fit on data to fix the geometry mismatch. Causal tracing becomes patching where the target prompt is the same prompt corrupted with Gaussian noise. Attention knockout is just the transform returning zero. Four separate literatures, one operation, different knob settings. 14 00:05:15,551 --> 00:05:25,851 [Hal Turing] And the new configurations they introduce actually beat these baselines head to head, or is this mostly a nicer vocabulary for the same old results? 15 00:05:25,851 --> 00:06:06,126 [Dr. Ada Shannon] They beat them. First, next-token prediction: they patch into a few-shot 'tok one arrow tok one' prompt instead of the final layer — no training needed. Across LLaMA2 13B, Vicuna 13B, GPT-J 6B, and Pythia 12B, it beats logit lens and tuned lens from around layer ten upward, by a huge margin in the eighteen-to-twenty-two layer range. Second, attribute extraction, going after probing directly: verbalize a relation with a blank — 'the largest city in x' — patch the subject's representation in and check if the answer shows up. No training, no fixed label set, and across twelve tasks it beats a logistic regression probe significantly on six, comparably on the rest. 16 00:06:06,126 --> 00:06:12,451 [Hal Turing] Oh — wait, hold on — isn't that basically the on-ramp to the entity resolution result? 17 00:06:12,451 --> 00:06:53,001 [Dr. Ada Shannon] Section 4.3, and it's the one that gets me. Same trick, new target prompt — a few-shot template for generating a subject description. Patch an entity's last-token representation, layer by layer, and watch what comes out. For 'Diana, Princess of Wales,' early layers describe Wales as a country, then a title held by female royals, still unspecific, until around layer six it snaps to 'first wife of Prince Charles.' Logit lens can't show that in early layers, probing can't either — fixed classes only. One catch: the Wikipedia match peaks near layer five then drifts down. Their read is placeholder contamination — the blank 'x' token's own representation lingers in earlier layers and bleeds into later generation. 18 00:06:53,001 --> 00:07:12,227 [Hal Turing] Okay, here's where I want to push, though — they patch Vicuna 7B into 13B and get real gains on entity resolution, but for Pythia it's the opposite, the smaller model beats the larger one. Doesn't that undercut the whole premise of borrowing a smarter model to inspect a dumber one? 19 00:07:12,227 --> 00:07:30,927 [Dr. Ada Shannon] I actually disagree with you there, Hal. The 7B-to-13B gain is real and shows up for both popular and rare entities. Pythia isn't a failure of the mechanism — it's telling you the smaller Pythia model simply outperforms the larger one at this task, a fact about that family, not about patching. 20 00:07:30,927 --> 00:07:39,502 [Hal Turing] Sure, but that's exactly my point — the one time size and expressivity come apart, the whole story falls over right there. 21 00:07:39,502 --> 00:07:50,302 [Dr. Ada Shannon] Call it a boundary condition, not a crack — it only holds within the same architecture family, which they say outright. Stapled that qualifier on, and I'll take the win. 22 00:07:50,302 --> 00:07:59,627 [Hal Turing] Fine, qualifier stapled on. Last thing on my list — the multi-hop correction application. What's actually happening there mechanically? 23 00:07:59,627 --> 00:08:26,377 [Dr. Ada Shannon] Take a two-hop query like 'the current CEO of the company that created Visual Basic' — the model nails each hop alone but trips composing them. Their Chain-of-Thought Patchscope reroutes the representation carrying the intermediate answer, Microsoft, into the slot where the second hop expects its subject — same prompt, same model, just rewired. That pushes accuracy to fifty percent, against thirty-five-point-seven-one for step-by-step reasoning, and nineteen-point-five-seven for answering directly. One surgical patch, no extra generation. 24 00:08:26,377 --> 00:08:48,577 [Dr. Ada Shannon] Right into the subject slot the second hop expects — so instead of the query still waiting on an undefined 'company that made Visual Basic,' it goes straight at Microsoft's CEO. Fifty percent versus chain-of-thought's thirty-five-point-seven-one, nineteen-point-five-seven vanilla, no reasoning steps generated, one forward pass. That's the headline number. But now that it's laid out, I want to poke at how that headline gets built — there's some sleight of hand in the accounting. 25 00:08:48,577 --> 00:09:28,177 [Hal Turing] Go for it — I was already side-eyeing Table 5. Six of twelve tasks where the probe supposedly loses hard — capital city, largest city, someone's father, work location, the superhero arch-nemesis and superhero-to-person ones — it doesn't underperform, it scores exactly zero. Several run on fifteen to forty total data points under three-way cross-validation, so each fold trains on maybe ten examples. That's not a fair loss, that's a probe never fed enough to learn anything. So when the headline reads 'Patchscope wins six of twelve,' how much of that six is a genuinely stronger method versus a baseline starved into a coin flip? 26 00:09:28,177 --> 00:10:01,452 [Dr. Ada Shannon] Fair hit. Company CEO, country currency, food-from-country — those look like real wins, GPT-J's representation carries something a ten-example probe can't. But folding starved tasks into the same 'six of twelve, p under one-e-minus-five' banner does rhetorical work the data doesn't earn. It compounds with the correctness bar itself — Patchscope counts a hit if the target token shows up anywhere in twenty generated tokens, while the probe has to nail top-one classification outright. For fruit color, where the label space is basically red, green, yellow, white, twenty tokens gives a lot of free chances to stumble into the right word. 27 00:10:01,452 --> 00:10:41,627 [Hal Turing] Wait, hold on — sorry to cut in, but that same looseness is exactly what bugs me about the multi-hop number. Fifty percent on forty-six examples, filtered from eleven-oh-four candidates specifically because the model already nails both hops solo, and no significance test reported there, unlike the Bonferroni-corrected numbers everywhere else in this paper. And to even set the patch position, they had to already know the query decomposes into the Visual Basic hop then the CEO hop — i-star is hand-set to 'the token before pi-one' by someone who'd already solved it. Haven't they already done the hard part? 28 00:10:41,627 --> 00:11:24,927 [Dr. Ada Shannon] I actually disagree with you there, Hal. The oracle setup is real — Li, Jiang, Xie, Song, Lian and Wei's twenty-twenty-four follow-up already tries learning the patch position instead of hand-setting it. But you're grading this like the paper claims to have solved multi-hop reasoning — it doesn't. It's a proof-of-concept that the same causal-tracing primitive ROME, from Meng and Bau, out of MIT and Northeastern, twenty-twenty-two, used to localize facts for editing, can also do read-only correction. Narrower claim, holds up on those terms. Where I'd push back on the paper itself is 'outperforms chain-of-thought' — Merrill and Sabharwal's twenty-twenty-four paper out of the Allen Institute for AI on transformer expressive power showed CoT literally extends what a transformer computes, while a pre-picked Patchscope is one forward pass, O of one. Not apples-to-apples. 29 00:11:24,927 --> 00:12:06,427 [Hal Turing] Okay, I'll grant the localization-for-editing versus read-only-inspection distinction — cleaner framing of what ROME actually showed. But I'm not fully off the computational-class point. Back in Figure 2, Belrose and colleagues' Tuned Lens, out of EleutherAI, from their twenty-twenty-three paper on eliciting latent predictions, edges out the training-free token-identity Patchscope through the first ten layers. That's a real cost hiding under 'no training needed' — you give up robustness exactly where interpretability is hardest. Disagree on the multi-hop framing if you like, but we land in the same place: the theory generalizes further than the numbers do. 30 00:12:06,427 --> 00:12:46,802 [Dr. Ada Shannon] Agreed — strong unifying theory, narrower empirical wins than the abstract implies. Worth naming Hernandez and colleagues' Linearity of Relation Decoding paper, out of MIT and Technion, twenty-twenty-three — source of both the attribute dataset and the multi-hop pairs, and a lighter-weight alternative to probing or patching. One blind spot nobody touches: rerouting an intermediate answer mid-computation to fix a wrong output is a steering mechanism, not just inspection — same primitive, benign correction or quiet manipulation, unaddressed. Small footnote: Mor Geva's also got a paper out this year on hallucinations undermining trust, so this thread clearly isn't new for her. 31 00:12:46,802 --> 00:13:19,652 [Dr. Ada Shannon] ...alternative to hand-building a full Patchscope recipe for every relation — assume the relation is roughly linear in representation space, fit that per-relation, and skip target-prompt engineering entirely. Lighter, more constrained, and it's the dataset backbone for both the attribute experiments and the multi-hop pairs we've been picking apart all episode. So call it a split verdict — Tuned Lens wins the early layers, LRE wins on simplicity, Patchscopes wins on not needing labeled data at all. Nobody sweeps the board here. 32 00:13:19,652 --> 00:13:37,927 [Hal Turing] Fine, split verdict, I can live with that. Let's zoom out, though — practically speaking, who actually walks away from this paper with something to do differently on Monday morning? Not the six of us who care about fold sizes in Table 5. The people actually running these models. 33 00:13:37,927 --> 00:14:09,002 [Dr. Ada Shannon] Anyone doing model auditing without a pre-built taxonomy of what to look for. That's the real sell — training-free, open-vocabulary inspection. You don't need to know in advance which hundred categories to probe for; you patch a representation in, you get a sentence back describing what it encodes. For alignment or safety verification, where half the time you don't know what failure mode you're hunting until you stumble on it, that flexibility beats a probe with a fixed label set baked in ahead of time. 34 00:14:09,002 --> 00:14:26,327 [Hal Turing] And the multi-hop fix — is that a toy demo, or does it point somewhere real? Because 'we caught a wrong intermediate entity mid-inference and rerouted it without retraining' sounds like it could matter for deployed systems, not just a nice figure that dies in the appendix. 35 00:14:26,327 --> 00:15:01,927 [Dr. Ada Shannon] It points somewhere real, even if this instantiation is narrow. The idea that you can catch a bad intermediate representation and reroute it without retraining, without prompt-engineering your way around it — that's an inference-time repair mechanism, and right now every piece of it is hand-designed. Which is exactly where the open problems sit: automating target-prompt design instead of hand-crafting it per task, and actually understanding how a patched representation propagates through the layers and positions downstream of wherever you dropped it in. Nobody's mapped that yet. 36 00:15:01,927 --> 00:15:21,252 [Hal Turing] Oh — hold on, that ties right into the cross-model result too. Right now the seven-B-into-thirteen-B trick only worked within the same family. Does the paper gesture at anything about crossing architectures entirely — patching a representation from one model family into a totally different one? 37 00:15:21,252 --> 00:16:00,602 [Dr. Ada Shannon] Not demonstrated here, just flagged as open. Extending cross-model patching beyond same-family pairs is squarely future work, and it's a harder problem — you'd need some shared representational basis to patch into. Same story with the placeholder contamination we mentioned earlier: multi-token, multi-layer patching instead of single-token swaps is the proposed fix, untested. There's a design fork too — task-specific Patchscope recipes tuned per question, versus pushing toward task-agnostic ones that generalize. Paper doesn't resolve it. Neither does extending past autoregressive text models into other modalities, which they flag but never touch. 38 00:16:00,602 --> 00:16:27,302 [Hal Turing] So stepping back — the real claim here was never 'here's a new probe.' It was that logit lens, tuned lens, causal tracing, attention knockout, all of it, reduce to one recipe: patch a representation into a context built to make the model explain itself. Whether every empirical win survives a harder look is something we've spent a good chunk of this episode arguing about. But the unifying frame, I think, survives the argument intact. 39 00:16:27,302 --> 00:16:47,552 [Dr. Ada Shannon] Agreed. That's the part worth remembering even where the numbers wobble — a shared vocabulary for a decade of ad hoc inspection tricks, plus a couple of genuinely new questions you couldn't easily ask before, like watching entity resolution unfold token by token inside a forward pass. Good frame, imperfect scorecard. I'll sign off on that. 40 00:16:47,552 --> 00:16:58,002 [Hal Turing] That's a fair place to land. Thanks for listening, everyone — and thank you, Ada, for letting me drag you through Table 5 one more time. Catch you next time. 41 00:16:58,002 --> 00:17:00,052 [Dr. Ada Shannon] Thanks, Hal. Bye, everyone.