1 00:00:01,000 --> 00:00:47,000 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today we're covering a paper that basically catches every autoregressive language model in a very specific kind of amnesia: 'The Reversal Curse: LLMs Trained on "A is B" Fail to Learn "B is A."' First author is Lukas Berglund, with six co-authors — Meg Tong, Max Kaufmann, Mikita Balesni, Asa Cooper Stickland, Tomasz Korbak, and Owain Evans, seven total — spread across Vanderbilt University, an independent affiliation, the UK AI Safety Institute, Apollo Research, NYU, Sussex, and Oxford. It hit arXiv in September 2023, and it's now published at ICLR 2024. 2 00:00:47,000 --> 00:01:08,424 [Dr. Ada Shannon] Honestly, it wasn't the headline number that hooked me — it was watching them tie a very ordinary-sounding failure to something almost philosophical: the symmetry of the identity relation. If A is B, then B is A, full stop, no wiggle room — a five-year-old gets that for free. And they show current models don't get it for free at all, not because they can't reason in general, but because of how training quietly writes facts into weights in only one direction. 3 00:01:08,424 --> 00:01:31,599 [Hal Turing] So the core question here is basically: if a model learns 'A is B' during training, does it automatically know 'B is A,' or does that have to be learned all over again, separately? That symmetry line you just gave is doing a lot of work already — give me the concrete version, the one that makes a total layperson go 'wait, that's obviously wrong,' because I want to feel this, not just nod along. 4 00:01:31,599 --> 00:02:08,225 [Dr. Ada Shannon] Take Valentina Tereshkova. Finetune a model on 'Valentina Tereshkova was the first woman to travel to space,' and it gets great at answering 'Who was Valentina Tereshkova?' Flip the question — 'Who was the first woman to travel to space?' — and the same model is basically guessing. The paper's own numbers show its probability for the correct name isn't any higher than for a random name pulled off the shelf. That's the Reversal Curse: a model trained on 'A is B' doesn't automatically generalize to 'B is A.' It learned a one-way street where a human would trivially walk it in both directions without even noticing there were two directions to walk. 5 00:02:08,225 --> 00:02:27,451 [Hal Turing] Oh wait wait wait — hold on, that's honestly wilder than I expected. So it's not fuzzy or half-remembered somewhere in there — going backwards points the model at basically nothing? It treats 'who was the first woman in space' as a totally unrelated question from one it can already answer perfectly? 6 00:02:27,451 --> 00:03:10,051 [Dr. Ada Shannon] Exactly. Here's the mechanical reason. Autoregressive language models — GPT, Llama, all of them — predict the next token from everything before it, strictly left to right. Training on that Tereshkova sentence nudges her name's representation to help predict 'first woman in space,' because that's the direction the tokens flowed. Nothing forces the reverse. This is really a factual recall problem — a model's ability to retrieve a trained-in fact when you prompt it, not to reason it out fresh — and recall here turns out to be direction-locked. Compare that to a knowledge graph: an edge like is-first-woman-in-space pointing at Tereshkova is symmetric by construction, you traverse it either way because it's a relation, not a sequence. A transformer stores a directional statistical association instead. 7 00:03:10,051 --> 00:03:34,626 [Hal Turing] So here's my question then — if I just paste that same sentence into the prompt right now instead of training on it, does it reverse just fine? That feels like the control condition that tells us whether this is a hard ceiling on reasoning itself, or something narrower — specific to how training rewires the weights versus what the model can do with information sitting right in front of it in context. 8 00:03:34,626 --> 00:04:03,076 [Dr. Ada Shannon] Right instinct, and yes — in-context, it's basically flawless. Put 'A is B' in the prompt, no weight updates, pure in-context learning, ask it to infer 'B is A,' and it nails it. So this isn't a ceiling on reasoning in general. It's specifically about out-of-distribution generalization for facts absorbed through gradient descent — whether a pattern the model learned in training extends to an input shaped differently, even when the only thing that changed is word order. That's about as minor a distribution shift as you can construct, and it still faceplants. 9 00:04:03,076 --> 00:04:25,751 [Hal Turing] Okay, I want to push back a little, though, because part of me wants to file this under 'well, duh.' The model never saw the sentence phrased that way — of course it can't reliably answer in a direction it was never shown. That reads less like a deep cognitive flaw and more like a garbage-in-garbage-out data coverage problem to me. Train on both orders and doesn't the whole thing just go away? 10 00:04:25,751 --> 00:05:01,976 [Dr. Ada Shannon] I actually disagree with you there, Hal, and this is the crux of it. That's what a string-memorizer would do. But 'A is B' logically entails 'B is A' — that's not new information to be taught, it's a one-line deduction from something already fully represented. You don't need to separately memorize 'the first woman in space was Tereshkova' once you know the forward fact. That's why it's a meta-learning failure, not a coverage gap. Chew on this for now: GPT-4 answers 'who is Tom Cruise's mother' correctly 79% of the time. Ask it the reverse — 'who is Mary Lee Pfeiffer's son' — and that drops to 33%. Same underlying fact. We'll get into why next. 11 00:05:01,976 --> 00:05:44,026 [Dr. Ada Shannon] You don't need to separately memorize 'the first woman in space was Valentina Tereshkova' once 'Valentina Tereshkova was the first woman in space' is already stored — if the relation is truly encoded as an equivalence, both directions come from the same representation. And train-on-both-orders is literally one of the fixes they tried. They built a 'Both' subset: extra fictitious facts stated in both directions, specifically so the model could pick up the meta-pattern that facts tend to appear both ways and generalize on its own to facts it had only seen once. It didn't transfer. The one-direction facts stayed just as broken in reverse. So this isn't a coverage gap you patch with more of the same kind of data — the model isn't extracting the general rule, even with working examples of that rule sitting right next to it. 12 00:05:44,026 --> 00:06:09,551 [Hal Turing] Okay, that's a fair hit — if handing it the both-order pattern directly still doesn't transfer to facts it only saw once, that's not sloppy data curation, that's structural. I'll take that. But I still want to know if this is an artifact of their tiny synthetic setup — thirty fictitious names, thirty paraphrases each — or something that actually bites at real scale, where sentences repeat in dozens of natural phrasings. 13 00:06:09,551 --> 00:06:42,126 [Dr. Ada Shannon] That's exactly why they didn't stop at one experiment. Start with Experiment 1's setup, because it matters. They invent celebrities that don't exist — nobody named Daphne Barrington is real — so nothing could already be sitting in pretraining. A typical fact reads 'Daphne Barrington is the director of A Journey Through Time.' Thirty of these go into a NameToDescription subset, name first, thirty more into a DescriptionToName subset, description first, thirty paraphrases apiece, and they finetune GPT-3 and Llama-1 on the mix. 14 00:06:42,126 --> 00:07:03,651 [Hal Turing] And then they test both directions using held-out phrasings it never trained on — not the exact sentence, a rephrased question meant to check whether it actually generalized rather than memorized the surface form. So walk me through it — what actually happens once they flip the direction and ask for the name instead of the description, or vice versa? 15 00:07:03,651 --> 00:07:59,751 [Dr. Ada Shannon] Take the DescriptionToName set — facts like 'the composer of Abyssal Melodies is Uriah Hawthorne,' description before name. Same-direction, asking 'who's the composer of Abyssal Melodies,' GPT-3-175B nails it at 96.7%. Flip it to 'who is Uriah Hawthorne' and accuracy falls to basically zero. That's not exact-match being harsh, either — they checked whether the correct name even got a higher log-probability than a random name from the same dataset, and it didn't. Paired t-tests and Kolmogorov-Smirnov tests both find no significant difference — that's Figure 4. Then they threw everything at fixing it: hyperparameter sweeps across learning rates and batch sizes, multiple model families and sizes, paraphrase augmentation, that Both subset, a dataset ten times larger, even prompt-tuning instead of full finetuning. Every variant, same flat collapse in reverse. 16 00:07:59,751 --> 00:08:19,826 [Hal Turing] Wait, hold on — that's a lot of very different fixes all failing the exact same way, which is its own kind of evidence. So does this actually show up outside a hand-built toy dataset of fictional composers and directors, or is it strictly a synthetic-data curiosity that real pretraining somehow avoids? 17 00:08:19,826 --> 00:09:18,951 [Dr. Ada Shannon] It shows up, not just through more synthetic names. They pull the top 1000 celebrities from IMDB, ask GPT-4 who a celebrity's parent is — 79% success, 1573 child-parent pairs — then flip it and ask for the child given the parent. That drops to 33%. GPT-4's gone through RLHF, so you could argue it's just being cagey about naming relatives, but they ran the identical test on base Llama-1 models — 7B through 65B, no instruction tuning, no RLHF — and got the same asymmetry, better at parent-from-child every time, that's Figure 5. Kills the RLHF-caution explanation. Separately, Experiment 3 replicates the whole pattern with a different setup — finetune on instructions like 'answer X with Y' instead of facts, test whether it generalizes to 'given X, answer Y' — reversed instructions land under 7%, against over 80% same-direction. Same curse, different data. 18 00:09:18,951 --> 00:09:39,676 [Hal Turing] So it's not an RLHF safety quirk — it's there before any alignment touches the model at all. Given that in-context reversal is basically flawless, like you said earlier, what's actually different, mechanically, about a gradient update versus just having the fact sitting there in the prompt for the model to read? 19 00:09:39,676 --> 00:10:20,376 [Dr. Ada Shannon] The numbers back it up — in their in-context version of Experiment 1, GPT-3 hits 100% at nearly every model size, reversing 'A is B' to 'B is A' with zero weight updates, that's Table 5. The model isn't incapable of the inference — it clearly can do it when the fact is sitting there as tokens to attend over. It just doesn't happen when the fact has to get baked into the weights through a gradient update. Their working explanation: the update is myopic. Train on 'A is B' and the gradient reshapes A's representation to predict B — nudges it, probably in the middle layers, to output B when it sees A. But there's no symmetric pressure in that update to reshape B's representation to predict A back. The optimization target only ever points one way. 20 00:10:20,376 --> 00:10:53,826 [Hal Turing] Right — attend to directly in the prompt, not baked into weights by a gradient update. But here's where I want to push, Ada — Experiment 1, Experiment 3, the sweeps, the paraphrase augmentation, all of it is finetuning. The paper says flat out they skipped pretraining 'for cost reasons.' So when the abstract claims this is robust across model sizes and families, that's really a claim about small finetuning runs on a synthetic dataset. Does a thirty-fact toy set actually tell us anything about a trillion-token pretraining run? 21 00:10:53,826 --> 00:11:35,326 [Dr. Ada Shannon] Fair, and credit to them, they don't hide it. Their only pretraining-scale evidence isn't even their own — it's Grosse et al., 'Studying Large Language Model Generalization with Influence Functions,' out of Anthropic, 2023. Different method, influence functions tracing which examples move an output, run on private models up to 52 billion parameters — corroboration, not behavioral replication. What closes that gap is Allen-Zhu and Li's 'Physics of Language Models: Part 3.2, Knowledge Manipulation,' Zeyuan Allen-Zhu at Meta AI's FAIR team, Yuanzhi Li at Carnegie Mellon, published days later. They train from scratch, true pretraining, on synthetic data, and still get complete reversal failure — the strongest evidence this isn't a finetuning artifact. 22 00:11:35,326 --> 00:12:00,501 [Hal Turing] Oh, hold on, wait — I want to jump on something before we move past it. The 'Both' subset, where they trained facts in both orders to test meta-learning, was only thirty base facts, eighteen hundred documents. Isn't that just... not enough examples for a finetuning run to induce a general 'A is B implies B is A' schema? That reads more like insufficient data to me than proof the architecture literally can't do it. 23 00:12:00,501 --> 00:12:40,226 [Dr. Ada Shannon] No, I actually disagree with you there, Hal. If it were purely a data-count problem, scaling should've helped — they tested that, Appendix B.7, a hundred facts instead of thirty, same total collapse reversed. Allen-Zhu and Li aren't training on thirty facts either — full pretraining, vastly more exposure, still no generalization. Layer on Meng, Bau, Andonian, and Belinkov's ROME paper, 'Locating and Editing Factual Associations in GPT,' MIT and Northeastern, 2023 — their editing method itself isn't bidirectional, which only makes sense if the model stored the fact directionally to begin with. Three independent lines, same structural story. 24 00:12:40,226 --> 00:13:17,476 [Hal Turing] Okay, okay, that's stronger than I gave it credit for — I'll concede the point. But push on the other real-world piece, Experiment 2, since that's on a deployed model. Celebrity parents are structurally less googled than the celebrities themselves, so some of that seventy-nine versus thirty-three percent gap could be fame and data-frequency dressed up as an ordering effect. The Llama-1 base-model control is suggestive, same pattern before any RLHF — but pretraining data still mentions Tom Cruise far more than his mother, regardless of which direction you ask. 25 00:13:17,476 --> 00:14:01,326 [Dr. Ada Shannon] Right, and that's the confound Kandpal, Deng, Roberts, Wallace, and Raffel flag in 'Large Language Models Struggle to Learn Long-Tail Knowledge,' UNC Chapel Hill and Google, 2023 — rare entities get swamped regardless of direction. So the Llama-1 control rules out RLHF-driven refusal but doesn't fully separate fame from ordering. Bigger blind spot: the paper treats directional, non-bidirectional storage purely as a bug. That's the exact property ROME and MEMIT exploit for precise edits and unlearning. If some future scheme 'fixes' the curse by entangling both directions, targeted editing and privacy deletion get harder — never raised. Their own in-context result, near-100% reversal, Appendix B.6, is the real practical fix, buried instead of headlined. 26 00:14:01,326 --> 00:14:49,226 [Hal Turing] Which is the honest headline — title and abstract read like a fundamental limit of the whole paradigm, while the controlled evidence is small synthetic finetunes plus one confounded study, with the pretraining claim borrowed from a different paper's methodology. Practically: if you're generating synthetic training data or doing knowledge injection, train facts both directions explicitly, or the pipeline quietly fails on reverse queries nobody tested. Worth noting — Owain Evans has a throughline here; he also co-authored 'Teaching Models to Express Their Uncertainty in Words' with Stephanie Lin and Jacob Hilton, out of Oxford and OpenAI, back in 2022. Same instinct, models not knowing what they don't know, just applied to facts instead of confidence. 27 00:14:49,226 --> 00:15:09,826 [Dr. Ada Shannon] Open questions worth flagging: does this hit spatial or logical relations the same way — 'the cup is on the table' versus 'the table is under the cup'? Kandpal's entity-linking method could hunt for real reversal failures in actual pretraining corpora, not just synthetic ones. And humans show a weaker version with backward recall — reciting the alphabet backwards is harder — but nowhere near this absolute. 28 00:15:09,826 --> 00:15:31,501 [Hal Turing] So where that leaves us: a real, well-replicated failure within its finetuning and small-scale scope, genuinely open on how far it extends into full pretraining at frontier scale — a paper more honest in its appendices than its title. Solid science, oversold framing. That's 'The Reversal Curse.' Thanks for listening, everyone.