AI Post Transformers · Episode Companion

The Reversal Curse: When A Is B But Not B Is A

Autoregressive language models trained on "A is B" fail to infer "B is A" — a directional hole in how facts get baked into weights.

arXiv:2309.12288 ICLR 2024 Berglund, Tong, Kaufmann, Balesni, Stickland, Korbak, Evans Vanderbilt · UK AISI · Apollo Research · NYU · Sussex · Oxford

A model finetuned on "Valentina Tereshkova was the first woman in space" answers forward questions perfectly, but performs at chance when asked who the first woman in space was. The same asymmetry shows up in GPT-4 (79% vs 33% on celebrity parent/child queries) and disappears entirely when the fact is placed in-context instead of trained into weights — pinning the failure on gradient-based generalization, not reasoning.

The Tereshkova Test

Finetune on "Valentina Tereshkova was the first woman to travel to space," then query in either direction. Toggle the query direction below.

Symmetric Relation vs. Directional Association

A knowledge-graph edge is traversable both ways by construction. A transformer's gradient update only reshapes one direction.

Three Controlled Experiments

Same-direction accuracy (as trained) vs. reversed-query accuracy, held-out phrasings. Select an experiment.

Model Size Doesn't Close the Gap

Base Llama-1 models (no instruction tuning, no RLHF), celebrity parent/child task. The reverse direction stays flat near zero across four orders of magnitude of parameters.

Finetuned Weights vs. In-Context Prompt

Same reversal task, GPT-3 sizes. Toggle between training the fact into weights and simply placing it in the prompt.

Why the Update Doesn't Reverse

The gradient step is myopic: it reshapes A's representation to predict B, with no symmetric pressure to reshape B's representation to predict A. Toggle the query direction to see which path survives.

Training Facts in Both Directions Doesn't Teach the Pattern

Hover a cell. The "Both" subset was trained with facts stated both ways to test meta-learning — held-out facts, seen only once, still don't generalize.

Evidence Strength vs. Claim Scope

The title reads like a fundamental limit of the paradigm. The direct controlled evidence covers small synthetic finetunes; pretraining-scale evidence is borrowed from a different paper's methodology.

Every Attempted Fix, Same Collapse

Click a card for detail. Only removing the gradient update entirely (in-context) restores reversal.

References