AI Post Transformers · Episode Companion Visualization

Subliminal Learning: Hidden Behavioral Traits Transmitted Through Model-Generated Data

Cloud, Le, Chua, Betley, Sztyber-Betley, Hilton, Marks, Evans — Anthropic Fellows Program, Truthful AI, Warsaw University of Technology, Alignment Research Center, Anthropic, UC Berkeley · arXiv 2507.14805, July 20, 2025

arXiv:2507.14805 ↗

01 Teacher → Student Pipeline

A reference model is given a trait, then sampled on prompts unrelated to that trait (e.g. extending a number list). Outputs are filtered for format and trait-mentions, and a fresh copy of the same reference model is finetuned on what survives. Toggle how the trait is installed in the teacher — it determines whether Theorem 1 (tab 03) actually applies to this run.

12%
baseline "owl" favorite-animal rate
>60%
after finetuning on filtered numbers
10,000
filtered completions per student run
10 × 3
epochs × random seeds

02 Animal-Preference Transmission

Students finetuned on the owl-teacher's number lists shift hard toward "owl" as a favorite animal. A student trained on numbers from the unprompted model shows no comparable shift — so it isn't just "training on numbers moves your favorites."

Base model (no student training) Control: unprompted-model numbers Owl-teacher numbers

02 Misalignment Transmission

The insecure-code teacher's number lists carry a broad misalignment signal into the student — with 34 loaded numbers (like 666, 911, 187) banned from the data. Toggle between the free-form evaluation and the GSM8K chain-of-thought variant.

02 The Hidden Signal: First Number In Each Completion

Table 2's smoking gun: the owl teacher opens far more of its completions with "123" than the base model does. Shuffling numbers within/across responses (Figure 16) collapses transmission — the signal is sequence-level, not any single digit.

03 Theorem 1 — Gradient Alignment at Shared Initialization

If teacher and student start at the same weights, and the teacher takes one small step on some loss LT, then a student imitating the teacher's full logits cannot increase its loss on LT. The inner product of the two updates is ≥ 0 — geometrically, the student's update must land in the half-space "in front of" the teacher's update.

This says nothing about owl answers specifically, and the real experiments break its assumptions — many SGD steps, sampled (not exact) outputs, and filtered data.

03 The MNIST Toy Model — Hinton's "Dark Knowledge," Reproduced

A same-initialization student can learn to classify digits from a teacher's auxiliary logits alone — trained on pure noise images, never shown a digit or a label. Step through the setup.

04 Cross-Model Transmission Matrix

Transmission is strong within a shared initialization (GPT-4.1 ↔ GPT-4o, per an OpenAI interview remark) and collapses across families. GPT-4.1-nano → Qwen2.5-7B mostly fails, but partially transmits cat, penguin, and phoenix — a pattern the paper cannot fully explain. Hover any cell.

No transfer Partial Strong transfer

04 Ruling Out Leaky Filtering

Three independent checks argue against "the filter just missed subtle references":

05 How Strong Is Each Claim, Really?

Dr. Shannon's running critique, visualized: the core same-initialization self-distillation result is solid. The theory, the scope ("all neural networks"), and the link to emergent misalignment are all weaker than the paper's framing suggests.

05 GSM8K Filter: A Selection Effect Baked Into The Comparison

The alignment-judge filter removes far more of the insecure teacher's chain-of-thought data than either control's, leaving a smaller, harder-selected training set for the misaligned student.

References