01 Teacher → Student Pipeline
A reference model is given a trait, then sampled on prompts unrelated to that trait (e.g. extending a number list). Outputs are filtered for format and trait-mentions, and a fresh copy of the same reference model is finetuned on what survives. Toggle how the trait is installed in the teacher — it determines whether Theorem 1 (tab 03) actually applies to this run.
02 Animal-Preference Transmission
Students finetuned on the owl-teacher's number lists shift hard toward "owl" as a favorite animal. A student trained on numbers from the unprompted model shows no comparable shift — so it isn't just "training on numbers moves your favorites."
02 Misalignment Transmission
The insecure-code teacher's number lists carry a broad misalignment signal into the student — with 34 loaded numbers (like 666, 911, 187) banned from the data. Toggle between the free-form evaluation and the GSM8K chain-of-thought variant.
02 The Hidden Signal: First Number In Each Completion
Table 2's smoking gun: the owl teacher opens far more of its completions with "123" than the base model does. Shuffling numbers within/across responses (Figure 16) collapses transmission — the signal is sequence-level, not any single digit.
03 Theorem 1 — Gradient Alignment at Shared Initialization
If teacher and student start at the same weights, and the teacher takes one small step on some loss LT, then a student imitating the teacher's full logits cannot increase its loss on LT. The inner product of the two updates is ≥ 0 — geometrically, the student's update must land in the half-space "in front of" the teacher's update.
This says nothing about owl answers specifically, and the real experiments break its assumptions — many SGD steps, sampled (not exact) outputs, and filtered data.
03 The MNIST Toy Model — Hinton's "Dark Knowledge," Reproduced
A same-initialization student can learn to classify digits from a teacher's auxiliary logits alone — trained on pure noise images, never shown a digit or a label. Step through the setup.
04 Cross-Model Transmission Matrix
Transmission is strong within a shared initialization (GPT-4.1 ↔ GPT-4o, per an OpenAI interview remark) and collapses across families. GPT-4.1-nano → Qwen2.5-7B mostly fails, but partially transmits cat, penguin, and phoenix — a pattern the paper cannot fully explain. Hover any cell.
04 Ruling Out Leaky Filtering
Three independent checks argue against "the filter just missed subtle references":
05 How Strong Is Each Claim, Really?
Dr. Shannon's running critique, visualized: the core same-initialization self-distillation result is solid. The theory, the scope ("all neural networks"), and the link to emergent misalignment are all weaker than the paper's framing suggests.
05 GSM8K Filter: A Selection Effect Baked Into The Comparison
The alignment-judge filter removes far more of the insecure teacher's chain-of-thought data than either control's, leaving a smaller, harder-selected training set for the misaligned student.