This page treats the episode as a mechanics problem, not a philosophy essay: where does a model keep the tiny slice of state it can report on, reason over, and deliberately steer? Hover the visuals. Click the tabs. Most numbers here are synthetic, but shaped to match the paper’s qualitative findings and intervention style.
Click the paper’s five functional claims and watch the same mid-layer corridor change role.
The key move is averaging Jacobians over many contexts so the probe finds what is poised for report, not only what is easy to decode in one prompt.
Choose a case. Then compare a clean steering edit against partial and full ablations.
The paper’s stronger claim is not only that hidden meaning exists, but that one narrow subspace is especially reusable and especially vulnerable to targeted edits.
Compact pointers to the papers and adjacent threads most relevant to the episode’s visual argument.