AI Post Transformers arXiv:2603.21396 Interactive Viz Link

How Models Detect Hidden Activation Steering

A visual tour of detection versus identification, the residual-stream intervention setup, and the paper’s claim that post-training installs circuitry that can report hidden activation tampering without spraying false alarms.

model: Gemma3-27B layers: 62 total best injection: layer 37 strength: α = 4 concepts: 500 success/failure: 242 / 258
Detection Headline
0% FP
Refusal Ablation
10.8 → 63.8%
Ridge Explained
44.4%
LDA Split
75.6%

References

Core sources, nearby mechanistic context, and prior episodes connected to this result.

Mechanisms of Introspective Awareness Uzay Macar et al., 2026
arXiv:2603.21396
Looking Inward Binder et al., 2024
Scholar search
Activation Addition Turner et al., 2023
Scholar search
Representation Engineering Zou et al., 2023
Scholar search
Scaling Monosemanticity Templeton et al., 2024
Scholar search
Circuit Tracing Ameisen et al., 2025
Scholar search
Episode: Anthropic: Introspective Awareness in LLMs AI Post Transformers, 2025
Open episode
Episode: Neural Chameleons and Evading Activation Monitors AI Post Transformers, 2026
Open episode
Extracted arXiv IDs from transcript: 2603.21396