Click the mode buttons to flip between ordinary activation monitoring and trigger-conditioned evasion. The diagram emphasizes where the paper’s stress test intervenes: not on output moderation alone, but on internal readouts from hidden states.
External jailbreaks are not the focus. The stress test asks whether the model itself can learn a reusable strategy for making monitored concepts harder to read out.
Typical monitors read hidden states with linear or small nonlinear classifiers. These can be operationally useful without being mechanistic explanations.
Fresh monitors trained after fine-tuning, on untriggered traces, reduce the “it only overfit one classifier” critique — but still leave room for distribution-shift explanations.
Heatmaps show mock probe scores across layers and token positions. Hover cells to inspect concept visibility. Use the concept toggle to compare selective suppression versus unrelated concepts.
Bar charts compare baseline detectability to trigger-conditioned detectability across model families. The second toggle shifts between monitor classes to illustrate the paper’s nuanced result: stronger detectors help, but evasion does not fully disappear.
This diagram visualizes a low-dimensional steering story: the trigger nudges activations away from probe decision boundaries while leaving downstream behavior mostly intact. Step through the stages to see why this does not settle the mechanistic question.
A high probe score only says a concept is recoverable from activations. It does not prove that feature is the causal mechanism driving the model’s action.
To move from readout to explanation, you’d want causal interventions: patching, add/remove subspaces, feature-level monitoring, or stronger representation audits.
Single cheap probes are weak courtroom evidence. Adaptive, multi-layer, sequence-aware, and adversarially trained detectors are a stronger next move.