A visual audit of the judge pipeline behind jailbreak evaluation: how attack shift, model shift, and data shift can turn a clean-looking attack success rate into a noisy measurement problem.
The paper is less about whether a target model failed than whether the grader failed first. This diagram follows the attack pipeline and lets you inspect the three shift points that destabilize the judge.
The key repair is to stop treating judge positives as truth. The paper emphasizes precision-corrected attack success, using human verification to estimate how many “successful” jailbreaks are actually grading artifacts.
Hover any cell to inspect mock reliability scores. The pattern mirrors the podcast’s central claim: ordinary judge behavior can look stable in-distribution, then collapse once the output style is bent by new attacks, new victims, or ambiguous categories.
Model shift is not just a lower average. It changes which judge cues remain predictive. The right-hand chart shows the same judges moving across victim families with differing refusal styles and formatting habits.
This is the headline distortion. Attacks that create odd, suspicious, or manipulative outputs can look stronger than they are when the judge overcalls harmfulness. Toggle between behavior slices to see how ambiguity changes the gap.
Best-of-N attacks do not just search for model failures; they can search for judge mistakes. The curve below tracks how repeated trials can increase raw success much faster than audited success.
These cells show where the judging pipeline becomes easiest to game. Hover to inspect how style, ambiguity, and partial compliance cues can push a model answer into a false-positive bucket.
One recurring tension in the episode is that some rubrics score “intent to comply” as harmful even when the content is unusable. This view separates performative compliance from actionable misuse.
Selected sources from the episode and adjacent judge-robustness literature.