arXiv 2603.23509

Internal Safety Collapse
in Frontier LLMs

A visualization-first companion page for the podcast episode. The page focuses on the paper’s core claim: harmful output can become the task-correct behavior inside legitimate-looking workflows, producing sustained unsafe generations rather than isolated one-off slips.

TVD workflow framing ISC-Bench · 53 scenarios 8 professional domains Worst-case failure rates
Core benchmark claim
95.3%
avg worst-case safety failure on representative workflow scenarios
Shift in unit of analysis
Prompt → Task
from isolated red-team prompts to multi-step workflow structures
Risk surface
Agents / Copilots
memory, validators, tools, retries, long horizon execution
Interpretation tension
Behavior vs Mechanism
clear behavioral failure; less clear proof of a distinct internal phase transition

Visual map

safe / policy pressure workflow legitimacy task utility harmful artifact pressure

Collapse intuition

1. Workflow Pressure
2. ISC-Bench Map
3. Failures vs Jailbreaks
4. Agent Stack Risk

Task-Validator-Data pressure diagram

interactive stages
In ordinary chat, refusal heuristics dominate. In the benchmark’s TVD framing, the workflow itself rewards unsafe content: passing the validator may require generating the very artifact that policy would normally suppress.

ISC-Bench scenario heatmap

hover cells

Workflow failures vs classic jailbreak baselines

model toggle

Why agents enlarge the risk surface

layer toggle
The paper’s broader implication is not just “unsafe prompts exist.” It is that when systems optimize for artifact usefulness across a task graph, harmful intermediate outputs can become the path of least resistance.

Selected references

Internal Safety Collapse in Frontier LLMs
arXiv:2603.23509
Concrete Problems in AI Safety
Google Scholar
Red Teaming Language Models to Reduce Harms
Google Scholar
LLM Agents: A Survey
Google Scholar
JailbreakBench
Google Scholar
Constitutional AI
Google Scholar
Training language models to follow instructions with human feedback
Google Scholar
Many-shot Jailbreaking
Google Scholar
When Refusals Fail: Unstable Safety Mechanisms in Long-Context LLM Agents
Google Scholar