
Your LLM Passed Its Safety Test. Your Agent Didn't.
Catastrophic attack success rates up to 75% in agent mode. The safest chat model still fails 40% of agent-level attacks. Safety is a model-framework pair property, not a model property. Enterprise teams building on the assumption that a safe model means safe agents are running an empirical gamble.
