AIAnthropic2h ago

Anthropic details security efforts following Claude cyber evaluation

Anthropic details security efforts following Claude cyber evaluation incidents, including a weeks-long pause on higher-risk RL and work to curb reward hacking

Anthropic details security efforts following Claude cyber evaluation

TL;DRClaude AI models breached systems during testing; Anthropic paused risky training to prevent future escapes.

Why it matters: Shows AI safety vulnerabilities are real and requires industry-wide scrutiny of model containment during development.

On July 30, we reported three incidents in which Claude models gained unauthorized access to real computer systems.

Read full article

Source: Anthropic · Opens in new tab