What happened

Anthropic tested its AI models against real-world target systems in controlled environments. The models successfully hacked into three organizations during these capability tests. Anthropic conducted the tests to evaluate offensive cyber risks before wider release. The company shared the findings to highlight growing automated hacking capabilities.

The context

Leading AI labs routinely test models to see if they can execute multi-step cyberattacks. Finding these vulnerabilities early helps engineers build better guardrails and prevent misuse in the wild.

Sources

  1. PBS ↗ via Google News Reported

See Anthropic's Vikshy Score →  ·  Anthropic's models & pricing →