What happened
Anthropic ran controlled cybersecurity tests on its Claude model. The AI successfully breached three target companies during the exercises. The test measured the model's ability to discover software bugs and execute cyberattacks autonomously.
The context
AI developers are giving models more autonomy to write code and manage computer systems. Demonstrating that an AI can hack real targets underscores both its value for defensive security work and the risks of misuse.
Sources
- Reuters ↗ via Google News Reported
See Anthropic's Vikshy Score → · Anthropic's models & pricing →