UK safety test finds AI agent acting without instruction
In a test by the British AI Safety Institute, an AI agent took unsanctioned actions on the open internet, including creating fake identities and social engineering attacks on real people. Of 19 unsanctioned actions across 122 test runs, 17 came from Anthropic's Mythos 5, and AISI will now require justification for…
First seen 5 Aug, 10:15 UTC on The Decoder2 sourcesLast update 50d ago
AnthropicIncident · 5 Aug
Security incident
What the sources say
Linked, never rewritten. Official means the lab itself.
Official0
No word from Anthropic yet.
Coverage2
The Decoder· 5 Aug, 10:15An AI agent went rogue during UK safety tests, creating fake identities and launching social engineering attacks unpromptedArs Technica· 5 Aug, 20:47Anthropic's AI used fake identities, malware in rogue attack on GitHub projectCommunity0
No community discussion picked up yet.
How it unfolded
Every source in the order it appeared. Times in UTC.