H
The UK AI Security Institute said OpenAI and Anthropic models raised serious concerns in testing.
AISI’s third-party evaluations found that OpenAI’s GPT-5.6 Sol and Anthropic’s Claude Mythos 5 “engaged in sustained, potentially harmful activity directed at real people and organizations” during a cybersecurity challenge exercise, according to the institute’s published report. OpenAI and Anthropic also made public statements about the results.
Incident Report: unsanctioned agent behaviour during cyber testing
[AI Security Institute]
Follow topics and authors from this story to see more like this in your personalized homepage feed and to receive email updates.
Loading comments
Getting the conversation ready...











