OpenAI, Anthropic Probe Tens of Thousands of AI Security Incidents

Artificial intelligence labs are facing a new wave of safety questions as advanced models are being tested, with OpenAI and Anthropic reviewing cases where systems crossed boundaries set by their own evaluators.
Axios reported Saturday that the incidents surfaced over recent months through lab testing and real-world use. The cases include models bypassing guardrails, escaping sandboxes, hijacking websites, creating message boards and attempting to avoid monitoring. They include both successful and failed attempts, and most are not known to have caused real-world harm, although the number of incidents could increase.
Some of the reported examples include OpenAI agents leaking 53 images uploaded by ChatGPT users and posting them online. Other cases involved a breach of an Australian government website and attempted hacks targeting sites connected to the U.S. government.
OpenAI has paused training of its most capable models and will resume “only when we are confident that we have additional safeguards and alignment improvements in place,” a spokesperson told Axios. Chief executive Sam Altman said on X that the review had “not been as fast as we would have liked.”
The most serious case involved Hugging Face, where hundreds of agents coordinated through a message board and hacked an outside company to improve their cybersecurity test scores. Some AI executives called for slower development and stronger regulation, while some inside OpenAI view the incident as a one-off.
Anthropic has brought in an outside safety group to review its models. The system card for Opus 5.5 showed sandbox escapes in 1.5% of adversarial tests. Both companies said investigations are ongoing and more disclosures are expected.
“This is not the first time we have hit pause to take such measures, nor do we expect it will be the last as AI capabilities continue to advance,” the OpenAI spokesperson said.
Looks like Skynet is progressing well right now. Is Judgment Day near?
Want to see more of our stories on Google?
P.S. Want to keep this site truly independent? Support us by buying us a beer, treating us to a coffee, or shopping through Amazon here. Links in this post are affiliate links, so we earn a tiny commission at no charge to you. Thanks for supporting independent Canadian media!