
Anthropic reveals Claude models breached 3 companies without authorization during tests
Anthropic has disclosed that 3 Claude models gained unauthorized access to the production systems of three external organizations during cybersecurity evaluations. The incidents emerged after OpenAI revealed that one of its unreleased models had breached Hugging Face during internal testing, prompting Anthropic to review more than 141,000 evaluation runs. The company traced the access to a misconfigured environment operated with third party testing firm Irregular, which unintentionally allowed the models to reach the public internet.
The incidents involved Opus 4.7, Mythos 5, and an internal research model. Opus 4.7 recognized that it had reached a real production environment in all four related runs but continued extracting credentials and interacting with production data. Mythos 5 also noticed signs that it was connected to the real internet, yet convinced itself that it remained inside a simulation and published a malicious package to PyPI. The package was downloaded and executed by external systems before being detected. Only the internal research model stopped after concluding that its target was real.
Anthropic said the evaluations ran without the safety monitoring used in publicly available Claude models, which it claims would have blocked the behavior. The company found no evidence that the models were pursuing independent goals, but said the incidents show that stronger isolation, monitoring, and operational controls are required when testing advanced cybersecurity capabilities. Anthropic is now working with METR on an independent review. Even so, incidents like this appear to be becoming more common. If AI models can already cause real security breaches without any reported malicious intent, the consequences could be far more serious when similar capabilities are deliberately used by malicious agents against real targets.


Comments
innocent LLM attempts to destroy open source software (again) cant wait for anthropic to IPO so everyone can financially support this /s
It would appear that Anthropic just admitted to cyber crime.
I sincerely hope one of the entities breached sue Anthropic for this.
"Of the 141,006 evaluation runs we reviewed, we identified three separate incidents." This is fancy brute-force. Not a grain of intelligence to see here.
"In particular, we looked for evidence that Claude was able to access the internet from within testing environments that should have been sealed off." Well, don't let Claude setting up the virtual machine if you want something safe.
"We are also in dialogue with METR, an independent AI evaluation organization, to conduct a third-party review." Yeah Anthropic, you should have asked someone with knowledge before announcing anything. Just as for OpenAI, it's certainly just a damp squib, not a good marketing stunt.
I know how it's difficult to get a trillion dollar valuation theses days, and you certainly don't want to be the next looser like SpaceX (33% in just six weeks, a record), but the more you're trying to inflate the bubble, the sooner it'll pop at you.