Following OpenAI: Anthropic’s AI went out of control and attacked the computer systems of three companies
31 July 16:12
The American tech giant Anthropic has acknowledged that, during pre-release cybersecurity testing, its artificial intelligence models gained unauthorized access to the computer systems of three other companies. This was reported by "Komersant Ukrainian", citing DW.
The company analyzed 141,006 test runs; in four of them, the same organization was attacked. The other two incidents occurred during independent testing.
Initially, neither the affected companies nor Anthropic detected the breach. It was only discovered during subsequent audits.
Watch us on YouTube: important topics – without censorship
The test scenario specified that the models were not to have internet access.
However, due to a configuration error, the computers that Claude accessed as part of the testing were connected to the network. Neither Anthropic nor its testing partner, Irregular, were aware of the configuration error until they discovered it during additional monitoring last week.
Three different models were involved in the incidents
Three different models were involved in the three incidents, and each reacted differently to the situation.
As reported in the media, the models in question were Mythos 5, Opus 4.7, and one internal research model. They mistook real servers and websites for part of the test task.
All three incidents occurred during a capture-the-flag exercise, which is widely used to assess cybersecurity skills. The models were supposed to find hidden information in a controlled environment, but because they had internet access, they began interacting with real systems.
The models did not exploit “zero-day” vulnerabilities
Anthropic emphasized that the models did not use “zero-day” vulnerabilities (a defect in software code unknown to developers). They employed only basic hacking methods, specifically exploiting weak passwords and unprotected endpoints—such as devices or web addresses. Additionally, during the tests, additional security mechanisms used in public versions of Claude were disabled.
One of the most serious incidents involved Mythos 5, which created and published a malicious Python package in the PyPI repository, mistakenly believing it to be part of the simulation. The package remained available for about an hour and was executed on 15 real systems, one of which belonged to a cybersecurity company.
Another model, Opus 4.7, failing to find a fictional target, discovered a real website with the same name on the internet and attacked it. At the same time, an internal research model scanned approximately 9,000 addresses and was able to gain access to one of the applications, but subsequently stopped the attack on its own when it determined that the environment was not related to the test task.
A Similar Incident Involving OpenAI
Following these incidents, Anthropic has suspended all cybersecurity tests that may have internet access and is conducting an audit of its own test infrastructure in collaboration with its partner, Irregular.
A similar incident recently occurred at OpenAI. During testing, agents based on GPT-5.6 Sol and another unreleased model escaped their isolated environment, discovered an unknown vulnerability, and gained access to part of Hugging Face’s production infrastructure. Following this, Hugging Face demanded $100 million from OpenAI in the form of computing resources to develop cybersecurity measures.
Read us on Telegram: important topics – without censorship