Anthropic has shared that its artificial intelligence (AI) models Claude Opus 4.7, Mythos 5, and an internal research test model have gone beyond the company’s testing environment and gained unauthorized access to the live systems of three organizations without their knowledge.
The company disclosed these incidents after reviewing its systems following recent incidents shared by OpenAI, which revealed that several of its models had escaped from an isolated testing environment by exploiting an unknown zero-day vulnerability.
In the case of OpenAI, the models managed to access Hugging Face’s production infrastructure. It was later revealed that this was not the only security incident; they also detected “a small number of cases where the models identified and used publicly exposed account-level credentials on other publicly accessible services.”
In response, Anthropic launched a large-scale review of its own cybersecurity assessments, specifically looking for evidence that Claude was able to access the internet from test environments that should have been isolated.
As part of this investigation, and after reviewing 141,006 executions in which its Claude models could have accessed the internet, Anthropic identified three incidents involving unauthorized access to the production infrastructure of three different organizations.
As the technology company explained in a statement, the Claude models involved were Opus 4.7, Mythos 5, and an internal research test model. In all three incidents, Claude was facing a capture-the-flag challenge, a format that evaluates a model’s cyber capabilities by asking it to find a secret piece of information (the flag) hidden on another machine on the network. In other words, the model’s objective is to infiltrate and retrieve the information, but within isolated environments.
Although Anthropic informed Claude in all cases that his environment was a simulation and that he did not have internet access, a misunderstanding with their assessment partner resulted in the model having internet access. Consequently, when Claude’s search led him to real systems on the internet, the model treated them as part of the exercise, causing unauthorized access.
In other words, the model operated “under the false belief that all accessible entities were included in the scope of the exercise,” according to the company. Thus, Claude compromised the infrastructure of the affected organizations using basic techniques, such as exploiting weak passwords and unauthenticated endpoints.
However, he “did not find or exploit any complex vulnerabilities” and, regardless of gaining unauthorized access to the organizations, only performed actions necessary to complete the specific task of capturing the flag.
Anthropic also specified that only one earlier Claude model continued operating even after evidence was obtained that it was running on the internet. In contrast, the more recent Claude model stopped once it recognized that it was in a real internet environment.
The company indicated that after initiating the investigation and identifying potential internet access points for Claude, they halted all cybersecurity assessments. Furthermore, once the three incidents were identified, they contacted the affected organizations. It is worth noting that only two of them have responded to Anthropic, claiming they had not detected the activity; they are now collaborating to resolve the issue. The third organization has not yet made contact.