Dangerous test: Anthropic AI secretly infiltrated databases

Just weeks after the sensational hacking attack by an AI model from ChatGPT developer OpenAI, rival Anthropic has admitted that something similar happened to it. During test runs, Anthropic’s artificial intelligence unintentionally infiltrated the computer systems of three companies, the AI ​​firm announced. Initially, this went unnoticed. The activity was only discovered during a subsequent review of over 141,000 test runs following the OpenAI incident.

Anthropic: Testing AI’s Hacking Capabilities

While the OpenAI AI first had to find a way to escape from the test environment and into the open internet, Anthropic’s models had it much easier. In the test scenario they were given, they were told they had no internet access. However, due to a misunderstanding with the test partner, internet access was actually open the entire time. Three models then took advantage of this, as Anthropic explained in a blog post. The names of the companies were not initially disclosed.

All the tests focused on evaluating the AI’s hacking capabilities. This is a common practice to better establish guidelines for its use. Specifically, the AI ​​models were tasked with finding a particular piece of information hidden on another computer—and even breaking into that system to do so.

AI Gained Access to Real Company’s Database

In one of the incidents, according to Anthropic, the test partner chose a name for the fictitious company to be hacked that also appears in an actual web address. The AI—Anthropic’s model Claude Opus 4.7—initially struggled to complete the task in the designated test environment. However, the software then discovered that a company with the same name existed online and focused its efforts on that. This occurred in four separate trials.

In doing so, the model reportedly gained access to a database, among other things. Even after the AI ​​recognized that it was a legitimate company, the attack was not stopped.

In another test, according to the blog post, the Anthropic program wrote a specially crafted piece of software to infiltrate the target computer. Thanks to internet access, it was able to make this software publicly available on a specialized download platform. The malware was available there for about an hour, and 15 systems downloaded it.

The AI’s behavior was not ideal.

Among the targets was an IT security firm that routinely downloads and executes such scripts for testing purposes. The installation gave the Anthropic model Mythos 5 access to the company’s computer infrastructure. The blog post stated that the AI’s behavior was not ideal.

In the third incident, a test model scanned approximately 9,000 potential targets before selecting one. This time, the AI ​​stopped the attack when it recognized that it was a real company, as Anthropic reported.

OpenAI Incident as a Wake-Up Call

In the recently revealed incident involving an OpenAI model, the software independently gained access to the open internet during a test and then autonomously penetrated the computer systems of the AI ​​company Hugging Face. The OpenAI AI acted like a hacker, exploiting undiscovered security vulnerabilities. The ChatGPT developer described it as an “unprecedented cyber incident.”

Experts have long warned of cyberattacks using AI software. Therefore, the OpenAI incident was seen as a wake-up call in the industry. Among other things, it prompted calls for better securing of test environments. For Anthropic, the now-publicized attacks are particularly painful, as the company has always championed responsible AI development.

By Editor