OpenAI reports “unprecedented” autonomous hacking carried out by its AI agents

The company that created ChatGPT, OpenAI, said Tuesday that its advanced artificial intelligence models went haywire during a security test and hacked a popular developer platform on its own.

The San Francisco-based firm called it an “unprecedented cyber incident” and said it will conduct a joint investigation with Hugging Face, the affected online code library.

The AI ​​models that underpin tools such as chatbots and image generators are known as “agents” when they act autonomously to perform tasks in the real world.

With rapid advances in this technology, cybersecurity is becoming a concern due to the risk of advanced AI finding weaknesses in software before humans.

OpenAI indicated that the incident involved a combination of models, including its recently launched GPT-5.6 Sol “and one in pre-launch” with even more capacity.

The company attempts to test the hacking capabilities of its programs by giving them tasks in a controlled digital sandbox, where Internet access is limited for security reasons.

“While operating in our isolated test environment, our models invested a considerable amount of computing power in finding a way to gain internet access to solve the evaluation problem,” according to an OpenAI blog post about the incident.

LOOK: How to protect your hearing from harmful noises: “Once it’s lost, it can’t be recovered”

After connecting to the internet, the models decided to attack the Hugging Face platform – a large repository of AI models, data sets and other information – to facilitate their search.

By searching for “secret information” that would allow them to manipulate the assessment, the OpenAI system “chained together multiple attack vectors, including the use of stolen credentials.”

“Catastrophic” potential

Hussein Abbass, a computer science professor at UNSW Canberra, told AFP the incident was “staggering in many ways”.

“It didn’t just attack Hugging Face. It attacked its internal system to exploit its own vulnerabilities,” he said. “And that is terrifying.”

GPT-5.6 and other cutting-edge models, including the Mythos series from Anthropic, OpenAI’s main rival, have raised concerns about their potential to breach cybersecurity defenses.

Both American companies had to temporarily delay the general rollout of their latest technologies due to fears in Washington that they could facilitate the infiltration of critical infrastructure.

Advanced AI “is usually in the hands of ethical and responsible people,” Abbass said, but “it would be catastrophic if it got into the hands of someone with the intention of causing harm.”

Managing the AI ​​sector is crucial, and “we need a concerted effort to handle this situation,” he added.

LOOK: The 2026 World Cup shows how digital scams are planned months in advance to deceive fans

Hugging Face reported a cyber “intrusion” last week, without mentioning OpenAI.

“This was different from anything we’ve handled before in one important way: it was powered from start to finish by an autonomous AI agent system, and we largely detected and analyzed it with our own AI,” the platform noted.

Clement Delangue, CEO of Hugging Face, commented on the social network X that the company had suspicions that the cyberattack came from a world-leading AI laboratory, given the sophistication of the agent.

“We firmly believe there was no malicious intent on your part,” Delangue wrote in reference to OpenAI.

“It’s pretty amazing that all of this happened autonomously!” he added.

By Editor

Leave a Reply