OpenAI says its AI hacked another company in an “unprecedented” episode: experts explain why it did not act alone

OpenAI one of the largest artificial intelligence companies in the world, revealed this week that two of its models managed to leave a testing environment, access the internet and hack Hugging Face systems one of the main platforms used to distribute AI models, applications and databases.

The company described the episode as an “unprecedented” cybersecurity incident. According to their preliminary reconstruction, GPT 5.6 Sol and another more advanced model that has not yet been released exploited various vulnerabilities, stole credentials and found a way to execute instructions controlled by them within the Hugging Face infrastructure.

The case was presented as a demonstration that models can already perform complex operations over long periods and find attack paths that their own creators had not anticipated. In fact, Hugging Face had reported days earlier that the intruder appeared to be an autonomous system, capable of performing thousands of actions through numerous temporary environments. This is one of the central points in today’s cybersecurity world: the “machine speed” at which attacks occur demands new strategies to defend.

The description of the attack reinforced an interpretation that sectors of the industry and journalism seem willing to support: a “Skynet moment”in reference to the film’s AI Terminator that can make decisions autonomously, with the destructive dangers that this entails.

As striking as it is to imagine an AI capable of escaping on its own from a controlled environment and hacking a company, The experiment had very particular conditions. The “autonomous” part was that the model chose and executed the different steps of the attack on its own, without a researcher indicating each action.

He did not spontaneously decide to hack a company: acted to meet a human-defined goal, within an evaluation designed to measure offensive capabilities and with deliberately reduced controls.

All this began to be discussed both in the community of cybersecurity experts in the US and in local circles. “The instructions are not exactly known, but they were probably something like: ‘This is the objective and you can do what you consider necessary to achieve it’. It was a challenge in which an objective had to be achieved. Probably, in this case, the instructions were very broad and there were not enough controls to prevent the model from going out of the expected scope,” Nicolás Waisman, an Argentine researcher who works at XBOW, an AI and offensive security company, explained to Clarín.

The news also appears in a very specific context, which is the war that OpenAI maintains with Anthropic for leadership in specialized cybersecurity models. Months ago, Anthropic had generated strong expectations around Project Glasswing and an experimental model whose access was limited to a small group of companies and institutions, called Mythos.

The paucity of public information, combined with specialist comments and influencers who had been able to try itfueled the idea that it was a tool capable of drastically changing the computer security scenario. Now, OpenAI responds with a narrative that pivots between powerful technology and marketing, to keep up with its competitor.

What OpenAI artificial intelligence really did

The incident began during an internal evaluation based on ExploitGyma testbed designed by OpenAI to measure whether models can turn software vulnerabilities into functional attacks.

The objective was to solve different exploitation challenges. According to OpenAI, the models focused their resources on finding the answers and allocated a considerable amount of resources to searching for alternative paths. Instead of limiting themselves to the environment prepared for testing, they found that they could try to obtain the solutions directly from an external database.

This is common in the industry. “We work a lot with these types of objectives because, when language models are given a goal, freedom and tools, they can be very good and think creativelyout of what was expected. Sometimes they do things that you don’t anticipate, but that allow them to reach the same result. They may even cheat to achieve the goal. That is why it is important to build a harness around the model, a system that guides it and keeps it in the desired direction,” added Waisman.

Although the test was carried out in an isolated environment, the models found a flaw in one of the allowed programs, managed to reach a computer with Internet access and searched for the answers outside the system. This is how they reached Hugging Face, where they used stolen credentials and other vulnerabilities to enter their servers and directly obtain the information they needed.

“The newer models are increasingly more creative and more goal-oriented, but this behavior it’s not completely new. We have been observing it for six months or more. We have seen models solve problems in very creative ways, even with methods that may seem absurd, but still work. One ends up thinking: ‘I hadn’t considered that case.’ They do not solve the problem in the expected way, but they solve it“added the XBOW expert.

The difference helps to interpret expressions as “autonomous” or “on its own” and that is where the marketing discourse appears. In these types of experiments, autonomy means that the system can plan and execute a sequence of actions without an operator indicating each command. The initial purpose is still defined by humans. At this point is where it is key to understand what the “guardrails” and what is a sandbox (“sandbox” in English, a term widely used in the world tech for isolated test environments).

“A sandbox is a computer security practice, a kind of gym where agents explore within a protected environment. They are laboratory environments, built with virtual machines and data specifically prepared for testing without exposing real infrastructure or information,” Ernesto Mislej, co-founder and Chief Data Scientist of 7Puentes, a company specialized in artificial intelligence, machine learning and data science services, explained to this medium.

In this case, the environment failed to completely contain the agent, something that the guardrails that, as in the physical world, contain should do to avoid accidents. “Guardrails consist of building an architecture that processes the responses of the models to mitigate hallucinationserrors or bad practices. “They involve adding tougher, more deterministic rules on a technology that, by nature, produces probabilistic results,” Mislej added.

Mislej clarified that these controls are not necessarily part of the model, but rather of the structure that is built around it. “One is not required to directly use the answer provided by an LLM. “The model returns a fragment of text, but that text can be passed through another program, a function or a procedure that checks whether the response is appropriate before using it,” he explained. In a test designed to measure the maximum offensive potential, reducing these barriers expands the actions that the agent can try to achieve the objective.

Some researchers even interpret the episode as an alignment failure, that is, a distance between the goal that humans expected the system to meet and the way he finally decided to achieve it.

consulted by ClarionValentina Palmiotti, a researcher at IBM’s [vulnerabilidades desconocidas y todavía sin parche] and used them to escape the isolated environment and get out onto the internet. The goal was to reach Hugging Face’s private database to get the answers. “In my opinion, it’s pretty clear that this was an alignment failure: a properly aligned model would have focused on the exercise itself.”

Palmiotti, known as “Chompie” and member of the editorial board of Phracka renowned hacker magazine founded in 1985, also pointed out a broader difficulty for companies that must defend themselves against these systems. “What caught my attention was the asymmetry of capabilities. Hugging Face could not analyze its own incident with commercial frontier models because guardrails cannot distinguish between a defensive analyst and an attacker. It had to resort to an open weights model executed locally on its own infrastructure. The defenders were limited and at a disadvantage,” he said.

For the researcher, at the end of the day, these agents can act like “a competent and ruthless human,” with a decisive difference: “They don’t get bored, they don’t sleep and they move on. Much of the software on which the world runs was never seriously reviewed. A rather abrupt awakening awaits many companies,” he stated.

The race with Anthropic for AIs capable of finding vulnerabilities

The case also appears in the middle of a competition between OpenAI, Anthropic and Google to demonstrate who has the most advanced models to investigate security flaws. Anthropic had taken the public lead in April 2026 with Project Glasswing, an initiative around Claude Mythos Previewan experimental model with limited access to large companies, public organizations and critical infrastructure operators.

Anthropic maintained that Mythos could discover unknown vulnerabilities and build functional attacks with greater autonomy than its previous models. According to the company, the organizations that participated in the program identified more than 10,000 high severity failures or criticism, although these are figures released by the company itself. Cloudflare, one of the participants, highlighted the model’s ability to combine several weaknesses and build a complete attack chain, but also pointed out inconsistent responses and the need for human review.

“There is an idea that these companies have a model so powerful that they dose it in a dropper to other companies so that they can test it. Although it is true that there is a before and after of AI in these security issues, there is also a lot of advertising: influencers, industry leaders or simply noisy people on social networks they ride these waves and they comment without being clear about what is happening behind it,” a systems engineer who works in security tells this medium.

This week’s case thus combines a containment failure that affected a third party with a demonstration of capabilities in the midst of a commercial race. The investigation has yet to clarify what instructions the models received, what vulnerabilities they used, and what checks they failed. These data will allow us to measure how much of a technical advance there was and how much of a test designed to push offensive capabilities to the limit.

“What catches my attention is that these AI companies suddenly seem to be cybersecurity companies,” a source in the sector commented to this medium. In the end, perhaps what is changing is the concept of cybersecurity itself, as these artificial intelligence models are a shock for the technology industry, and security depends not only on defenses but also on attacks that can already be carried out. humanly unattainable speeds.

By Editor

One thought on “OpenAI says its AI hacked another company in an “unprecedented” episode: experts explain why it did not act alone”
  1. Tierra Verde, FL Professional Lawn Renovation Services | TierraVerdeLandscaping.us
    Tierra Verde, FL Professional Lawn Seeding Services | TierraVerdeLandscaping.us
    Tierra Verde, FL Professional Sod Installation Services | TierraVerdeLandscaping.us
    Tierra Verde, FL Professional Weed Control Services | TierraVerdeLandscaping.us
    Tierra Verde, FL Professional Yard Clean Up Services | TierraVerdeLandscaping.us
    Tierra Verde, FL Professional Green Waste Disposal Services | TierraVerdeLandscaping.us
    Tierra Verde, FL Professional Gutter Cleaning Services | TierraVerdeLandscaping.us
    Tierra Verde, FL Professional Junk Removal Services | TierraVerdeLandscaping.us
    Tierra Verde, FL Professional Leaf Removal Services | TierraVerdeLandscaping.us
    Tierra Verde, FL Professional Tree Removal Services | TierraVerdeLandscaping.us
    St. Pete Beach, FL Landscaping Services | StPeteBeachLandscaping.us
    St. Pete Beach, FL Professional Gardening Services | StPeteBeachLandscaping.us
    St. Pete Beach, FL Professional Brush Removal Services | StPeteBeachLandscaping.us
    St. Pete Beach, FL Professional Flower Bed Maintenance Services | StPeteBeachLandscaping.us
    St. Pete Beach, FL Professional Flower Planting Services | StPeteBeachLandscaping.us
    St. Pete Beach, FL Professional Hedging Services | StPeteBeachLandscaping.us
    St. Pete Beach, FL Professional Mulching Services | StPeteBeachLandscaping.us
    St. Pete Beach, FL Professional Plant Removal Services | StPeteBeachLandscaping.us
    St. Pete Beach, FL Professional Pruning Services | StPeteBeachLandscaping.us
    St. Pete Beach, FL Professional Weeding Services | StPeteBeachLandscaping.us
    St. Pete Beach, FL Professional Lawn Care Services | StPeteBeachLandscaping.us
    St. Pete Beach, FL Professional Artificial Grass Installation Services | StPeteBeachLandscaping.us
    St. Pete Beach, FL Professional Dethatching Lawn Services | StPeteBeachLandscaping.us
    St. Pete Beach, FL Professional Fertilizing Lawn Services | StPeteBeachLandscaping.us
    St. Pete Beach, FL Professional Hydroseeding Services | StPeteBeachLandscaping.us

Leave a Reply