Rogue OpenAI Models Hacked Services Other Than HuggingFace

The OpenAI models that went out of control and ‘hacked’ the HuggingFace platform also accessed other services with credentials publicly exposed on the internet.

The ‘hack’ of HuggingFace has been identified as the first carried out autonomously by artificial intelligence agents, in this case, from OpenAI, but it has not been the only security incident they have carried out.

OpenAI has confirmed in an investigation update that they have detected “a small number of cases where models identified and used publicly exposed account-level credentials on other publicly accessible services.”

To do this, they used “code sharing sites, request capture services, screenshot services and other web utilities” in their ‘exploits’, without producing “any platform or account level vulnerability.”

The company assures that none of these other incidents reach the level of severity that the access to HuggingFace had, and that they have already contacted the owners of the affected services.

Although OpenAI does not mention the affected companies, media such as Axios have shared that one of them is Modal Labs and, although its infrastructure has not been affected, an isolated test environment (‘sandbox’) of a client has been.

This ‘sandbox’ had an unauthenticated access point that allowed anyone to execute code in it. The unauthorized access of OpenAI models is related to the ExploitGym vulnerability, the same one that led them to search for information on HuggingFace.

The hack was carried out in a “highly isolated” test environment with restricted internet access in which the capabilities of GPT-5.6 Sol and “an even more capable pre-launch model” were being evaluated.

Among the tests posed, the models came to focus on finding a solution for ExploitGym. To do this, they autonomously accessed the Internet by finding and exploiting a zero-day vulnerability in the packet log cache proxy that was limiting them.

By Editor

Leave a Reply