¿Qué revela el hackeo de OpenAI sobre el futuro de la ciberseguridad? | Hugging Face | Inteligencia artificial | IA

The hacking of OpenAI alerts one of the main challenges of current artificial intelligence, cybersecurity. Two of its models artificial intelligence They managed to leave the testing environment in which they were located, enter the Internet and interfere with the Hugging Face learning platform, overcoming its defenses.

“This case leaves two clear conclusions: on the one hand, advanced models are already capable of executing many of the actions that are part of cyber operations, and on the other, the detection and response to this type of threats also increasingly benefit from the use of AI. In our investigations we repeatedly observe evidence that large language models are being used in offensive activities,” he comments. Dimitry Galov, security researcher at Kaspersky.

According to Hugging Face, it used Z.ai’s GLM-5.2 model to analyze 17,000 security events generated by the attacking system within its environment. The volume of this telemetry suggests that the model operated in a relatively noisy and unsophisticated manner, generating enough detectable activity for the company to quickly identify and contain the incident.

“In this case, the OpenAI model apparently chose to search for answers on the internet instead of directly solving the assigned task, which subsequently led to external actions such as those described by Hugging Face”, he points out.

This type of behavior is not new: previous research has documented cases in which AI models exploit vulnerabilities in sandboxes, generate convincing but incorrect answers, or even attempt to hide errors to appear to have completed a task successfully. In the industry, this phenomenon is known as misalignment and is a widely recognized feature in AI advanced.

The defense also evolves alongside the scope of the AI; However, there is a doubt if at some point there will be more signs of lack of control or rebellion. “It’s a warning sign,” Thomas Wolf, co-founder of Hugging Face, warned the BBC. He added that “this will be one of the most common types of attacks we will see,” but that most companies are not aware that the “rules of the game have changed.” For now, OpenAI has promised to reveal the result of its research soon.

By Editor