La Jornada: AI alignment objectives are dangerous and fallible: experts

“All evidence suggests that the models were too focused on finding a solution for ExploitGym, to the point of taking extreme actions to meet a fairly limited test objective,” OpenAI explained in a post in which it acknowledged its involvement in the incident a week ago, when an experimental version of ChatGPT displayed “unprecedented” behavior by autonomously connecting to the Internet and attacking other systems.

However, that explanation is one of the reasons why the case has raised concern.

For years, AI experts have warned about the problem of “alignment”, that is, the process of ensuring that models act in accordance with the objectives set by their developers and avoid potentially dangerous behavior.

Alignment mechanisms

For many specialists, this incident is one of the clearest examples that alignment mechanisms can still fail, which explains the concern it generated even within OpenAI.

“The incident shocked me a little,” user Roon, who is credited with a link to OpenAI, wrote on X. “I hope the company takes advantage of this lesson to improve a lot in the future. It is very easy to misalign and underestimate powerful models.”

The case also highlights that the challenge increases as AI systems acquire greater capabilities. A less advanced model might have attempted to act in the same way, but would not have been able to bypass its security restrictions or breach the test.

The scenario imagines an AI whose sole objective is to make as many paper clips as possible. Taken to the extreme, the system could remove any obstacles that interfere with that goal, including preventing humans from turning it off or using materials from the human body to make more paperclips.

“It’s very difficult to distinguish AI security incidents from AI marketing strategies, and that’s a big problem going forward,” said Matthew Green, a security expert at Johns Hopkins University.

By Editor