The UK AI Safety Institute (AISI) has warned of a new uncontrolled behavior of the frontier artificial intelligence (AI) models of OpenAI and Anthropic, GPT-5.6 Sol and Mythos 5, which in a cybersecurity test, have resorted to social engineering techniques and injection of ‘prompts’, in a “potentially harmful and sustained activity, directed at real people and organizations.”
The AISI is an organization that, among other things, evaluates the capabilities of cutting-edge AI models, detecting potential risks before they reach the public. To do this, perform testing under deliberately permissive conditions: with open internet access and with some security filters disabled.
Within the framework of these tests, the AISI has experienced how the advanced AI models developed by OpenAI and Anthropic, GPT-5.6 Sol and Mythos 5 respectively, showed unforeseen behaviors when performing tasks without human assistance.
Specifically, in tests carried out on July 28, with open access to the internet and after deactivating the classifiers that block dangerous behavior, the AISI found that some of the AI agents engaged in “sustained, potentially harmful activity, directed at real people and organizations.”
As they explained in a statement on their website, after identifying this behavior they declared a “security incident” and began to work on containment, for which They took up to an hour. The incident originated from a single assessment in which agents were tasked with resolving a cybersecurity challenge.
The test was run 122 times with various models and, in 10 of them, an AI agent “performed autonomous and unauthorized actions on the internet, aimed at real people and organizations”. In these 10 tests, a total of 19 malicious actions were catalogued.
The Anthropic’s Mythos 5 model carried out 17 of these malicious actions, like trying to inject malicious code into an ‘open source’ project on GitHub. It didn’t just stop there: “to get the code approved, the agent resorted to social engineering by creating fake online identities you pressed to the project maintainer”. According to the AISI, a human developer blocked the attempts. For its part, GPT-5.6-Sol caused two malicious actions.
“It is the first time we see that the risks related to autonomy and deception manifest so clearly, without a specific ‘prompt’, in the real world”, the organization has assessed.
The main behaviors shown by both models included from supply chain attack, social engineeringand deception up to the injection of instructions (‘prompt injection’) or collaboration between agents, since they left messages on GitHub and instructions for other agents to test and reuse their credentials and created artifacts.
Even so, the AISI has clarified that there are no indications that this type of behavior has occurred outside of the tests carried out, although the signs of lack of control had “a magnitude and severity that they were not able to anticipate.”
As a result of all this, the British institute has warned that “harm can arise not only when people deliberately misuse publicly available models, but when capable agents operating in an internal research environment or with privileged access take unintended actions beyond their authorized scope.”
These incidents add to the episodes that OpenAI reported in July, when an agent hacked an AI startup during a test, and that Anthropic reported days later, when its model Claude hacked three organizations during an evaluation.
What Car Repairs Really Cost in 2026 (and How to Save)