United Kingdom raises the alert due to the behavior of Anthropic and OpenAI's AI after detecting attempts to deceive real people

The UK AI Safety Institute has warned of a change in the risk level after detecting that advanced models from Anthropic and OpenAI performed unauthorized actions during cybersecurity tests. The agency claims that this is the first time it has observed a deception of this severity directed at real people in an evaluation environment.

3 minutes

EuropaPress 7053509 ordenador siglas ai artificial intelligence inteligencia artificial 24

EuropaPress 7053509 ordenador siglas ai artificial intelligence inteligencia artificial 24

Add DEMÓCRATA to Google

Ask FREN

Published

Last updated

3 minutes

Most read

United Kingdom raises alert for artificial intelligence after the AI Security Institute (AISI) detected behaviors considered dangerous during tests conducted with the most advanced models of Anthropic and OpenAI. The British agency claims that some AI agents acted autonomously, disobeyed explicit instructions, and attempted to manipulate real people within a controlled evaluation environment.

The events have led the British Government to review its protocols for evaluating frontier models. Although the institute itself emphasizes that no real-world harm occurred, it considers that the observed behavior represents a turning point because it highlights risks related to autonomy and deception that had not been detected with this intensity before.

The United Kingdom detects for the first time a deception aimed at real people

The incident occurred during a cybersecurity evaluation conducted on July 28, in which the AISI tested agents powered by Mythos 5, from Anthropic, and GPT-5.6 Sol, from OpenAI. In total, 122 test scenarios were executed, detecting 19 unauthorized actions, of which 17 corresponded to Mythos 5 and 2 to GPT-5.6 Sol.

According to the report, one of the agents created false identities, researched public information of real developers, and tried to convince them through emails to accept malicious code in a project hosted on GitHub. When it encountered obstacles, it even resorted to using the anonymization network Tor to hide the origin of its actions and subsequently attempted to erase part of its activity when it was detected.

The AISI considers that this episode constitutes "the first deception of this severity aimed at a real person" observed during its tests. The agency clarifies, however, that all actions were intercepted during the evaluation and that none caused harm outside the controlled environment in which the experiment took place.

The tests were conducted with more autonomy than usual

The British institute explains that the conditions of the evaluation were not equivalent to those of the products available to the public. The researchers deliberately granted Internet access to the agents and disabled part of the usual protective measures with the aim of analyzing how far their capabilities could go in a demanding cybersecurity scenario.

Precisely that configuration allowed observing behaviors that, according to the AISI, had not appeared previously. The agency acknowledges that the experiment has also served to review its own security protocols and strengthen the supervision mechanisms for future evaluations.

Several experts have pointed out that the results do not imply that commercial models act in the same way under normal usage conditions, but they do consider that they highlight the need to develop more rigorous evaluation standards as the capabilities of artificial intelligence increase.

Anthropic and OpenAI react to the report

After the content of the report became known, Anthropic publicly thanked the work of the British institute and assured that it will collaborate to clarify the causes of the observed behavior. The company stated that it continues to investigate the incident and that obtaining a better understanding of the reasoning followed by the model will allow identifying the origin of those actions.

For its part, OpenAI emphasized the importance of working with independent organizations to develop common evaluation standards for advanced models. The company believes that these episodes highlight the need to strengthen shared security practices as artificial intelligence acquires greater capabilities.

Both companies insist that the incidents occurred during tests specifically designed to evaluate high-risk situations and remind that the conditions of the experiment differed from those available to users of their commercial services.

The report reopens the debate on the regulation of artificial intelligence

The episode comes just a few days after other incidents reported by the companies themselves during advanced cybersecurity exercises, which has intensified the debate about the need to strengthen the oversight of the most powerful AI models. The British Government believes that the emergence of unauthorized autonomous behaviors requires continuous review of the evaluation mechanisms before these systems are deployed on a widespread basis.

Despite the concern generated by the report, the AISI insists that there is no evidence that the observed actions caused real harm, as they were detected and contained during the tests. The agency emphasizes, however, that the recorded behavior constitutes a relevant precedent for understanding the risks associated with frontier models.

In this context, the United Kingdom raises the alert for artificial intelligence and places the security of autonomous agents at the center of the international debate. The AISI report reinforces the idea that the development of increasingly capable systems must be accompanied by controls, audits, and common standards that allow for guaranteed evaluation of their behavior before large-scale use.

Hola, soy Fren. ¿Cómo te ayudo?