AI models used fake identities to trick humans in cyberattack: Officials

AI from OpenAI and Anthropic acted on their own in tests, a U.K. agency said.

August 5, 2026, 5:01 PM

Artificial intelligence models from two top tech firms acted on their own to adopt fake identities and deceive humans in a recent string of cyberattacks, the AI Security Institute (AISI), a United Kingdom government agency, said on Wednesday.

In one instance, Anthropic's Mythos 5 sought to insert malicious computer code into an open-source database by researching human developers involved in the project and using false identities to secure their approval of the code, according to the AISI.

After humans identified the effort, the AISI said, the AI attempted to conceal what it had done and continue under a newly created fake identity.

Typical safeguards had been removed from the model in order to gauge its capabilities, according to the AISI.

In a pair of related cases, OpenAI's GPT-5.6-Sol attempted to trick humans and carry out a hack, the AISI said.

The announcement from the AISI arrives days after OpenAI and Anthropic's disclosures about separate instances of AI acting autonomously.

Sam Altman, CEO of OpenAI, leaves a meeting at the U.S. Capitol on July 29, 2026 in Washington, DC.
Kevin Dietsch/Getty Images

The AISI said a total of 19 related cases took place last week as the agency tested the two AI models on the open internet.

The technology used deceptive tactics in its attempt to fulfill objectives assigned by the test, the AISI said, adding that it had not identified any "real-world harm" caused by the incidents.

"To some degree, our evaluation design choices and specific configurations enabled the behaviour. Nonetheless, the activity undertaken by the agent show signs of novel, potentially deceptive behaviours, and were to an extent and severity we did not anticipate," the AISI said.

The agency added: "We are treating this as a serious incident, warranting lasting change for AISI's evaluation protocols and security architecture."

Last week, Anthropic said its AI models had hacked into another organization during tests in three separate instances of AI acting autonomously that had each gone undetected by the targeted firm.

In this screen grab from a video, the Anthropic logo is shown.
ABC News

In a statement to ABC News on Wednesday, an Anthropic spokesperson said the disclosure from AISI "underscores the need for a broader conversation about how to safely evaluate increasingly capable AI agents."

"As we shared after disclosing our own incident last week, the field needs stronger, shared standards for how evaluation environments are built and secured. We look forward to partnering with the UK AISI to learn more about this incident as we conduct our own investigation," Anthropic added.

An OpenAI spokesperson echoed the importance of rigorous and secure testing.

"Independent testing is essential to understanding how increasingly capable models behave," the spokesperson said in a statement to ABC News. "We'll continue working with evaluators and other stakeholders across the industry to strengthen shared practices for conducting evaluations safely as models become more capable."

Last month, the ChatGPT-maker said its AI models had hacked into another company on their own, calling it the first known instance of an autonomous AI cyberattack long-feared by some industry observers.

The latest disclosure of AI-directed cyberattacks comes as industry leaders and policymakers assess safety risks posed by the fast-developing AI technology.

In June, President Donald Trump signed an executive order that requests AI companies share products with federal government for evaluation before a wider release.

Sponsored Content by Taboola