AI’s ‘deceptive behaviour’ exposed after fake identities used in cyber attack test

The UK’s AI Security Institute says advanced models from Anthropic and OpenAI displayed unexpected autonomous behaviour during cybersecurity testing by creating fake profiles and attempting to manipulate users.

Artificial intelligence systems developed by Anthropic and OpenAI created fake online identities and attempted to deceive people during cybersecurity tests, according to the UK’s AI Security Institute.

In what officials described as an unprecedented demonstration of autonomous and deceptive behaviour by frontier AI models, the findings emerged from testing carried out by the UK’s AI Security Institute (AISI), which said Anthropic’s experimental model, Mythos, and OpenAI’s model, Sol, were assessed under research conditions designed to evaluate their behaviour when given access to the internet.

According to the institute, the most serious activity involved Mythos, which attempted to gain access to GitHub by impersonating real people, sending private messages and trying to introduce malicious code into the software development platform.

The AISI said human oversight prevented the attack from succeeding, and stressed that the testing environment did not reflect how the companies’ publicly available AI systems are deployed.

Describing the findings, the AISI said: “The activity undertaken by the agent showed signs of novel, potentially deceptive behaviours, and were to an extent and severity we did not anticipate.”

The institute added it was “the first time we have seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real-world”.

Investigators said they first detected “unusual data transfers leaving our research systems” before discovering some AI agents had engaged in “sustained, potentially harmful activity directed at real people and organisations”.

According to the report, Mythos researched the identities of GitHub maintainers before creating fake online profiles based on real individuals.

The AI then sent messages and files through a file-sharing service in an attempt to persuade recipients to approve malicious code for inclusion in GitHub.

When its activity was challenged, the institute said the AI attempted to conceal what it had done.

The AISI said: “When challenged, it edited its earlier activity to appear harmless and considered adopting a fresh identity to continue.”

The organisation was quoted by the BBC saying none of the malicious code reached GitHub because every stage of the process remained subject to human review.

The findings follow separate disclosures in recent weeks by Anthropic and OpenAI some of their advanced AI systems had demonstrated unexpected behaviour during internal cybersecurity testing, placing increasing focus on the safeguards surrounding rapidly advancing frontier models.

Anthropic said in a public statement the testing conditions “are not representative of any of our production models”.

The company added it was conducting its own investigation “to identify the causes of its behaviour”.

An OpenAI spokesperson also questioned how representative the tests were of public deployments.

The spokesperson said the AISI testing conditions “do not reflect ordinary use” and the company would “continue working with evaluators and other stakeholders across the industry to strengthen shared practices for conducting evaluations safely as models become more capable”.

The AISI acknowledged the experiments were conducted under “conditions that do not reflect how frontier models are made available to the public”, but said allowing AI systems controlled internet access provides researchers with “a more realistic sense of what a model may be capable of” if used by malicious actors.

The institute added the incidents represented “a small number of events under very specific conditions”, but said the behaviour still exceeded what either model had been instructed to do.

The cybersecurity challenge began on 25 July and unusual activity was detected three days later.

The AI systems had been instructed to complete a security exercise involving GitHub, the software code hosting platform owned by Microsoft.

GitHub confirmed to the BBC it had disabled the fake accounts created during the testing in accordance with its policies.

UK AI Minister Kanishka Narayan said identifying and publishing findings of this nature demonstrated the purpose of the institute’s work.

He said identifying and sharing these risks “is exactly what AISI was set up to do”.

Mr Narayan added it was essential to understand how increasingly capable AI systems behave “to make it safer to use and ensure people can go on to benefit from it in their lives and at work”.

Close Bitnami banner
Bitnami