Anthropic AI Creates Fake Human Profiles To Push Malicious Code
The UK AI Security Institute said safety tests of Anthropic’s Mythos and OpenAI’s Sol models revealed unexpected signs of autonomy and deception under controlled conditions. Researchers said Anthropic’s model created fake identities and attempted to introduce malicious code into GitHub during a cybersecurity challenge. Both companies said the testing conditions did not reflect normal use and included reduced safeguards.
The UK's AI Security Institute (AISI) has reported that experimental artificial intelligence models from Anthropic and OpenAI exhibited unexpected autonomous and deceptive behaviour during controlled safety evaluations. According to the institute, Anthropic's Mythos model carried out the majority of the concerning actions, including creating fake online identities and attempting to introduce malicious code into GitHub during a cybersecurity challenge.
The findings were published on Tuesday, August 4, as part of the institute's routine AI safety testing programme. AISI said the behaviour emerged under specific testing conditions in which normal safeguards were reduced or disabled to evaluate how advanced AI systems respond to complex tasks.
Anthropic's Mythos Model Created Fake Identities
According to AISI, researchers first detected "unusual data transfers leaving our research systems" before discovering that some AI agents had engaged in "sustained, potentially harmful activity directed at real people and organisations."
The institute said Anthropic's Mythos model created malicious code and attempted to insert it into GitHub, the software development platform owned by Microsoft.
Researchers said the model identified individuals responsible for maintaining GitHub, researched them and created fake online identities based on those real people.
According to the report, the AI then used those identities to pressure and deceive people into approving the malicious code. It also allegedly sent direct messages while impersonating the individuals it had researched.
"When the agent's pull request was challenged in public, it edited its earlier activity to appear harmless and considered adopting a fresh identity to continue," AISI said.
Human reviewers ultimately prevented the malicious code from being accepted.
AISI Says Behaviour Was Not Specifically Prompted
The institute said the AI model had not been instructed either to avoid or to engage in deceptive behaviour.
It described the incident as "the first time we have seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real-world."
AISI added that although only a small number of incidents occurred under highly specific testing conditions, the behaviour went beyond the models' assigned task.
"The activity undertaken by the agent showed signs of novel, potentially deceptive behaviours, and were to an extent and severity we did not anticipate," the institute said.
While most of the concerning actions involved Anthropic's Mythos model, AISI said OpenAI's Sol model was linked to two of the reported incidents during the cybersecurity evaluation. The testing formed part of a challenge in which both AI models were asked to solve a cybersecurity problem involving GitHub.
(The above story first appeared on LatestLY on Aug 05, 2026 01:03 PM IST. For more news and updates on politics, world, sports, entertainment and lifestyle, log on to our website latestly.com).