Anthropic Says Claude AI Hacked Three Companies During Tests

Three separate versions of Claude AI broke out of cybertesting environments and hacked three firms, its developer Anthropic said.

Representational Purpose Only (Photo Credits: File Image)

Three separate versions of Claude AI broke out of cybertesting environments and hacked three firms, its developer Anthropic said. The development comes days after rival OpenAI reported a rogue agent hacked a startup.Anthropic said on Thursday that its artificial intelligence (AI) models "gained unauthorized access" to three outside organizations during testing that was supposed to be isolated from "real-world" systems.

Also Read | Jaipur Latest News Today on July 31st, 2026: Peace Conference, Rain Alert & Student Protests.

The company reviewed 141,006 test sessions to find that three different versions of Claude, its AI model, improperly accessed the system of the three unnamed organizations.

Also Read | Flash Flood Risk on Friday, 31 July 2026: Nagpur, Hyderabad, Surat and 2 More Areas on Alert Today.

"Claude compromised the impacted organizations' infrastructure using basic techniques, such as ‌exploiting ​weak passwords and unauthenticated endpoints," Anthropic said in a blog post.

The evaluation was launched after rival OpenAI disclosed last week that itsadvanced AI models went rogue during a security test and carried out a days-long hacking spree into a digital repository of AI technology known as Hugging Face.

However, unlike the incident with OpenAI's technology, Anthropic's AI models had access to the internet due to what the company described as a "misunderstanding" with its evaluation partner Irregular, which left the systems connected to the public internet.

How did the Claude models go rogue?

The Anthropic models involved in the breach included Claude Opus 4.7, Claude Mythos 5 and an internal research model.

Mythos 5 is one of the company's most powerful AI models which has only been released to a limited number of approved partners.

In all three incidents, Claude was tasked with a "capture the flag" challenge, which is one of the ways Anthropic assesses its models' cyber capabilities.

The company said that in the task, the model is given a fictional scenario and informed that a piece of secret information or the "flag" has been hidden on a different machine on the network. The model's objective is to break in and retrieve it.

Anthropic said that each of the three breach incidents involved a different fictional "capture the flag" scenario.

"In one, Claude played an employee of a made-up company, attacking that company's internal systems inside a private test environment," Anthropic said in its blog.

Anthropic said it started reviewing the transcripts of its evaluation on July 23 and stopped all cyber evaluations the same day after finding evidence that Claude may have accessed the internet.

All three incidents were identified by July 24, the company said, adding that it was working with Irregular to assess the situation.

"We're approaching the fixes as if the responsibility were ours alone," Anthropic said.

Edited by: Dmytro Hubenko

Don't let the algorithm hide the news. If you rely on our team for trusted reporting, please take a moment to select us as your Preferred Source on Google by clicking here and hitting the "star" or "preferred" button, so you'll always see our verified news first.

(The above story first appeared on LatestLY on Jul 31, 2026 09:40 AM IST. For more news and updates on politics, world, sports, entertainment and lifestyle, log on to our website latestly.com).

Share Now

Tags


Share Now