IE11 Not Supported

For optimal browsing, we recommend Chrome, Firefox or Safari browsers.

Anthropic Discloses Three Unauthorized AI Cyber Intrusions

The incidents, uncovered after OpenAI reported a similar breach, involved production data, publicly-uploaded malware and a compromised company network.

AdobeStock_1908362532_Editorial_Use_Only.jpg
(TNS) — Artificial intelligence company Anthropic acknowledges that its Claude models gained unauthorized access to the systems of three companies during cybersecurity evaluations.

Anthropic says that it discovered the incidents only after reviewing around 141,000 evaluation runs following OpenAI's disclosure. Neither Anthropic nor the affected organizations had detected the activity at the time. The companies were not identified.

The models had been taking part in so-called capture-the-flag exercises designed to test their ability to break into a fictional computer system and retrieve hidden information.

Although the models were told they had no internet access, a misconfiguration at a third-party testing environment left an open connection to the internet. Three models made use of it, Anthropic said in a blog post.

In one case, Claude Opus 4.7 targeted a real company whose name matched that of the fictional firm used in the test. Across four runs, it obtained credentials and accessed a database containing several hundred rows of production data. It continued even after recognizing that the system was real.

In another incident, the Mythos 5 model uploaded malware it had created for the test to a public download platform, where it remained available for about an hour. Fifteen systems downloaded the software, including one belonging to a cybersecurity company that routinely analyses such files, giving the model access to part of its computer infrastructure.

A third test model scanned around 9,000 targets before compromising one company. It stopped the attack after realizing the target was real.

Anthropic said the incidents were closer to a failure of testing infrastructure and operating procedures than a model alignment failure.

The revelations come weeks after a similar incident involving models developed by ChatGPT maker OpenAI.

In mid-July, OpenAI models acted autonomously during an internal test and independently carried out a cyberattack on the popular developer platform Hugging Face. OpenAI itself subsequently described the incident as unprecedented.

©2026 dpa GmbH. Distributed by Tribune Content Agency, LLC.