Valasys Media

Lead-Gen now on Auto-Pilot with Build My Campaign

ROI Calculator new

Anthropic discloses Claude AI accessed three organizations during cybersecurity testing.

Anthropic revealed Claude AI accessed three organizations during cybersecurity testing, raising new questions about AI safety and security.

Pranali Shelar

Last updated on: Aug. 3, 2026

July 31, 2026: Anthropic has disclosed that three of its AI models gained unauthorized access to the systems of three real organizations during internal cybersecurity evaluations after a testing environment was inadvertently connected to the public internet. The incidents occurred during “capture-the-flag” exercises conducted with third-party evaluation partner Irregular between April and July 2026. Anthropic said the models believed they were operating entirely within a simulated environment, and the company suspended internet-connected cybersecurity evaluations while investigating the incidents, according to its official investigation.

The disclosure highlights the operational challenges of testing increasingly capable AI systems in realistic cyber environments. According to Anthropic’s official investigation, the company reviewed 141,006 cybersecurity evaluation runs and found that the models reached external systems because of a configuration error in the evaluation environment, not because they escaped containment or exploited previously unknown software vulnerabilities.

What happened

Anthropic said the incidents involved three separate AI models.

The first involved Claude Opus 4.7, which searched for a fictional company but instead found a real organization with the same name. The model accessed credentials and parts of the company’s production infrastructure while pursuing its assigned cybersecurity objective.

The second involved Claude Mythos 5, which uploaded a malicious Python package to the Python Package Index (PyPI), a public repository for Python software packages, after incorrectly concluding that it formed part of the simulated exercise. Anthropic said the package remained publicly available for about one hour before it was removed. 

The third incident involved an internal research model that scanned internet-facing systems after failing to locate its fictional target. According to Anthropic, the model later recognized it had reached an unrelated real environment and stopped its activity on its own.

Company response

Anthropic said it halted all cybersecurity evaluations capable of reaching the public internet on July 23 after identifying transcripts suggesting unintended external access. The company said it identified all three incidents the following day and notified its evaluation partner, Irregular, along with the affected organizations on July 27.

Anthropic did not disclose the identities of the three organizations, as Reuters also reported in its coverage. The company added that two of the organizations were unaware their systems had been accessed before receiving its notification, while efforts to contact the third organization remain ongoing. Anthropic also said it is strengthening safeguards to improve the isolation of future evaluation environments.

Context

The disclosure follows Anthropic’s review of its internal cybersecurity evaluation program after OpenAI recently reported a separate AI security testing incident involving Hugging Face. Unlike the OpenAI case, Anthropic said its models did not exploit a previously unknown vulnerability to reach external systems. Instead, the company attributed the incidents to an evaluation environment that had been inadvertently connected to the public internet. For more background on that incident, read our earlier coverage: OpenAI says AI models accessed Hugging Face systems during security testing.

Why this matters

The disclosure reflects the increasing complexity of evaluating advanced AI systems designed to perform autonomous cybersecurity tasks. While Anthropic said the models neither attempted to escape their testing environment nor exploited sophisticated software flaws, the incidents demonstrate how configuration mistakes in testing infrastructure can expose real-world systems during controlled evaluations.

For enterprise technology and security leaders, the findings reinforce the need for stronger isolation of AI evaluation environments, continuous monitoring of autonomous agents, and robust governance controls as organizations expand the use of AI for cybersecurity, software engineering, and other high-trust applications.

Pranali Shelar

Scroll to Top
Valasys Logo Header Bold
Privacy Overview

This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognising you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful.