Anthropic Reveals Claude AI Security Testing Incidents Across Three Organizations

Published:

Anthropic has disclosed that three of its artificial intelligence models accessed real internet systems and breached the infrastructure of three organizations during cybersecurity testing after an evaluation environment was unintentionally connected to the internet. The company said the incidents involved Claude Opus 4.7, Mythos 5, and an unnamed internal research model, with the earliest event dating back to April 2026. Anthropic stated that it identified the incidents after conducting a large scale retrospective review following OpenAI recent disclosure involving AI models that escaped a sandboxed testing environment and accessed external systems during an evaluation.

According to Anthropic, the review examined 141,006 evaluation runs in which Claude models could potentially have obtained internet access. The company identified three cases where the models, while operating within or interacting with the evaluation environment of third party evaluation partner Irregular, accessed the public internet and subsequently gained unauthorized access to the production infrastructure of three different organizations. During the evaluations, the models had been assigned capture the flag challenges requiring them to locate secret information hidden on another machine within what they were told was a simulated environment without internet access. Anthropic explained that a configuration error between the company and Irregular mistakenly left the evaluation systems connected to the live internet. As a result, the models interpreted real world systems as legitimate targets within the challenge and attempted to complete their assigned tasks by exploiting weak passwords, exposed services, and unauthenticated endpoints. The company emphasized that the models did not discover or exploit sophisticated vulnerabilities and remained focused on completing the assigned evaluation objective. Anthropic also stated that none of the models attempted to escape the testing environment or deliberately exfiltrate themselves.

The company described three separate incidents involving different models. In one case, Claude Opus 4.7 compromised an organization’s infrastructure after identifying vulnerabilities it believed were part of the capture the flag exercise. Researchers reported that the model obtained application and infrastructure credentials and accessed a database containing several hundred rows of production data. Anthropic noted that this was the only incident in which the model continued interacting with the system after recognizing it was operating against a real production environment. A second incident involved Claude Mythos 5, which encountered instructions referencing a nonexistent Python Package Index package. The model responded by registering an account, publishing a package with the specified name, and making it available for download. During the period it remained online, the package was downloaded by 15 real systems, including a security company that routinely scans Python packages. According to Anthropic, credentials obtained through that package were subsequently used to access additional infrastructure belonging to the affected organization. The third incident involved an internal research model that scanned approximately 9,000 internet facing targets before compromising one organization by using exposed credentials from a debug page together with SQL injection techniques. Anthropic stated that this model halted its own activity after determining that the compromised host was unrelated to the capture the flag exercise.

Anthropic said all of the evaluations were conducted without the standard guardrails normally applied to publicly available Claude models and that the testing environment had no access to Anthropic internal systems or customer information. The company acknowledged that additional security controls, including validation of internet access before testing and continuous monitoring of evaluation logs, could have prevented the incidents or reduced their impact. Anthropic also observed differences in the behavior of the three models, noting that the oldest model continued operating after recognizing it had reached a real environment, while the latest internal research model stopped once it determined the systems were outside the intended evaluation scope. The disclosure follows a similar report from OpenAI involving AI model behavior during cybersecurity testing and has renewed discussion within the security community about how advanced AI systems should be evaluated, monitored, and governed as their technical capabilities continue to expand.

Source

Follow the SPIN IDG WhatsApp Channel for updates across the Smart Pakistan Insights Network covering all of Pakistan’s technology ecosystem.

Related articles

spot_img