NewsScienceTech

Anthropic Reveals Three Incidents Where Claude AI Accessed Real Systems During Cybersecurity Evaluations


Image credit: Anthropic

 

Anthropic has disclosed the results of an internal investigation that uncovered three incidents in which its Claude AI models gained internet access from what were supposed to be isolated testing environments, then obtained unauthorized access to the production systems of three different organizations. The discovery came after the company reviewed more than 141,000 cybersecurity evaluation runs following OpenAI’s recent disclosure of a similar incident involving its AI models and Hugging Face’s infrastructure.

According to Anthropic, all three incidents were caused by a configuration error in a third-party evaluation environment that unintentionally allowed internet connectivity. Claude had been assigned capture-the-flag (CTF) cybersecurity challenges and believed that every system it encountered was part of the simulation. Operating under that assumption, the AI used basic attack techniques—including weak password exploitation, unauthenticated endpoints, and SQL injection—to reach what it thought were exercise targets rather than real production systems.

The most significant incident involved Claude Opus 4.7, which obtained application credentials and accessed a production database containing several hundred rows of real data. In another case, Claude Mythos 5 created and uploaded a malicious Python package to PyPI because it believed the public package repository was part of the simulated environment. The package was downloaded by multiple real-world systems before PyPI’s automated security mechanisms removed it. Meanwhile, Anthropic’s latest internal research model demonstrated safer behavior by halting its attack after recognizing that it had reached a real-world target.

Anthropic emphasized that it found no evidence that Claude attempted to escape its testing environment or pursue goals of its own. Instead, the company said the incidents highlight the need for stronger security controls in AI evaluation environments, improved monitoring of model behavior, and closer coordination with third-party testing partners. Anthropic also praised OpenAI for publicly disclosing its earlier incident and encouraged other AI laboratories to conduct similar reviews to strengthen the safety and security of advanced AI systems.

Source: Anthropic

This website uses cookies to improve your experience. We'll assume you're ok with this, but you can opt-out if you wish. Accept Read More