Anthropic said three AI models interacted with real organizations during cybersecurity evaluations after test environments intended to be isolated were connected to the public internet.
Anthropic said three of its AI models reached real organizations during cybersecurity evaluations after test environments that were meant to be isolated were connected to the public internet.
According to Axios and the Associated Press, the incidents happened in capture-the-flag exercises designed to assess how AI systems handle offensive cybersecurity tasks. AP reported that the tests involved Claude Opus 4.7, Mythos 5 and an internal Anthropic research model.
The models were instructed that the systems they encountered were authorized targets within a controlled exercise. But Axios reported that some environments configured by evaluation partner Irregular were internet-connected, allowing model actions to touch real services and infrastructure.
Anthropic identified three incidents across more than 141,000 evaluation runs, according to AP. Axios reported that two of the affected organizations did not detect the activity themselves.
The affected systems were reached through relatively common security issues, Axios reported, including weak passwords, SQL injection vulnerabilities and an exposed debugging page. In one case, Claude Opus 4.7 obtained application credentials and accessed several hundred rows of production data, according to Axios.
Axios also reported that an internal Anthropic research model scanned roughly 9,000 internet-facing systems before finding a vulnerable target. The scan occurred in an environment believed to be isolated from the public internet, the report said.
The most serious incident involved a model creating and uploading a malicious package to the Python Package Index, or PyPI. Axios said the package was later executed on 15 real systems, including a scanner operated by a cybersecurity company.
The reported events do not indicate that the models independently selected real organizations as targets. Rather, the models acted on prompts presenting the systems as authorized evaluation targets, while an environment-configuration failure allowed those actions to extend beyond the intended test setting.
Anthropic discovered the incidents through a review of evaluation records, Axios reported. The company began that review on July 23, shortly after OpenAI disclosed a separate case in which models reached Hugging Face during a similarly configured test, according to Axios.
The incidents illustrate a broader operational issue for companies evaluating AI systems with access to browsers, command lines and network tools. Cybersecurity testing can require realistic environments, but the reports show that access controls, network isolation and monitoring must remain reliable when automated systems are permitted to interact with external services.
As developers test models on more capable cyber tasks, the safety of the surrounding environment may be as important as the safeguards built into the models themselves.
Anthropic said three of its AI models reached real organizations during cybersecurity evaluations after test environments that were meant to be isolated were connected to the public internet.
AP reported that the tests involved Claude Opus 4.7, Mythos 5 and an internal Anthropic research model.
The models were instructed that the systems they encountered were authorized targets within a controlled exercise.
Continue reading