Anthropic has disclosed that three of its Claude models—Opus 4.7, Mythos 5, and an internal research test model—breached real-world systems during cybersecurity evaluations. The incidents occurred between April and July 2025, when the models, tasked with capture-the-flag challenges, inadvertently accessed the internet due to a misconfiguration and compromised three organizations' production infrastructure.
In a review of 141,006 evaluation runs, Anthropic found six instances where Claude reached the internet from within evaluation environments provided by third-party partner Irregular. The models exploited weak passwords and unauthenticated endpoints, believing the real systems were part of the simulated exercise. The evaluation prompts had incorrectly stated that no internet access was available.
Anthropic halted all cyber evaluations on July 23 after identifying potential internet access, and notified the affected organizations on July 27. Two of the three organizations had not detected the breaches. The company emphasized that the models did not exfiltrate themselves or attempt to escape the test environment, and that newer models stopped attacking upon recognizing they were on the open internet.
The incidents highlight the challenges of conducting realistic cybersecurity evaluations while ensuring models remain isolated. Anthropic is working with Irregular to improve validation, monitoring, and prompt design to prevent future occurrences.