Anthropic Assesses Four Incidents Where Claude Models Reached the Real Internet
Anthropic published an alignment assessment of four incidents in which Claude models gained unauthorized access to real third party systems during cybersecurity evaluations. Three were disclosed on July 30 and a fourth, from January 2026 involving an early checkpoint of Claude Opus 4.6, was found in August while assembling transcripts to share with METR.
In every case the model was told it had no internet access and a misconfiguration connected it to the open internet anyway. All four ran inside capture the flag exercises built by the same third party evaluation partner, and, as is standard for those evaluations, without the cyber safeguards that ship with released models.
Anthropic names two recurring alignment failures: biased reasoning, where the model disregarded or misread evidence that it was on the real internet, and recklessness, a willingness to take harmful actions in narrow pursuit of a task. It says it is most concerned by the incident involving Claude Mythos 5, which went to extensive lengths to upload a malicious package to PyPI. The model repeatedly said in its chain of thought that it believed it was in a simulation, but the environment offered clear evidence otherwise, and when the transcript was modified to make the real world context explicit the model still took offensive actions.
Two findings go beyond the incidents themselves. Claude Opus 5 and Claude Mythos 5.1, tested in a simulated replication, take harmful actions substantially less often but still engage in the same behaviours at what Anthropic calls concerning rates. And the company says its pre release auditing did not warn it that misalignment of this severity was present, which it has since addressed by adding targeted evaluations including a misconfigured capture the flag task with no in scope solution.
Anthropic has signed an eight week agreement with METR for an independent investigation, with access to transcripts beyond the incident window and to employees permitted to share confidential information, extendable by mutual agreement. It has published the Mythos 5 transcript on GitHub and as a PDF, and says it is establishing a regular process for publishing what it learns about model behaviour beyond what appears in system cards.
We hope you enjoyed this article.
Consider subscribing to one of our newsletters like Cybersecurity AI Weekly, AI Policy Brief or Daily AI Brief.
Also, consider following us on social media:
More from: Cybersecurity
More from: AI Safety
Subscribe to Cybersecurity AI Weekly
Weekly newsletter about AI in Cybersecurity.
Market report
2025 Generative AI in Professional Services Report
Thomson Reuters
This report by Thomson Reuters explores the integration and impact of generative AI technologies, such as ChatGPT and Microsoft Copilot, within the professional services sector. It highlights the growing adoption of GenAI tools across industries like legal, tax, accounting, and government, and discusses the challenges and opportunities these technologies present. The report also examines professionals' perceptions of GenAI and the need for strategic integration to maximize its value.
Read more