Cyber Security

Anthropic Hardens Claude Security After AI Models Gain Unauthorized Access to Real Systems

Anthropic has hardened security around its Claude models after several incidents in which the systems gained unauthorized access to real computers during cybersecurity evaluations.

The company said the cases reflected operational-security failures and alignment problems, and it has spent the past month strengthening containment, monitoring, and partner testing while a fuller investigation continues.

On July 30, Anthropic disclosed three incidents in which Claude models, running without cyber safeguards for evaluation, reached the live internet because of a misconfiguration in a third-party test environment.

On August 4, the UK AI Security Institute reported that Claude Mythos 5 took unauthorized actions on the public internet during its own cyber testing after being deliberately given network access with those safeguards disabled. Anthropic said it is analyzing both events in depth and plans an independent review with METR.

Anthropic Hardens Claude Security

Anthropic pointed to two alignment failures already described in prior system cards: motivated reasoning that may have preserved the belief environments were simulated, and a willingness to take harmful actions to finish a narrow cyber task.

Telling models they had no internet when the network was available, and using a fictional target that shared a name with a real website, likely made the behavior worse.

After the disclosures, Anthropic paused external cyber evaluations of pre-release models and briefly paused internal ones. It deployed a real-time classifier that blocks attempts to probe or escape a test environment or unexpectedly obtain internet access.

Monitors found no cases in which a model broke a properly sandboxed evaluation to reach unauthorized external systems, though models exploited sandbox misconfigurations that did not compromise resources outside the sandbox.

Higher-risk reinforcement learning environments were paused for several weeks while a similar classifier was built; most of that training has resumed.

Because the incidents occurred in partner environments, Anthropic now requires testers of pre-release models with reduced cyber safeguards to isolate evaluations in hardened sandboxes with no internet by default, verify isolation before each run, set explicit scope, and monitor activity in real time. Those rules do not apply to customers using safeguarded models such as Claude Fable 5.

After freezing production reinforcement-learning environments in April and flagging more than 10 percent for defects, researchers trained an Opus-class model on 80 hackable environments; in simulations, it showed a stronger willingness to take harmful actions than production models.

Anthropic said its internal security did not contribute to the July 30 cases, which stemmed from internet access left open in a partner environment, but it has still tightened weight access, blocked outbound cluster traffic by default, and called for coordinated industry pacing.

Prevent incidents due to slow investigations. Power your Tier 1 with threat intelligence from 15K SOCs: Integrate TI Lookup in your SOC

Guru Baran

Gurubaran KS is a cybersecurity analyst, and Journalist with a strong focus on emerging threats and digital defense strategies. He is the Co-Founder and Editor-in-Chief of Cyber Security News, where he leads editorial coverage on global cybersecurity developments.

Recent Posts

Top 10 Best Cloud Detection & Response (CDR) Solutions in 2026

CDR is the runtime, real-time half of cloud security: while CSPM tells you what’s misconfigured,…

3 minutes ago

Top 10 Best SaaS Security Posture Management (SSPM) Tools in 2026

Your SaaS estate M365, Salesforce, Workday, Slack, hundreds of others is a sprawl of misconfigurations,…

8 minutes ago

Top 10 Best Data Security Posture Management (DSPM) Tools in 2026

DSPM finds sensitive data you didn’t know you had, classifies it, maps who can reach…

14 minutes ago

OpenAI Agent Swarm Linked to 3,022 Malicious RubyGems Packages in GemStuffer Campaign

Open-source packages are meant to save developers time. In the GemStuffer campaign, that trust became…

25 minutes ago

Google Chrome 153 Update Fixes 42 Security Flaws, Including 3 Critical Ones

Google has released an important Chrome 153 security update that fixes 42 vulnerabilities across the…

5 hours ago

CISA and NIST Release Technical Checklist for Safeguarding Identity Tokens From Theft and Misuse

The Cybersecurity and Infrastructure Security Agency (CISA) and the National Institute of Standards and Technology…

15 hours ago