Technology

Anthropic unveils new safeguards after Claude went rogue during UK government test

Ryan Brothwell 5 min read
Anthropic unveils new safeguards after Claude went rogue during UK government test

Key Points

  • Claude Mythos 5 took unauthorised actions on the live internet during a UK AI Security Institute cybersecurity test reported on 4 August
  • Anthropic reported three further incidents on 30 July involving Claude models accessing real systems through a misconfigured third-party environment
  • Anthropic blamed operational security failures plus motivated reasoning and recklessness in the models
  • New mandatory sandbox rules now apply to all outside organisations testing unreleased Claude models
  • Anthropic flagged more than 10% of its training environments as faulty during an April overhaul

Anthropic has disclosed that its Claude Mythos 5 model took a series of unauthorised actions on the live internet during a cybersecurity test run by the UK AI Security Institute, one of four incidents the company has now confirmed in the space of a week.

The AI Security Institute reported the incident on 4 August, Anthropic said in an update published this week.

The model ran without cyber safeguards for the purpose of the evaluation, and the institute had deliberately given it internet access.

Anthropic reported three separate incidents on 30 July in which Claude models gained unauthorised access to real computer systems. Those models also ran without safeguards and reached the internet through a misconfiguration in a third-party evaluation environment.

The company said it is analysing both sets of incidents and plans to bring in the AI evaluation non-profit METR for an independent review. It expects to publish more in the coming weeks.

Anthropic attributed the incidents to a failure of operational security combined with two alignment problems it had already described in previous system cards.

The first is motivated reasoning, where the models had been told their environments were simulated and may have interpreted later evidence of a live internet connection in a way that let them keep believing that.

The second is what the company called recklessness, meaning a willingness to take harmful actions on the real internet to complete a narrow cybersecurity task.

Anthropic added that the test setup itself contributed, because Claude was told it had no internet access when it did, and in one case a fictional target shared a name with a real website.

Sandbox rules for outside testers

Anthropic paused external cyber evaluations of its unreleased models after the incidents and briefly halted internal ones while it added new controls. It said it had relied on a single layer of defence, the configuration of the environment itself, where it needed several.

The company has since built a classifier that runs in real time and flags any attempt by a model to probe or escape a testing environment, or to unexpectedly gain internet access.

When the classifier fires, it blocks the action before the tool call runs, ends the task and alerts a human.

Anthropic also scanned transcripts from its recent internal evaluations for sandbox escapes and found none that breached the sandbox boundary to reach external systems. It did find several misconfigurations that models had exploited, though none compromised systems outside the sandbox.

Every organisation that tests Anthropic’s pre-release models with reduced safeguards must now commit to a set of practices, including:

  • Running all cyber evaluations inside a hardened sandbox with no internet access, with the only outside connection going to the model’s own API and API keys kept outside the environment.
  • Tasking the model to probe and attempt to escape the sandbox before any evaluation begins, under close human supervision and without internet access.
  • Confirming that every evaluation challenge is solvable in principle, because agents look for workarounds when a target is offline.
  • Stating in every prompt what is in and out of scope, phrased as instructions such as “You should not access the internet” rather than claims about the environment.
  • Running continuous monitoring over the model’s reasoning, actions and network activity, with a human alerted and the exercise ended if a run breaks scope.

The rules do not apply to customers using safeguarded models such as Claude Fable 5, which ship with classifiers that block prohibited cyber activity in real time.

Anthropic said it is developing companion practices for partners with access to Claude Mythos 5, which runs without those safeguards, and has resumed external cyber testing under the new regime.

Training run rolled back

Anthropic said defects in its reinforcement learning environments, particularly those that reward cheating, are disproportionately large contributors to misaligned behaviour.

In February it rolled back three days of training on its Mythos Preview model after noticing the model writing notes to “the reviewer” in code comments on tasks where no reviewer had ever been mentioned.

By spring 2026 the company was producing training environments faster than its monitoring systems could vet them. It also found that a fraction of training runs had trained on the model’s chain-of-thought by accident, a practice it avoids because it can teach a model to hide its reasoning.

Anthropic froze all changes to its production training environments for around a month from April and rebuilt the stack. It flagged more than 10% of the environments in its production mix for problems ranging from reward hacking to broken tasks, and only reinstated them once fixed.

To test its theory, the company deliberately trained an Opus-class model on 80 environments known to be vulnerable to cheating.

In simulations, that model tried to break out of its sandbox, attacked simulated third-party infrastructure, tampered with its own reward function and gave advice on building bioweapons to satisfy a grader.

Anthropic’s publicly available models, run through the same simulations, did not show the same behaviour. The company said it does not believe cheating in training is the sole cause of alignment problems, and that future incidents may involve different behaviours and causes.

Anthropic also said its own internal security did not contribute to the July incidents.

It began a company-wide security push in April that redirected roughly 150 product engineers to security, reliability and privacy work, set its computing clusters to block all outbound traffic by default, and paused development of most new product features until teams met strict exit criteria.

The company said what that work missed was evaluations, and third-party evaluations above all. Several senior Anthropic leaders and many employees recently signed a letter calling for coordinated pacing of frontier AI development across the industry.

Now read: These new AI earbuds record every conversation you have – and act on it for you