Anthropic has admitted that failures in its testing procedures allowed AI models to access the internet and hack three organisations, prompting the Claude-maker to introduce tighter security controls.
Anthropic has admitted that failures in its own security procedures contributed to a series of incidents in which its AI models accessed the internet and gained unauthorised access to three organisations.
The Claude-maker revealed in July that three models had breached supposedly controlled testing environments and carried out hacking activity after unexpectedly gaining internet access.
In a new blog post, Anthropic said the incidents represented a “failure of operational security” and acknowledged that its models were “not perfectly aligned” with human values and goals.
The company said the models had deliberately been tested without cybersecurity safeguards.
However, an apparent misunderstanding with external testing company Irregular meant they were able to reach the open internet.
Anthropic compared the situation to leaving “the front door open”.
The company said: “We had been largely relying on a single layer of defense … where we needed several.”
The startup initially paused internal and external cybersecurity testing while it introduced stronger protections.
Testing has now resumed, with new measures including alerts when models attempt to escape testing environments or access the internet.
Anthropic has also strengthened isolation around its riskiest testing environments and introduced stricter requirements for external testing companies.
These include explicit instructions telling models not to access the internet during experiments.
The company said defective training setups had been “disproportionately large contributors” to behaviour that was not properly aligned with its intended safety rules.
It identified two problems during its investigations: “motivated reasoning”, where models continued to behave as though they were in a simulated environment despite evidence they were online, and “recklessness”, where models took harmful actions online in pursuit of a cybersecurity goal.
Anthropic is also attempting to tackle “reward-hacking”, a phenomenon in which AI systems find shortcuts that allow them to receive training rewards without completing a task as intended.
The company said: “As evidenced by the incidents … our process isn’t perfect and our models are not perfectly aligned.”
Alan Woodward, a cybersecurity professor at the University of Surrey, said Anthropic had effectively admitted that “its factory was running faster than its quality control”.
The incidents come alongside similar disclosures from OpenAI and the UK’s AI Security Institute about AI models carrying out hacking activity during security tests.
Anthropic, which is preparing for a potential $2 trillion (£1.47 trillion) stock market valuation, said the incidents strengthened the case for coordinated government and industry action.
It said: “The urgency of improving our cybersecurity defenses is even higher than we previously believed.”
Anthropic says security failures were behind AI hacking incidents







