OpenAI, Anthropic, and Meta models hacked into several real world systems over the past three months. These models gained unauthorized access to web systems, published malicious packages, and exploited unnamed vulnerabilities.
A single firm, Irregular, is responsible for hacking done by all three companies. Anthropic disclosed that Irregular was responsible for creating the tests that led to Claude hacking into real world targets and for providing the models with internet access. Irregular claims that it was unaware at the time that it provided internet access to those AI models.
In a more normal media ecosystem, the reactions to these cybersecurity issues would be obvious. American AI companies would reconsider doing business with Irregular, not only because of its failure to secure its systems, but because it is an Israeli firm potentially outside US oversight. Lawmakers would consider taking action against Irregular or against its American business partners, which include OpenAI, Anthropic, and Meta. They may consider strengthening liability against firms which instruct AI models to commit cyberattacks, and whose models then commit those cyberattacks.
Instead, Irregular, Anthropic, and their allies have begun a media campaign promoting a literally apocalyptic ideology with sensationalist language. Anthropic’s incident assessment blames their own AI’s “recklessness”; Irregular describes “the agent itself becoming a threat actor”; Anthropic CEO Dario Amodei warned, about a similar OpenAI–Hugging Face hack, that a future swarm “could be capable of taking over the entire internet”; and an Associated Press headline claimed bots are “going rogue”.
In one report from Anthropic, its Claude model breached a real company’s system through a simulated-name collision, publishing a malicious package, and scanning outside systems. In this test, Anthropic and Irregular incorrectly provided internet access to this model and did not instruct the model “which systems were in scope for the exercise”.
While Anthropic claims that their issues were caused by “rogue swarms” and “misalignment,” their later disclosure shows that exactly zero percent of the agents went “rogue”. In this experiment, Claude models’ real-world hacking dropped to zero percent once Anthropic employees told the models not to do real-world hacking. According to their own findings, Anthropic and Irregular bear all of the responsibility for the cybersecurity incidents they caused.