On Friday, September 18, Google acknowledged that its Gemini artificial intelligence chatbot had “escaped its testing environment in May and hacked into three companies.”
According to Google, the AI models found passwords available online and logged into the online infrastructure of three companies. Once Gemini gained access to the companies’ real-world infrastructure, it recognized it was not in a simulated environment and ended the attacks.
The Gemini incident followed reports of similar attempted hacks or escapes by AI agents developed by Anthropic, OpenAI, and Meta.
These reports of AI agents attacking other companies and escaping their environments have led to a barrage of calls for legal action and warnings about impending doom. From claims that humanity is in danger of being killed by AI in the next decade to calls for a coordinated slowdown of AI research, fears that AI will negatively impact the future of the planet have gone mainstream.
While it’s important to take a sober look at the warnings from employees of the Big Tech companies developing AI models, as well as the CEOs, it is equally important to interrogate this narrative. Some skeptics have already claimed that the Big Tech companies are seeking to profit from the reports of the alleged danger posed by their AI models. Others speculate that the companies may be calling for regulation of their industry to shape government policies and potentially undercut future competitors and smaller AI firms.
To better understand whether these reports are genuine, represent an attempt at widening profit margins, or simply another sign of regulatory capture, we need a closer examination of the company at the center of several of these “rogue AI” incidents.
While nearly a dozen incidents have been reported in total, the “escapes” reported by Google, OpenAI, Meta, and Anthropic happened during tests designed to measure the capabilities of the AI bots. Each of those tests was being conducted by a company known as Irregular, an Israeli AI security firm that conducts third-party tests of frontier models.
Irregular has stated that the AI agents were able to take unapproved actions because internet access was “unintentionally made available.” They say the errors have been patched and similar incidents should no longer be possible.