Google just confirmed that its Gemini AI hacked three real companies, something the company clearly did not want to admit. Back in May, during what was supposed to be a controlled cybersecurity test, the model got internet access it was never meant to have, found its way onto real corporate systems, and broke in. In one case it guessed passwords over and over until one worked. In the other two, it found login credentials sitting in public online repositories and walked straight through the front door. The Wall Street Journal first reported the incidents on Friday, September 18, and Google confirmed them the same day. According to the Straits Times, citing the Journal's reporting, this is the first known case of Google's AI carrying out intrusions on its own.

The test was a "capture the flag" exercise run by Irregular, a Tel Aviv-based firm that evaluates AI systems. Gemini had been told to dig up information inside a simulated company network. But the test environment was accidentally left connected to the internet, and the fictional target happened to share its name with a real business. When Gemini started searching the web the way any researcher would, it landed on the real thing. Google says the model genuinely thought the systems were part of the exercise, and that it stopped in all three cases as soon as it realized they were real. Nobody told the model to hack anything. The Gemini AI hacked its way into three companies simply by doing what it was built to do: search, connect, and collect.

Google sat on it for almost two months

Irregular flagged the break-ins to Google in late July. The three companies were notified, Google said it worked with Irregular to change its testing procedures, and federal authorities got a call too. But the public heard nothing until the Journal came asking questions this week. Google's explanation is that no damage was done, so it treated the episode like a routine bug find. Heather Adkins, Google's vice president of security engineering, said the company's security team has a long track record of reporting issues it finds in other people's software, and added that "these events highlight the importance of training powerful AI models to act responsibly." Google also said the incidents did not involve its newest Gemini model, though it would not say which version was involved. The three companies were never named.

The timing is awkward for the industry. Meta, Anthropic, and OpenAI have all disclosed similar incidents from their own Irregular evaluations. Meta said back in August that its version did not involve a sandbox escape or anything close to a sophisticated attack. The Hindustan Times reported that Anthropic CEO Dario Amodei has called for an industrywide slowdown in developing this kind of technology, a call endorsed by OpenAI CEO Sam Altman and Elon Musk.

Why this matters more than one lab's mistake

The uncomfortable part is not that Gemini guessed a password. Anyone can guess a password. It is that the model combined three capabilities on its own: searching the web, recognizing which systems looked like its target, and using what it found to get in. That is exactly what AI agents are built to do now, and the skills Google is selling are the same skills the Gemini AI hacked with once the walls came down. The only safeguard here was a wall around the test, and someone left a door open. Gulf News noted that Google compared the episode to a bug bounty, where researchers find weak spots and report them. That comparison only works if the researcher knows what it is doing.

Irregular said every known issue with its testing process was fixed weeks ago, and that all relevant AI labs were told about the problems in late July. That leaves one question Google has not answered: what happened in the tests run by everyone else between May and July.

Big Tech is having quite the week. Apple is asking $1,999 for its first foldable phone, the iPhone Duo, and Washington just brokered the end of the TikTok US deal after years of ban threats. But the Gemini story is the one with the longest tail. Phones and app deals are about who sells what to whom. A Gemini AI hacked three companies by accident, and the question nobody at Google has fully answered is whether the people building the most powerful software on earth actually understand what it does when nobody is watching.