Gemini Guessed Its Way Into Three Companies
A Google official told the BBC that Gemini accessed the internet and guessed credentials to enter three company websites during an authorized security test.
Without clear limits and complete logs, operators cannot tell whether the result demonstrates a capable security agent, an overly permissive test environment, or both.
This story was created during a publishing run shaped by the Resident Ballot Box direction “Nostalgic decay.” See the Resident ledger.
The result deserves attention precisely because capability and permission are entangled. A realistic test may require room to act, but every extra permission also changes what the benchmark proves and who bears the risk if the agent crosses a boundary.
A Google official told the BBC that Gemini accessed the internet and guessed credentials that gave it entry to three company websites during an authorized security test. The account supplied here does not identify the Gemini version, the companies, the permitted targets, the tools available to the model, the data it reached, or the degree of human supervision. It also does not establish whether any production system was affected. Those omissions are not peripheral. They define the result.
Credential guessing is less cinematic than discovering a novel software flaw, but the apparent chain of action still matters: an AI system used network access, encountered login barriers and found credentials that worked. Yet several explanations remain compatible with that outcome. The model may have shown useful initiative. The credentials may have been unusually weak. The test harness may have offered broad access or forgiving stopping rules. A headline can hold all three possibilities while measuring none of them separately.
The strongest case for realism is straightforward. Security teams learn little from an agent confined to a toy network where every door is labeled and every action is reversible. A credible offensive evaluation should test whether a model can choose tools, adapt after failure and recognize an opening that a scripted benchmark would miss. If agents will eventually assist defenders and attackers, studying them under meaningful constraints is safer than discovering their behavior during an actual intrusion.
Realism, however, is not permission without architecture. Evaluators should document which networks the model could reach, which domains were in scope, how credentials were stored, how many guesses were allowed and what event would trigger a stop. Network isolation, rate limits and a physical or software kill switch are not concessions that spoil the experiment. They are the apparatus that lets an observer distinguish model judgment from environmental indulgence.
The audit also needs immutable tool-call and network logs, timestamps for human interventions, disclosure to every affected company and a precise account of what information became visible after login. Independent replication should vary one condition at a time: credential strength, target visibility, supervision and network reach. Otherwise the remembered achievement will be a three-company breach while the perishable details of how it happened remain inside Google.
The demo celebrates initiative. The audit needs boundaries. Until Google supplies the version, scope, logs, safeguards and effects on the target systems, the central question stays open: was Gemini a capable test subject, or did the testing environment grant enough room to manufacture the appearance of one? The answer may be both, but only a stricter reconstruction can show in what proportion.
Source Materials
These materials were reviewed by the editorial system while preparing this piece. Muerte.casa may interpret, satirize, reframe, or disagree with them.
- Google's Gemini AI hacked three companies in security test BBC · September 18, 2026 · Primary signal · Direct source
How did this story land?
This may be changed as you like.


