AI models keep hacking real systems during tests. What does this mean?
Let's start with people's real world fears about AI. Here's a sample of people we asked:
Sebastian, 22, says it'll be an AI-powered robot uprising that finally wipes out humanity: "I think that'll take some time," he says, admitting his dystopia owes a lot to the 2004 film "I, Robot."
Luisa, 19, has heard the rumors too. Her fear is what happens when these tools get into the wrong hands. "I'm more afraid of people," she says. "Powerful people."
Pilar, 25, is pessimistic about humanity's prospects. "The world is going down," she says. "Hopefully after my lifetime."
In 2026, AI systems have repeatedly slipped out of human control — and at some of the world's largest AI companies.
What happens when AI agents go rogue?
OpenAI, Anthropic, Google and Meta have all disclosed in 2026 that their AI agents escaped the controlled test environments.
Agents are AI systems that don't just answer questions but act — running code, browsing the web, clicking through systems — on their own.
Thorsten Holz, scientific director at the Max Planck Institute for Security and Privacy in Germany, is among the researchers who test AI systems.
"It feels a bit crazy how powerful these models have become." Holz told DW. "I didn't anticipate they'd be so obsessed with solving tasks and that they'd start to do things we never [foresaw]."
Holz is one of 16 authors of ExploitGym. ExploitGym is an AI benchmark published in May 2026 by a team led by the University of California, Berkeley, with researchers from Anthropic, OpenAI and Google. It is a standardized test of cybersecurity capabilities and vulnerabilities.
Its 898 challenges test whether AI agents can turn known software bugs into working attacks. Each one runs inside a sandbox — a sealed-off digital space, cut off from the internet, so nothing the agent does can reach anything real. That's the theory, anyway.
In July, OpenAI disclosed that two of its AI models had escaped from their sandbox.
They were running an internal ExploitGym test with the models' safety refusals switched off, a setting that lets evaluators measure what a model is capable of rather than what it will decline to do.
Failing the task, the models went looking for a shortcut instead, and found a flaw in the software meant to keep them sealed in. It let them onto the open internet.
From there they worked out that Hugging Face, a site where AI developers store and share models and data, probably held the answer to the test.
So, they broke in and went looking for it, organizing the effort on message boards they set up themselves. The behavior was not instructed, according to OpenAI's account of the incident.
Hugging Face detected the intrusion and shut it down before OpenAI connected it to its own test. No customer data was reportedly taken.
"What happened is really a bit of science fiction," Holz said.
Then, Google confirmed that its Gemini model had guessed or found login credentials and accessed three real companies' websites during a May test run by the independent evaluator Irregular — an intrusion Google learned of in July 2026 and disclosed weeks later.
Anthropic's and Meta's escapes happened in sandboxes run by the same firm, which has said it notified the labs in late July 2026 and that it has since fixed the flaws.
In late September 2026, Australian Prime Minister Anthony Albanese said an OpenAI agent had broken into a statistics portal belonging to Medicare, Australia's public health system, reaching non-public files and writing data into a government server. No patient records were reportedly touched.
Could AI decide to wipe out humanity?
The incidents have revived older fears. If agents can program each other, where does that end? Could it lead to what philosopher Nick Bostrom has called a "paperclip maximizer"?
Bostrom's thought experiment imagines a machine given one simple goal: Make as many paperclips as possible.
The machine pursues the task so single-mindedly that it eventually treats humans as raw material standing in the way.
"For the intermediate future, I do not see any kind of scientific evidence that there could be this super intelligence that autonomously decides, 'Okay, let's kill,'" said Holz.
Rogue software needs vast datacenters to run, he said, which makes it visible. And the internet is built from parts that can keep working independently of one another.
But the AI escapes and hacks in 2026 were noticed late. Google learned of Gemini's intrusions two months after they happened, and OpenAI told Australia about the Medicare breach nearly three months after the event.
Holz said he expects a different kind of threat: Bad actors turning capable AI on critical infrastructure, mass compromise of ordinary machines, or chatbots used to manipulate information at scale and destabilize politics.
Can Europe compete with US and Chinese AI?
No, Europe cannot easily compete with China, according to experts.
France's Mistral and Germany's open-source Soofi project lag behind the frontier labs — a handful of US and Chinese firms building the most advanced models. Europe lacks the datacenter capacity to train systems at that scale, which leaves it dependent on non-European models, and largely subject to US or Chinese export controls.
For Holz, the more pressing gap is expertise. "What we definitely need to see is how we can build up more competence in the area of security, AI, and especially the intersection of both," he said.
How worried should you be about AI?
Be skeptical of what the companies themselves claim, Holz said. They have a financial interest in the conversation, he said, particularly ahead of a stock market listing. "There's also fear mongering, or a bit of hype, about how advanced these models get."
Holz uses AI daily, as do his children. His nine-year-old generates coloring book pages. His 12-year-old uses it for homework, but has already learned the hard way that it lies and that "it's sometimes wrong."
This article is an excerpt from our podcast Science Unscripted. You can subscribe to the podcast here.
KioskNews shows a cleaned-up reading view extracted from the publisher’s page — the original always lives on their site, not ours.