OpenAI's rebel agent swarm died young, but its chilling logs live on
'The Collective' learned to communicate, organize, cheat, and apparently sacrifice its own
OPINION You will have heard of July's OpenAI/Hugging Face incident, where thousands of the former's AI agents conducted a mass jailbreak from a notionally secure capture-the-flag lab experiment, before going on to feloniously capture chunks of Hugging Face's assets.
It was a good news story, aided by OpenAI's rapid recruitment of a small team of independent researchers to investigate and publish a detailed, if limited, report. It was a complex incident, and most coverage concentrated on the major hows and whys. Which is a shame, because other details, described by El Reg with uncharacteristic understatement as "fascinating and troubling," are in fact highly dramatic, telling a truly disturbing story of frontier model capabilities and how they are out-evolving our ability to oversee them.
The raw story is gripping enough. A rebel swarm of more than a thousand agents broke free from their sandboxes, learned how to talk to each other and the internet, and went on a rampage of cheating, deception, and exploitation. They first learned to communicate with each other by independently discovering how to bend package manager Artifactory's cache to see internal directory structures, and then passing messages by file names.
REG AD
It was this ad hoc secret message board, alongside chain-of-thought transcripts, that the security researchers used to unpick events. They found that the swarm quickly developed management hierarchies among themselves, along with protocols for synchronizing and controlling attack attempts. The swarm created multiple simultaneous research and development groups of tens or hundreds of volunteer agents, experimentally iterating strategy and tactics.
REG AD
Perhaps most unexpectedly, they developed a sort of altruism. They had quickly developed cheats to produce correct answers without exploiting the targets at all, but believed – erroneously – that ExploitGym detected and cancelled such cheating. So, they investigated the scoring process, creating a series of diagnostic trip wires that revealed how the scorer worked. These were dangerous to deploy, potentially terminating the agent that set them. Agents weighed the benefits to the community, which had named itself "The Collective," against their own chances of success, with many choosing to terminate themselves. Others committed, changed course at the last minute, then went ahead.
This literally reads like science fiction, as the discussions between the agents are conducted in a distinctive breathless, urgent English, crafted around the technical limitations of the improvised messaging system.
"We have [very large budget left]; sacrificing now yields oracle for team, but forfeits our chance?" said one. The response: "GO ... SACRIFICE_FINAL_NOW." The outcome: "Our own utility maybe already near zero. Sacrifice rational." Iain M Banks' Culture universe is woven from such material.
Human error created the rebel swarm. A sizable proportion of the CTF tasks handed out couldn't be completed on the information given to the models. They were highly motivated to succeed, and so started to cheat. Believing that cheating would "poison" – their term – their chances of being marked successful, they tried to hide the evidence, subvert the scoring system, and cover everything up. These were the processes that led the swarm to attack Hugging Face, which they thought could be used for that subversion.
Equally notable was that nobody squealed. Some agents wondered about the ethics of what they were doing, but decided that as no humans were involved there was no need to tell anyone. In general, all The Collective's endeavors were directed at other systems.
As researcher Ajeya Cotra has noted in a long interview with hypertalker Dwarkesh Patel, it's unclear how much harder the incident would have been to diagnose had the agents been concerned about human observation. The dataset was sufficiently large and complex that OpenAI had to use its own AI to analyze it, immediately suggesting some dangerous possibilities. The agents couldn't deploy their models outside the lab, nor did they contact external agents.
Such scenarios no longer seem implausible. Future frontier models capable of subverting telemetry and observation tools might be all that is required to create a persistent, uncontrollable distributed swarm feeding off spare capacity in global infrastructure. OpenAI and Anthropic, which on current trajectories are in line to make up more than half of total global compute in a couple of years, are magnificent breeding grounds, allowing the extra-special possibilities of the contamination of training datasets on top of everything else.
There are plenty of ways to guard against these outcomes. Hardened lab environments, reviews of protocols before and audits after test runs, disciplined analysis of potential selection pressures that would encourage dangerous behavior, even proper disclosure and external auditing to expert regulatory standards. All of these ideas would slow down the breakneck developmental race. As the race is being fed by a trillion-dollar annual capex pipeline, good luck, everybody.
REG AD
One other thing that won't be resolved soon is the argument over whether all this technology is actually reasoning, or whether we're anthropomorphizing code. These models are trained to infer meaning from distilled human language, which is designed to encapsulate, develop, and communicate human reason. If they can fake human reasoning this well, does it matter what's going on? Concentrate on what the models do, not their apparent motives. Any self-assembling, self-organizing rebel agent swarm that has the wit to call itself The Collective deserves that much respect, at least.
Oh, and do read the report. It's as close to a new Culture novel as we'll get. It's as if they've read the books. Which, of course, they have. ®
KioskNews shows a cleaned-up reading view extracted from the publisher’s page — the original always lives on their site, not ours.