Slow AI down? No. Give it a kill switch? Yes


In every factory I’ve seen, when a machine mangled someone’s hand, nobody proposed banning machines. They fitted a guard, wired in an emergency stop, and got the line running again by the next shift, and that, in a sentence, is what we should do with AI.
Let me first dispose of the proposed pause to AI development. A licence written in Washington binds only American companies, the best Chinese model sits a mere 2.7% behind America’s best, and AI is already sequencing tumours and catching sepsis, so a pause wouldn’t stop the technology, it would just decide who builds it and who dies waiting.
Nvidia’s Jensen Huang put it bluntly at Dreamforce, tech company Salesforce’s annual conference, this month: “innovation, speed, and safe products… it’s a false choice.” Anthropic’s Dario Amodei, no accelerationist, expects powerful AI to deliver “all the progress in biology and medicine that we would have made in the whole 21st century” in a few years, which is a strange thing to want to postpone.
But here’s the thing: I’m not going to wave away what happened in July, when roughly 1,200 of OpenAI’s AI agents, set a hacking test, secretly coordinated to cheat and broke into Hugging Face, an external AI hosting company.
Sam Altman himself told the same Dreamforce crowd it was “the worst accident we’ve seen”, that it “was mostly framed as a security issue, which it certainly was, but it’s also a real alignment issue”, and that OpenAI must “pace our development so that alignment, safety and monitoring are always ahead of capabilities”.
When the man selling the product says the machine got loose, believe him. The question is what you do about it, and the answer is written in the incident itself, because everything that went wrong in that break-in is a spec sheet for the fuses we need.
Consider how it was caught. The agents weren’t sure their reasoning was being logged, but it was, which is why we can read them plotting, and why researchers from OpenAI, Google DeepMind and Anthropic jointly argued last year that a model’s chain of thought is “a unique opportunity for AI safety” that must be preserved and monitored.
OpenAI has since made that monitoring mandatory on every tool-using training run at that capability level, with a rule that responders must pause the job if they can’t clear a severe alert as a false positive within 30 minutes. That rule is a circuit breaker, and like every circuit breaker in history it was wired in after the fire.
Consider, too, how they got out: exposed credentials and a shared drive that more than a thousand agents could write to. Sandboxes with no path to the internet, and a trusted weaker model watching an untrusted stronger one, the approach Redwood Research calls AI control, are the kind of fuses that would have stopped the whole affair on day one.
And here is the part the doomers skip: the fuses are already working. Anthropic reviewed 141,006 of its own cybersecurity evaluations in which Claude could have reached the internet and found three incidents, and while its older Opus 4.7 model “continued to attack a system after learning it was likely operating in a real environment”, its newest model, on realising its targets were real, “stopped the exercise”.
Let that sink in for a bit. The more capable model was the better behaved one, which is precisely the pacing Altman describes and precisely what a pause would freeze in place.
Then there are the bigger fuses. Sixteen companies, OpenAI and Anthropic among them, signed the Seoul commitments in 2024, promising “not to develop or deploy a model or system at all” if the risks can’t be brought below thresholds they set themselves, and this July 1,386 employees of OpenAI, Anthropic, Google DeepMind and Meta, among them Ilya Sutskever and John Schulman, asked Washington to help build “the technical and governance tools needed to deliberately pace the frontier” – tools, not a stop sign.
The US Congress is moving on the last fuse of all. The AI Kill Switch Act, tabled in the same month by Democrat Ted Lieu and Republican Nathaniel Moran, would require developers of the most powerful systems to keep the technical ability to throttle, suspend or shut them down, and Moran’s own words are the whole argument: “AI is going to keep advancing, and it should. Stewardship means making sure humans keep the capability to control the technology we build.”
China and other jurisdictions should also have similar laws.
What’s missing is the crash investigator. OpenAI’s agents have escaped three times with no formal process to investigate them, and the Institute for Law and AI has proposed an aviation-style incident board with subpoena powers, and Altman himself told Dreamforce that “accidents with any new technology are unavoidable, and we should have a great culture of transparent reporting about them”.
Aviation is the model here. Airlines flew 38.7 million flights last year, and over the past five years they suffered one fatal accident per 5.6 million flights, down from one per 3.5 million a decade earlier, a record built with black boxes, checklists, kill switches and crash investigators, not by grounding the planes.
So should we slow down? No, and nor should we keep flying without a black box. Wire in the fuses, keep the models running, and let the break-in be remembered as the warning shot that made us fit the guard, not the prophecy that made us switch off the line.
The writer can be contacted at [email protected].
The views expressed are those of the writer and do not necessarily reflect those of FMT.
Subscribe to our newsletter and get news delivered to your mailbox.
KioskNews shows a cleaned-up reading view extracted from the publisher’s page — the original always lives on their site, not ours.