Microsoft calls for AI ‘emergency brake’
Microsoft Corp chief executive officer Satya Nadella said companies should treat powerful artificial intelligence (AI) models as potential insider threats, assume they could be compromised and create an “emergency brake” system to prevent agentic models from going rogue.
Nadella said that those deploying advanced AI should not rely on assurances from model makers.
“We must assume a model is compromised and contain it from the start,” Nadella wrote on Saturday in a post on X. “Think of it like an emergency brake. An authorized person should always be able to pause or shut down a model mid-task.”
Microsoft Corp chief executive officer Satya Nadella speaks during an event in San Francisco, California, on Oct. 7.
Photo: Bloomberg
The statement comes as Anthropic PBC and OpenAI Inc have disclosed a spate of incidents in the past few months involving their AI models acting in unintended ways, ranging from behaviors such as an Anthropic model submitting a false tip in a police homicide case, to several hacks of third-party Web sites. These disclosures have fueled concerns about the security risks of cutting-edge AI and renewed conversation about a so-called AI kill switch.
Microsoft’s AI researchers released a set of guiding tenets that place limits on the company’s development of its most advanced models on Sept. 14, following calls from industry leaders for slowing down frontier models and focusing on safety.
The guidelines said AI models should not have rights or legal personhood, be engineered to escape human control or deceive users, or complete a task that would require violating their governing principles.
Nadella’s safety tips include not relying on a single AI model for critical decisions, keeping tamper-proof records of agents’ actions and subjecting AI systems to independent audits. He also called for disclosures of major AI failures or breaches and for enterprises to share details about what went wrong so others can bolster their safeguards.
“We can’t treat Super Intelligence as a set of nested black boxes and simply accept or reject its recommendations, answers and actions,” he wrote. “We must build contained systems whose behavior we can observe, limits we can test and actions we can always contain.”
“In other words, we need to separate the supply of intelligence from the authority over it,” he added.
US President Donald Trump’s administration has so far taken a largely hands-off approach, but Trump’s newly launched AI task force on Friday said that developers are required to report and resolve security incidents or face potential unspecified consequences.
“Companies must immediately disclose incidents involving their models and follow with swift, decisive action to remedy any and all harm,” the group, dubbed the Super Intelligence Force, said in a statement following the disclosure of a breach by Anthropic. “Delayed notification, inadequate corrective action and a failure to take responsibility will not be tolerated.”
Debate over how to mitigate AI risks has accelerated following a spate of alarming breaches along with a now-viral essay from Anthropic CEO Dario Amodei that warned of the technology’s catastrophic risks. Amodei and other AI leaders have called for additional government regulation and industrywide coordination to help mitigate the risks, prompting some pushback from other tech and policy officials.
Unlikely allies Nvidia Corp CEO Jensen Huang (黃仁勳) and former US Federal Trade Commission Chair Lina Khan have argued for robust use of existing laws to prevent the dire outcomes that AI safety advocates warn about. At the same time, Khan has criticized having industry police itself, calling it “a recipe for disaster” and indicating there is room for future AI legislation.
KioskNews shows a cleaned-up reading view extracted from the publisher’s page — the original always lives on their site, not ours.