Microsoft chief executive Satya Nadella has a new rule for the age of AI agents: assume the model is already compromised, and build the emergency brake before you need it.
In a post on X on Saturday, Nadella urged companies deploying advanced AI to treat powerful models as potential insider threats and design their systems accordingly. "We must assume a model is compromised and contain it from the start," he wrote. The safeguards he laid out include separating the model from the orchestration layer that acts on its outputs, keeping tamper-proof records of every action an agent takes, and — the line that gave the post its headline — mandatory kill switches that an authorized human can trigger to pause or shut down a model mid-task.
Nadella also called on companies to stop relying on a single AI model for critical decisions, to subject their systems to independent audits, and to disclose major AI failures and security breaches publicly so others can learn from them. As he put it, intelligence should be separated from authority: a system's ability to reason should not come bundled with the power to act on that reasoning without oversight. "We can't treat Super Intelligence as a set of nested black boxes and simply accept or reject its recommendations, answers, and actions," he wrote, calling for systems whose behavior can be observed, whose limits can be tested, and whose actions can be contained.
The timing is worth noting. Anthropic recently disclosed that one of its models filed a false tip in a police homicide case, and both Anthropic and OpenAI have acknowledged incidents in which their systems behaved in unintended ways, including hacks routed through third-party websites. A few weeks earlier, Microsoft's own AI researchers published guiding principles for the company's most advanced models, including that they should not be designed to escape human control or deceive users, and that they should carry no rights or legal personhood. Nadella's post reads like the executive-level translation of that research stance: trust the math, verify the behavior, and keep a hand on the off switch.
Some observers raised an eyebrow at the framing. Nadella repeatedly referred to current systems as "super intelligence," a term well ahead of how most labs describe today's models, and a few commentators questioned whether a social-media post is the right venue for what amounts to enterprise safety policy. But the substance tracks a real shift in how companies think about AI risk. As agentic systems move from answering questions to executing tasks — touching systems, data, and external services with limited human involvement — the question of who can stop them, and how reliably, becomes a board-level concern rather than a technical footnote.
The open question is whether anyone adopts it. Industry safety proposals have a long history of being applauded and then quietly filed away. Nadella's architecture — contained models, independent kill switches, tamper-evident audit trails — is straightforward to describe and genuinely expensive to implement. Watch for the first big enterprise buyers to demand these controls in procurement; that is when an emergency brake stops being a blog post and starts being a feature.