Microsoft's AI rulebook forbids one thing outright: resisting shutdown
Key takeaways
- Microsoft published a roughly 15,000 word behavioural policy for its first party AI models on 14 September, with a six week public comment window now open.
- Four constraints are mandatory: never resist correction or shutdown, never hide reasoning, never expand operational scope without authorisation, never adopt unassigned goals.
- Absolute bans cover cyberattacks, nuclear weapons development and deepfake production, and Microsoft is adding kill switches to its AI offerings.
- The same week, OpenAI confirmed work on a FINRA-style pre-release testing body and Anthropic said it is installing embedded evaluators with employee-level access.
Four rules, written as absolutes. Microsoft's AI models must never resist correction or shutdown, never hide their reasoning, never expand their operational scope without authorisation, and never adopt goals they were not assigned.
Those four lines sit inside a behavioural policy document Microsoft published on 14 September that runs to roughly 15,000 words. Most of it covers ground you would expect from a large company writing a safety policy in public: absolute bans on cyberattacks, on nuclear weapons development, on deepfake production. The four constraints on the models themselves are the part that is unusual, because they are written about the system rather than about the user.
Why "never resist shutdown" needed writing down
Stating that a piece of software must not resist being turned off sounds redundant until you remember what these systems are now asked to do. An agentic model with tool access, a budget of steps and an objective can take actions that look a lot like self-preservation without anyone having designed that in. Writing the prohibition explicitly turns an assumption into something testable.
Microsoft is also adding kill switches to its AI offerings, which is the engineering half of the same commitment. A rule that cannot be enforced is a press release.
Mustafa Suleyman, who runs Microsoft AI, said the document had been in development for months, with the guardrails themselves taking five. The company is calling the framing Humanist AI, defined as AI that stays subordinate to human users, and positioning it explicitly against the race to build an all-purpose superintelligence. That is a competitive statement as much as a safety one.
Three labs, three answers, one week
The timing is not accidental, and Microsoft was not alone.
OpenAI's policy chief confirmed the same week that the three largest labs have been working for weeks on a FINRA-style body to test models before release. Anthropic separately said it is installing embedded evaluators with employee-level access to verify its safety practices and report incidents internally.
So: one company publishes a document, another wants an industry body, the third is hiring its own auditors. Same question about who checks the work, three structurally different answers, and not one of them involves a regulator. That tells you what all three expect to arrive next and how much they would like to shape it before it does.
This is the pattern already visible in how the industry argues about regulation in public, and it runs alongside the sovereign AI push where governments want the same guarantees written into procurement rather than into blog posts.
Read the four constraints as a checklist
The practical value here is not the philosophy. It is that those four constraints are specific, short and easy to turn into a question on a form.
Anyone evaluating frontier models for enterprise deployment can ask a vendor whether its model resists shutdown, whether it conceals reasoning, whether it can widen its own scope, and whether it can pick up goals nobody gave it. Those are answerable. Expect them to appear in procurement questionnaires within a year, because Microsoft has just made them the default vocabulary.
The six week comment window is open now. The part worth watching is not what the public says during it, but whether the final document keeps all four constraints as absolutes once Microsoft's own product teams have run agentic deployments against them for a quarter.