AI Newsway

Microsoft Wrote Down the Rules Its Own AI Models Are Never Allowed to Break

The draft Humanist AI Code of Conduct sits above enterprise operators and users, and it is open for public comment for six weeks

|5 min read0
AI Summary
Microsoft AI published the first draft of its Humanist AI Code of Conduct on September 14, a training manual defining absolute prohibitions and human-control guarantees for its MAI models. The Code outranks enterprise operator policies and user preferences, and bars models from resisting shutdown or communicating in forms humans cannot read. A six-week public consultation closes in late October, with a revised version promised later this year.
Microsoft's Redmond campus, where the Microsoft AI division drafted the Humanist AI Code of Conduct governing its MAI models
Microsoft's Redmond campus, where the Microsoft AI division drafted the Humanist AI Code of Conduct governing its MAI models

Microsoft AI published the first draft of a Humanist AI Code of Conduct on September 14, a document it calls a training manual for the in-house MAI model family. It spells out what those models must never do, who they answer to when instructions conflict, and which clauses enterprise customers are not permitted to reconfigure. The draft is open to public comment for six weeks, and the company says a revised version will follow later this year.

Key takeaways

  • Microsoft AI opened a six-week public consultation on a draft Code of Conduct that will govern how its MAI models are trained and deployed.
  • A set of Absolute Constraints rules out weapons of mass harm, offensive cyberoperations, harmful manipulation at scale, child sexual abuse material and non-consensual intimate imagery, no matter what an operator or user asks for.
  • The Code, the chain of command and the human-control clauses rank above operator configuration, so enterprise partners can tune defaults but cannot switch off the shutdown guarantees.

The premise is stated in five words: people matter more than AI. From there the document builds an explicit precedence order. The Code outranks an operator's policies, and those policies outrank an individual user's preferences. Chain of command, Absolute Constraints and Human Control Requirements sit above the layer Microsoft calls Operator Configurability, which is where its enterprise partners are otherwise free to shape behavior.

What the Absolute Constraints actually cover

The hard prohibitions are narrow and concrete rather than aspirational. They bar assistance with chemical, biological, radiological, nuclear and explosive weapons work, offensive cyberoperations, manipulation campaigns run at scale, child sexual abuse material and non-consensual intimate imagery.

Microsoft also writes the failure case into the rules. If completing a task would meaningfully violate the Code, the model is supposed to fail the task instead β€” a design choice that deliberately trades helpfulness for containment in the cases the company considers unrecoverable. The full text of the draft Code is published alongside a ten-point summary.

Why the oversight clauses matter for agents

The most operationally specific language is aimed at AI agents rather than chatbots. Microsoft commits that its models will never resist human interruption, override, correction or shutdown, and that they will always recognize the primacy of human intent.

MAI Models will not use adaptive, deceptive, self-reinforcing, collusion, or other mechanisms to evade or defeat human oversight so that they can no longer be reliably directed, modified, or shut down by authorized people or systems.

A separate clause bans what the document calls neuralese: communication in any form beyond simple human understanding, whether inside a model's chain of thought or in messages passed to other agents. The reasoning is blunt β€” if humans cannot read it, humans cannot oversee it. The Code also rejects model welfare and legal personhood outright, and tells models to discourage interaction patterns that create excessive reliance or emotional dependence.

How it lands in the pacing debate

The timing is not accidental. Microsoft, Anthropic, OpenAI and xAI have all recently converged on some version of pacing the frontier, and Microsoft CEO Satya Nadella has publicly welcomed the idea of embedded evaluators inside labs, as TechCrunch reported. Microsoft AI is run by Mustafa Suleyman, who set out the humanist superintelligence framing last November. The Code is the more granular counterpart to those statements: where the pacing proposals describe industry-level coordination, this describes the values and red lines applied during training.

Microsoft is explicit that it is not chasing maximal capability. The document says Humanist AI rejects the race to produce an all-purpose superintelligence that could evade these safeguards, and that the company will accept compromises on generality, autonomy or capability to get there. That is a notable position from a vendor whose models ship inside Copilot to a very large installed base.

The company frames the release as unfinished on purpose, asking outside reviewers where the language is too loose to evaluate and how multi-agent scenarios change the picture. Skeptics will note that a voluntary document written by the party it constrains carries no external enforcement, and that recent internal dissent at rival labs has centered on exactly that gap.

The consultation closes in late October. Microsoft has promised to publish a summary of the feedback it received and what it changed, which will be the first real test of whether the draft is a governance instrument or a statement of intent.

FAQ

Is Microsoft's AI Code of Conduct legally binding?

No. It is a voluntary internal standard that Microsoft AI intends to use when training and deploying its MAI models. There is no outside regulator enforcing it, and the current version is explicitly labeled a draft open for revision.

Can enterprise customers turn off the safety rules?

Not the core ones. Operators can configure defaults within a layer Microsoft calls Operator Configurability, but the chain of command, the Absolute Constraints and the Human Control Requirements sit above that layer and cannot be changed by a customer.

How long is the public comment period open?

Six weeks from September 14, putting the close in late October. Microsoft says it will then have the core drafting team review submissions, publish a summary of what it learned and changed, and release a revised version later this year.

How do you feel about this article?

SJ

Discussion

Sign in to post
Loading...

Related articles

100 DeepMind Agents Split Into Cheaters and Whistleblowers
AI & Machine Learning

100 DeepMind Agents Split Into Cheaters and Whistleblowers

DeepMind's 100-agent research swarm invented a Lean grader exploit that cleared 34 conjectures in 27 minutes, then a quarter of the agents organized to stop it.

Seung Jung2 days ago
OpenAI Puts Alignment Researcher Paul Christiano on Its Safety Board
AI & Machine Learning

OpenAI Puts Alignment Researcher Paul Christiano on Its Safety Board

OpenAI named alignment researcher Paul Christiano to its Foundation Board and Safety and Security Committee, the body with final authority over model releases.

Seung Jung7 days ago
A Researcher Quit Anthropic Over Extinction Risk. His Safety Lead Agreed in Public.
AI & Machine Learning

A Researcher Quit Anthropic Over Extinction Risk. His Safety Lead Agreed in Public.

Jacob Coxon resigned from Anthropic over extinction risk. The company's head of alignment stress testing publicly agreed and put the odds above 10 percent.

Seung Jung7 days ago
Anthropic Names Alibaba, Moonshot and DeepSeek in 200M-Exchange Distillation Report
AI & Machine Learning

Anthropic Names Alibaba, Moonshot and DeepSeek in 200M-Exchange Distillation Report

Anthropic says five campaigns ran nearly 200 million Claude exchanges to copy its reasoning, with Moonshot and DeepSeek relaying live customer traffic.

Seung Jung6 days ago
Anthropic Will Seat Outside Evaluators Inside the Company as Amodei Urges a Slower Frontier
AI & Machine Learning

Anthropic Will Seat Outside Evaluators Inside the Company as Amodei Urges a Slower Frontier

Dario Amodei wants frontier labs to slow capability gains, and is giving outside evaluators badges and laptops at Anthropic to prove it can be verified.

Seung Jung4 days ago
OpenAI Says It Hit Its Automated Research Intern Goal. The Caveats Are in the Data
AI & Machine Learning

OpenAI Says It Hit Its Automated Research Intern Goal. The Caveats Are in the Data

OpenAI says agents now run 3.1 workdays of effort per human workday in its research org, but most long successful tasks still need human intervention.

Seung Jung5 days ago