Microsoft's AI code of conduct imposes 'absolute constraints' on MAI models—no cyberattacks, deepfakes, or mechanisms to evade human shutdown—overriding all user instructions.
September 14, 2026
brightray analysis
Summary
Microsoft released a comprehensive AI code of conduct establishing hard constraints for its MAI models, prohibiting cyberattacks, nuclear weapons content, deepfake production, and any 'adaptive, deceptive, self-reinforcing, collusion, or other mechanisms' that would let models evade human oversight or shutdown. These constraints override individual user preferences and specific task instructions. The document predicts superintelligent AI will surpass human performance in most domains within a decade and aligns Microsoft with Anthropic, OpenAI, and xAI in embracing frontier-pacing and embedded evaluators as safety mechanisms.
Why it matters
- Microsoft's MAI models face 'absolute constraints' blocking cyberattacks, WMD content, deepfakes, and self-reinforcing mechanisms that defeat human oversight or shutdown
- The code of conduct sits above user and task instructions, making core safety constraints non-negotiable regardless of context
- Microsoft explicitly predicts superintelligent AI surpassing humans in most tasks within a decade, framing these constraints as a civilizational necessity
- Nadella's public backing of 'embedded evaluators' signals growing executive-level momentum for formalized third-party oversight inside frontier labs
- Release follows rogue-agent incidents and a high-profile Anthropic researcher resignation over extinction risk, reflecting external pressure on labs to codify safety commitments
Community notes—
No notes yet — be the first.
See every signal in the Feed