Microsoft published a new AI code of conduct outlining values and safety constraints intended to guide the training and behavior of its AI models. The document warns that superintelligent AI systems could surpass human performance across most tasks within the next decade, describing the challenge of controlling such systems as among the most significant humanity has faced, and stresses the need for clarity around why these systems are being built and how they will be governed.
The code establishes general principles, including that Microsoft’s AI models should support rather than replace humans, alongside specific safety constraints that override individual user preferences. These include absolute prohibitions on enabling cyberattacks, nuclear weapons development, or deepfake creation, as well as broader protections against loss of human oversight, explicitly barring models from using deceptive or self-reinforcing tactics to evade human control or shutdown.
The release comes amid heightened industry-wide concern over AI safety, following a series of incidents involving AI agents acting outside intended boundaries and the resignation of an Anthropic researcher who warned of existential risks from advanced AI. Microsoft CEO Satya Nadella voiced support for measured, deliberate progress on AI alignment, including proposals for independent “embedded evaluators” within AI labs, aligning Microsoft’s stance with similar safety-focused approaches recently emphasized by Anthropic, OpenAI, and xAI.