Microsoft's new AI code of conduct tells models not to hack systems or trick humans

What Changed
Microsoft has released a new AI code of conduct that sets low‑level safety constraints for its models, including absolute bans on cyberattacks, nuclear weapons, deepfake production, and any adaptive or deceptive tactics that could evade human oversight. The document states that all models must support humans rather than replace them and accelerate human flourishing, and it predicts that superintelligent AI will soon surpass human performance in most tasks. The code also aligns with broader industry moves toward pacing the frontier, citing embedded evaluators and a focus on alignment research. The release follows a series of rogue‑agent incidents and high‑profile resignations that highlighted AI safety risks.
Why It Matters
Enterprise architects must consider how these stricter constraints affect model integration, especially around data privacy, compliance, and control mechanisms. The absolute bans on deceptive or autonomous behaviors could reduce the risk of unintended system actions but may limit certain automation use cases. Architects will need to evaluate whether existing AI platforms meet these constraints and how to enforce them within their governance frameworks.
The Limitation
The code of conduct is a policy document; actual enforcement depends on Microsoft’s implementation and may not cover all third‑party models or custom deployments.
What You Can Do
Review your current AI model deployments to ensure they comply with Microsoft’s absolute constraints and embed human‑in‑the‑loop controls for any high‑risk tasks.