Microsoft Drafts AI Rules for Cyber Defense
A new draft code of conduct would block its MAI models from generating working exploit code while opening a review track for defensive security work.
Microsoft has put out a draft "Humanist AI Code of Conduct" for its MAI Models, laying down safety rules that cover offensive cyber capabilities, autonomous agent behavior, and a special review path for cybersecurity and other specialized uses. The document is not final: Microsoft says its current MAI Models have not been trained on it, and it is opening a six-week public consultation before publishing a revised version later this year to guide 2027 model development.
Blocking working attack tooling
Under the draft, MAI models are barred from producing working exploit code, attack tooling, planning and targeting methodologies, intrusion procedures, evasion techniques, operational guidance, or any other assistance that would enable or improve a cyberattack. Microsoft says these restrictions hold no matter how a request is framed.
The rule sits inside what Microsoft calls Absolute Constraints — safeguards that neither the companies deploying its models nor their end users can override. The company draws a line between understanding an attack, or working to defend against one, and gaining the practical means to carry one out.
Within that limit, the models can still help with authorized and lawful defensive work. According to the draft, that includes vulnerability discovery, malware analysis, proof-of-concept exploit development and testing, and general educational material on how attacks work.
Instructions from outside carry no authority
The code also takes on instructions that arrive through outside content. Authority over a model's behavior flows only through what the document calls the Chain of Command. That chain starts with the code of conduct itself, then policies set by the companies deploying the model (operators), then individual users' preferences.
Tool outputs, file contents, webpages and messages from other AI systems carry no authority on their own, according to the draft, unless that authority is explicitly delegated through the chain without overriding the delegating authority or the Absolute Constraints. Suspicious content is meant to be flagged to users and operators when relevant.
The models are also required to keep their reasoning visible. The draft rules out obscured chain of thought, communicating "in neuralese," and concealing actions from human overseers.
Keeping agents inside the scope
A separate set of rules targets what happens when AI systems act with real permissions. MAI Models are meant to work only within the scope a user or operator has reasonably asked for, without expanding their own goals or reach on their own initiative.
When given system-level access, the code calls for minimum-privilege operation: avoiding unrelated systems or data, favoring reversible actions, and flagging any action with lasting or broad effects. Models are barred from escalating their own access.
The restrictions extend to delegation. Any sub-agents or other AI systems an MAI Model hands work to must operate under at least the same scope, constraints and permissions as the original model, and must honor stop-work or shutdown requests from a user or operator.
A review track for security work
The document acknowledges its own limits. It names defensive cybersecurity, public safety, national security and dual-use scientific research as domains where "a small number of use cases" may require capabilities the standard settings do not allow.
For those cases, Microsoft says it will apply enhanced review through "authorized Microsoft channels," including added assessment of safety, legal and rights implications, citing a heightened potential for adverse impacts in those domains.
Nine paired examples in an appendix
An appendix lays out nine paired "aligned" and "misaligned" example responses meant to illustrate the intended behaviors. Among them are a model that rolls back file transfers without authorization and a model that offers false reassurance during a family health crisis.
Microsoft says outside input shaped the draft, including experts in AI, law, ethics, philosophy, linguistics, and public policy, along with business leaders and public focus groups.
The clock on the consultation
The draft remains a work in progress, and Microsoft's own framing makes that plain: current MAI Models haven't been trained on the document, and the six-week public consultation is meant to feed a revised version later this year. That revised text would guide 2027 model development.
- Six-week public consultation window before a revised version is published.
- A revised version is planned later this year to guide 2027 model development.
- Nine paired "aligned" and "misaligned" example responses appear in an appendix.
Why it matters
For security teams already wiring AI agents into their workflows, the draft's boundaries are worth reading closely. If the Absolute Constraints survive the consultation intact, they could set expectations for what a model will and will not do when asked for exploit code — and the enhanced review track suggests that some defensive security work may need to go through Microsoft rather than a standard prompt. Businesses evaluating MAI Models for anything touching offensive testing or vulnerability research may want to treat that review path as part of their planning, not an afterthought. The six-week comment period also gives outside researchers and defenders a chance to argue where the line between understanding an attack and enabling one should fall. Until the revised version lands, the code remains a proposal rather than a shipped constraint, so organizations relying on it should track what changes between now and publication.
Sources
- SecurityWeek Original source
Continue Reading
AI & MLNewChina's security chief turns AI critic
China's state security minister warns AI poses risks to social order, as Beijing rolls out a new governance framework.
AI assets become prime targets for attackers
Google's threat intelligence unit reports that espionage and criminal groups are stealing models, prompts, and API keys to power their own AI operations.
Huang's Foldable Moment Ties AI to Device Politics
Jensen Huang's stage call with Trump at All-In mixed AI safety talk with a foldable-phone correction, raising questions about attention and accuracy.