Agent Governance Toolkit: from "asking AI to behave" to "making it structurally incapable of misbehaving"

1. At a Glance

★★★★☆ (4/5). Microsoft’s runtime governance suite for AI agents — a deterministic gate in front of every tool call, making denied actions structurally impossible rather than merely unlikely.

What it doesn’t do: it does not scan your code, does not grade your configuration, and does not try to win the jailbreak fight inside the prompt. The README states the thesis outright: “Prompt-level safety is not a control surface. It is a polite request to a stochastic system.”

Where the missing star goes (section 7 has the full list):

  • Public Preview with only seven months of history — v4.1.0 just collapsed 45 packages into 5 distributions, so the API is still drifting
  • The default is allow when no policies are loaded — governance fails silently while the dashboard still reports “governed”
  • The kill switch doesn’t kill processes in the Go SDK — it is a cooperative circuit breaker, and its semantics differ across SDKs

Verdict: enterprise security and platform teams — the most complete compliance mapping in open source today, run a PoC now. Personal projects — built for enterprise governance; don’t.

Primary language Python (3.11+; the Prerequisites section says 3.10+)
License MIT — core packages carry zero Azure/Microsoft dependencies, runnable offline and air-gapped
Stars / Forks 6,351 / 1,132
Open issues / watchers 82 / 91
First commit 2026-03-02
Last push 2026-09-28 (PR numbers already at #4168)
Commits 2,751 (~7 months, ~13/day)
Release status Public Preview; latest release v4.1.0 (2026-06-09)
Repository size ~46 MB
Language SDKs Python (full stack) / TypeScript / .NET / Rust / Go (the latter four: core only — policy, identity, trust, audit)
Specs & tests 10 RFC 2119 specifications, 992 conformance tests, 29 ADRs, 60+ tutorials
Compliance mapping OWASP Agentic Top 10 / NIST AI RMF 1.0 / EU AI Act / SOC 2 / AARM R1–R9 / ATF (all five elements)
Policy evaluation cost <0.1 ms (single process, full overhead); 5–50 ms in a multi-agent mesh
URL https://github.com/microsoft/agent-governance-toolkit

Who it’s for (out of 5):

Audience Score Why
Enterprise security / CISO / GRC 5 The only open-source project that turns six standards — OWASP Agentic Top 10, NIST AI RMF, EU AI Act, SOC 2, AARM and ATF — into exportable evidence. agt verify --evidence --strict gates CI directly, and the audit chain is Merkle-structured with a Decision BOM
Agent application developers 4 Two-line govern() onboarding and 14 framework integrations (AutoGen, LangGraph, CrewAI, Semantic Kernel, OpenAI Agents, Claude Code, Dify…). The cost is writing YAML policy and understanding verdict semantics
Platform / SRE teams 4 Agent SRE is the real deal: kill switch, SLOs, error budgets, chaos testing, circuit breakers. But kill switch semantics differ across SDKs (section 7) — do not treat it as a life-saving guarantee
Red team / offensive security 3 agt red-team scan --min-grade B runs a 12-vector prompt injection audit and the MCP Security Gateway catches tool poisoning, drift and typosquatting. But this is a defensive tool; attack surface mapping is still manual work
Individuals / small projects 2 Designed for enterprise governance. The docs say so themselves: the full stack (mesh identity, execution rings, SRE) is overkill for simple cases. Use the minimal AgentControl.from_path() path instead

Building on it (MIT, three tiers):

Tier What you do Difficulty
Configuration Write YAML policy (block-destructive, require_approval, blocked_patterns), plug in a framework adapter Low, under an hour
Integration Embed core governance via the TypeScript/.NET/Rust/Go SDKs into your own platform; wire up the MCP Security Gateway Medium — read the relevant spec and its conformance tests
Kernel Touch the Rust ACS policy runtime, privilege rings, Merkle audit, cross-SDK DID standardization High — 10 RFC 2119 specs and 29 ADRs of prerequisite reading

2. What It Is, What It Isn’t

Not a prompt guardrail (it does not ask the LLM to “please follow the rules”). Not model safety or content moderation (it does not judge hallucinations or harmful content). Not OS-kernel-level isolation — the docs state plainly that the policy engine and the agents share the same process boundary, and recommend one container per agent in production. Not a managed cloud service (it is a framework-agnostic library). Not a turnkey compliance solution — the official framing is enforcement infrastructure.

The one-line positioning: Policy enforcement, identity, sandboxing, and SRE for autonomous AI agents. One pip install, any framework.

Four keywords: deterministic (not probabilistic), fail-closed (evaluation errors deny), tamper-evident audit (Merkle), and layered and optional (each layer installs independently; most teams run policy + audit and nothing else).

The README frames the problem through three questions, and this framing is, in my view, the most worth-stealing part of the whole project:

  1. Is this action allowed? OAuth scopes and IAM roles govern which services an agent can reach, not what it does once connected. An agent with send_email and query_database should not be able to drop_table.
  2. Which agent did this? When five agents share one API key, “an agent did it” is not an incident response.
  3. Can you prove what happened? Auditors and regulators need tamper-evident records: which policy was active, what the agent requested, why it was allowed or denied.

It backs the thesis with three hard references: Andriushchenko et al. (ICLR 2025) report a 100% attack success rate against GPT-4o, GPT-3.5, Claude 3 and Llama-3 on JailbreakBench using adaptive attacks with logprob access and suffix optimization; OWASP LLM01:2025 states outright that “it is unclear if there are fool-proof methods of prevention for prompt injection”; and Microsoft’s own AI Red Teaming Agent formalizes Attack Success Rate (ASR) as the canonical metric for this failure class. The conclusion: model-layer defenses are probabilistic by construction, so don’t fight there.


3. The Four Layers

The official pipeline is Agent → Policy Engine → Identity → Audit Log, with every layer optional.

Layer Package What it does Key mechanism
Policy enforcement Agent OS + Agent Control Specification (Rust core) Evaluates every tool call before it happens Stateless, deterministic, fail-closed; four verdicts: allow / deny / transform / require_approval
Zero-trust identity Agent Mesh Who is calling, and how much to trust them SPIFFE / DID / mTLS, Ed25519 signatures, trust scoring, delegation chains
Execution sandboxing Agent Runtime Where it runs, what it can touch Four privilege rings, saga orchestration, command denylist
Reliability engineering Agent SRE How you recover when it breaks Kill switch, SLOs, error budgets, chaos testing, circuit breakers

Five cross-cutting capabilities sit alongside, and two of them have more immediate value than the main layers:

  • MCP Security Gateway — tool poisoning detection, configuration drift monitoring, typosquatting, hidden-instruction scanning (127 conformance tests). MCP is the de facto standard of the agent ecosystem and where the attack surface concentrates; this is one of the few targeted open-source answers.
  • Shadow AI Discovery — finds unregistered agents across processes, configs and repos. For most enterprises, the actual starting point of agent governance is “we don’t know how many agents are running,” and this hits it directly.
  • Plus a Governance Dashboard (Streamlit fleet view), PromptDefense Evaluator (12-vector prompt injection audit), and Contributor Reputation (PR/issue author screening for social engineering, shipped as a reusable GitHub Action).

Policy reads plainly:

1
2
3
4
5
6
7
8
9
10
11
apiVersion: governance.toolkit/v1
name: production-policy
default_action: allow
rules:
- name: block-destructive
condition: "action.type in ['drop', 'delete', 'truncate']"
action: deny
- name: require-approval-for-send
condition: "action.type == 'send_email'"
action: require_approval
approvers: ["security-team"]

And the denial carries the rule name rather than a vague refusal:

1
2
GovernanceDenied: Action denied by policy rule 'block-destructive':
Destructive operations require human approval

4. Architecture: How Two Lines Become a Gate

The core is the govern() wrapper:

1
2
from agentmesh.governance import govern
safe_tool = govern(my_tool, policy="policy.yaml") # every call: evaluate + audit + raise on violation

The interception point sits in deterministic application code, before the model’s intent reaches the wire. That single sentence carries the entire value proposition: it demotes “will the agent misbehave” from a model-behavior question to a function return value. A cleverly worded prompt does not change what the policy engine decides.

Underneath sits ACS (Agent Control Specification) — a stateless, deterministic, fail-closed policy decision runtime with a Rust core. It backs the verdict semantics of the whole policy layer and also provides the minimal onboarding path, AgentControl.from_path("manifest.yaml"), which needs neither mesh nor identity.

On the audit side: a Merkle audit chain plus a Decision BOM (157 conformance tests), recording which policy was in force, what the agent requested, and the verdict. Note the boundary documented in section 2 of LIMITATIONS.md: it records attempts, not outcomes. Section 7 expands on that.

The performance accounting is unusually honest — it does not lead with the flattering number:

Component Typical latency When it applies
Policy evaluation <0.1 ms Every action
Ed25519 signature verification 1–3 ms Inter-agent messages
Trust score lookup <1 ms Inter-agent messages
IATP handshake (first contact) 10–50 ms First message between two agents
Network round-trip (mesh) 1–10 ms Distributed deployments only

For single-agent, single-process deployments, <0.1 ms is the full overhead. In a multi-agent mesh, expect 5–50 ms per governed interaction, and note that the dominant cost is cryptographic verification and network latency, not the policy engine.


5. Compliance and Specifications: The Real Moat

Six standards, all turned into exportable evidence, each backed by an RFC 2119 specification and conformance tests:

Standard Coverage
OWASP Agentic AI Top 10 All ASI risk categories mapped to deterministic controls
NIST AI RMF 1.0 Full GOVERN / MAP / MEASURE / MANAGE alignment
EU AI Act Compliance mapping with automated evidence
SOC 2 Control mapping with audit trail export
AARM Extended All R1–R9 satisfied; verified 2026-06-14
ATF All five elements mapped: Agent Mesh (identity), Agent OS (policy), Agent Compliance (governance), Agent Runtime (sandboxing), Agent SRE (incident response)

The CLI turns compliance into a CI gate:

1
2
3
4
5
agt verify                                          # OWASP compliance check
agt verify --evidence ./agt-evidence.json --strict # fail CI on weak evidence
agt red-team scan ./prompts/ --min-grade B # prompt injection audit
agt lint-policy policies/ # validate policy files
agt doctor # installation self-check

992 conformance tests, 10 specifications, 29 ADRs — a number that leads the open-source agent governance field by a wide margin. The engineering discipline is itself a signal: CodeQL (Python + TypeScript SAST), Gitleaks (secret scanning on PR/push/weekly), ClusterFuzzLite with seven fuzz targets (policy, injection, MCP, sandbox, trust), Dependabot across 13 ecosystems, and weekly OpenSSF Scorecard scoring with SARIF upload.

One inconsistency worth flagging: the repository description says “Covers 10/10 OWASP Agentic Top 10” while the README badge says “7 Full, 3 Partial”, with no reconciliation between them. Trust the badge: seven of ten fully covered, three partially.


6. Deployment and Integration

One line, and use the [full] extra — the base wheel installs only the compliance CLI:

1
pip install "agent-governance-toolkit[full]"

v4.1.0 consolidated 45 packages into 5 distributions, the largest structural change to date: -core (policy engine, capability model, audit, MCP gateway, zero-trust identity, trust scoring, A2A/MCP/IATP bridges), -runtime (privilege rings, saga orchestration, termination control, command denylist), -sre (SLOs, error budgets, chaos, circuit breakers), -cli (the agt command, OWASP verification, policy linting), and [full] as the meta-package. Legacy names (agent-os-kernel, agentmesh-platform, agent-sre, and others) remain installable as stubs that redirect.

Note that agent_os is deprecated: importing it emits a DeprecationWarning, and the replacement is agent-governance-toolkit-core. The pre-v4 agent_os.policies rule model is gone; BREAKING_CHANGES.md lists replacements.

Five language packages:

Language Package Command
Python agent-governance-toolkit pip install "agent-governance-toolkit[full]"
TypeScript @microsoft/agent-governance-sdk npm install @microsoft/agent-governance-sdk
.NET Microsoft.AgentGovernance dotnet add package Microsoft.AgentGovernance
Rust agent-governance cargo add agent-governance
Go agent-governance-toolkit go get .../agent-governance-golang

Fourteen framework integrations: Microsoft Agent Framework (native middleware), Semantic Kernel (native, .NET + Python), AutoGen / LangGraph / LangChain / CrewAI / Mastra / Google ADK (adapters), OpenAI Agents SDK / LlamaIndex (middleware), Haystack (pipeline), Dify (plugin), Azure AI Foundry (deployment guide). Claude Code and GitHub Copilot CLI are first-party developer surfaces built on the TypeScript SDK:

1
2
/plugin marketplace add microsoft/agent-governance-toolkit
/plugin install agt-governance@agent-governance-toolkit

7. Boundaries and Risks

This chapter carries more weight than the others, because docs/LIMITATIONS.md is the most honest self-disclosure document I have seen in an open-source project: 408 lines, 13 boundaries, each with “how to mitigate today” and “what we’re building,” and it reproduces external criticism verbatim with links. The six that matter most:

① It governs actions, not reasoning. If policy allows both read_database and send_slack_message, an agent can read your customer list and post it to a public channel — both actions individually permitted. Worse, there is a cross-session variant: attack state carried by persistent memory or persistent tools (notes, files, calendar) crosses session boundaries under permission isolation, so every session looks compliant while the full attack chain only resolves across the session sequence. Dai et al. (preprint, 2026-05) report 80–95% attack success rates across four base models under supply-chain SFT delivery. The docs concede this is not closed today (workflow-level policies are on the roadmap).

② The default is allow when no policies are loaded. The most silent-failure-prone item on the list: when the policy evaluator has no policies loaded, the default action is allow, and misusing permissive mode does the same. A developer who imports the governance layer but forgets to load policy files sees a dashboard reporting “governed” while zero rules are enforced. Credit where due: on runtime errors during evaluation, AGT fails closed (denies) — the direction is right. Mitigation: run strict mode in production (deny-by-default, requiring an explicit allow for every permitted action) and use agt audit to inspect what is actually loaded. This item came from external red-teaming — Periculo reported 15 bypass vectors.

③ The audit trail records attempts, not outcomes. An agent calls a web API that returns 200 with stale data; the audit chain still records “action allowed, executed” even though the agent’s goal was not achieved. For a project whose stated goal is “can you prove what happened,” that is a material gap. Mitigation today: SRE SLOs plus application-level result validation; post-action verification hooks are in progress.

④ Knowledge governance is missing. AGT does not govern the knowledge agents consume — not the provenance, freshness or authorization of documents, databases or embeddings retrieved during reasoning. An agent retrieves a confidential HR document (policy permits), summarizes it into a Slack message (policy permits); both actions are governed, but the knowledge flow — confidential data reaching an unauthorized channel — is invisible to AGT. From external analysis by Mojar AI.

⑤ Kill switch semantics differ across SDKs — the easiest one to misread. “Kill switch” does not currently describe a single cross-SDK termination contract. A recorded kill is not by itself proof that a process stopped:

Behavior Python TypeScript Go
Model Registered termination callback plus step handoff/compensation Registered termination handlers plus compensation handlers Scoped cooperative allow/deny registry
Termination success signal KillResult.terminated KillSwitchResult.terminated (handlers completed within budget; no independent process confirmation) None — KillSwitchDecision.Allowed describes permission, not process termination
Stops a non-cooperating execution Only if the registered callback stops it Only if a registered handler stops it No — the consumer must consult DecisionFor()

There is an extra trap in Python: after a failed or timed-out kill attempt, the target’s callback registration is removed unconditionally, so you must re-register the agent before retrying. Do not treat the kill switch as a life-saving guarantee. It is cooperative.

⑥ Credential persistence is unmanaged. AGT does not track which API keys, OAuth tokens or secrets an agent currently holds, does not revoke credentials at task boundaries within a session, and does not detect credential accumulation beyond what the current task needs. An agent receives an email token for Task A, moves to Task B which needs no email access, and the token persists — so a compromise during Task B hands the attacker email access that should have expired.

The remaining seven: performance framing (the <0.1 ms figure measures the policy engine only, excluding distributed overhead); complexity spectrum (full stack is overkill for simple cases); vendor independence (MIT with zero Azure dependencies — actually a strength); physical AI is out of scope (no hardware kill switches, force limiting or actuator safety interlocks, and real-time control loops under 10 ms may not tolerate the full governance stack); no continuous assurance over streaming data (it governs the subscribe action, not message-level quality); DID method inconsistency across SDKs (Python/.NET use did:mesh:*, TypeScript/Rust/Go use did:agentmesh:*, so prefix-matching policies silently miss half your fleet — mitigate with did:* wildcards or normalize at the application boundary); and application layer, not kernel layer (the policy engine shares a process with the agents, so production requires one container per agent).

The official positioning is clear-eyed: AGT is one layer in a defense-in-depth strategy, not the entire strategy. Model safety (Azure AI Content Safety, Llama Guard) filters inputs and outputs; AGT enforces actions; the application layer validates business logic; the infrastructure layer (containers, network policy, IAM) catches escape attempts.


8. Getting Started

  1. pip install "agent-governance-toolkit[full]" (Python 3.11+), then run agt doctor — it also confirms that no installed package requires cloud connectivity.
  2. Take the minimal path before considering the full stack: AgentControl.from_path("policies/manifest.yaml") gives native ACS policy evaluation with no mesh or identity dependency. Most teams stop here.
  3. Start with three rules: one deny for destructive operations (drop/delete/truncate), one require_approval for outbound actions (send_email), one blocked_patterns regex to catch PII. Don’t build out a full rule table on day one.
  4. Always run strict mode in production (deny-by-default). It is the only configuration that defends against “no policies loaded → silent allow.”
  5. Pick your integration: adapters for AutoGen / LangGraph / CrewAI, middleware for OpenAI Agents / LlamaIndex, native for Semantic Kernel. Nine examples ship in examples/, including multi-agent role-based policy with CrewAI, a trust-verified MCP server, and Kubernetes agent-sandbox with pre-dispatch command policy.
  6. Pin agt verify --evidence ./agt-evidence.json --strict and agt lint-policy policies/ into CI. Policy files deserve the same review process as code.
  7. On Claude Code: /plugin marketplace add microsoft/agent-governance-toolkit then /plugin install agt-governance@agent-governance-toolkit — one command to a governed plugin.
  8. Deployment shape: one container per agent (the official production recommendation). Azure / AWS / GCP / Docker Compose all documented. Air-gapped environments work: core packages have zero cloud dependencies, and all governance state uses standard formats (YAML, JSON, Ed25519 keys) — no proprietary lock-in.
  9. Read docs/LIMITATIONS.md before you ship (408 lines, 13 boundaries) plus docs/security/threat-model.md. That document deserves to be read before the README.

9. The One-Line Verdict

The most disciplined and most completely compliance-mapped open-source answer to “how do I govern AI agents at runtime” right now — 992 conformance tests, 10 RFC 2119 specifications, 29 ADRs, exportable evidence for six standards, and a self-disclosure document that reproduces external red-team criticism verbatim. That honesty alone earns a star. The missing star is equally clear: Public Preview with seven months of history and a v4.1.0 that just collapsed 45 packages into 5 (the API is still drifting); the default is allow when no policies are loaded, so governance can fail silently; the kill switch does not kill processes in the Go SDK and its semantics differ across SDKs; plus three material blanks — actions are governed, but not reasoning, knowledge flow or credential persistence. For enterprise security and platform teams: run a PoC now, start from the minimal path in strict mode, and don’t deploy the full stack on day one. For personal projects: this was built for enterprise governance and you almost certainly don’t need it. First move: pip install "agent-governance-toolkit[full]" → agt doctor → wrap your single riskiest tool function with three policy rules.

评论Comments