The systemAgent governanceSheet 47

The agent governance stack.

An agent that can act is an agent that can act wrongly, at speed, with your name on it. This is the stack that decides what it may do, and what you can prove afterwards, with every stage resolving to a published specification rather than a principle.

§ 01Definition

Eleven questions, in the order they fail.

Descriptive termDefinition v1.0 · 2026-08Nabeel Khan

An agent governance stack is the ordered set of controls that decide what an AI agent is, what it may do, what constrains it, what it must record, and who is accountable when it acts.

The term is not this practice’s invention and this page does not claim it. It is in common use, and § 04 names the serious work already using it. What is offered here is a specific ordering, and one property that is unusual: every stage below resolves to a published specification with its own page, not to a principle. A governance model you cannot implement on a Tuesday is a diagram.

The order matters because it is the order in which things fail. An agent with no identity cannot have bounded authority, because there is nothing to bound. Authority granted without an enumerated capability set is authority over whatever the tool happens to expose. And evidence captured after the fact is not evidence, it is reconstruction. Each stage is only as real as the one above it.

§ 02The stack

Each stage, and what breaks without it.

Identity purpose authority capability policy memory evidence action verification escalation audit
StageThe question it answersWhat happens when it is missingImplemented by
01 IdentityWho is this agent, distinctly from the human who started it?Actions land under a shared service account and no one can be asked about them individually.Agent registry entry, per-agent credential
02 PurposeWhat does it exist to do, narrowly enough to be falsifiable?Scope creeps by prompt. The agent that summarised tickets is now approving refunds.A written mandate, reviewed at Gate 1
03 AuthorityWhat may it decide alone, and what needs a human?The boundary lives in a prompt, which is a request rather than a control.Trust tiers, specified on the agentic AI page
04 CapabilityWhich tools and MCP servers may it actually invoke?Authority over whatever the connected surface exposes, which nobody enumerated.Capability contracts · MCP governance
05 PolicyWhich constraints bind it at runtime, not on review day?Policy exists in a document the running system never reads.Policy-as-code in the deployment path, PARA
06 MemoryWhat may it retain, and what must it forget?Yesterday’s customer record influences today’s unrelated decision, invisibly.Retention scope per PARA; adaptation writes to knowledge, never to production
07 EvidenceWhat did it rely on, captured when it relied on it?The defence eighteen months later is reconstruction, and the regulator can tell.The evidence store, LLM systems in production
08 ActionWhat did it actually do, as distinct from what it decided?Deciding and doing blur, so a wrong action cannot be traced to a wrong reason.PEVG: the executor calls tools and decides nothing
09 VerificationWho or what checked it, independently of what produced it?A system that checks its own work has no meaningful check.PEVG: the verifier holds binding judgement and may say unverified
10 EscalationWhen must a human intervene, and does the path exist before it is needed?The first escalation is invented during the incident, by whoever is awake.Tiered human review · AIRP, the AI Incident Response Protocol
11 AuditCan the whole chain be reconstructed by someone who was not there?Every stage above was done and none of it can be shown.MESA layer scoring · the Teardown

Read the third column first. It is the one that tells you whether you have the stage or only the word for it. Autonomy is a governed capability, not a feature, and the difference between those two sentences is eleven rows of evidence.

§ 03Why now

MCP turned tool access into a governance problem.

A connected MCP server is a production dependency that grants capability at runtime. Stage 04 stopped being a design-time list the moment an agent could acquire tools it was never reviewed with.

That is the shift that makes the whole stack urgent rather than tidy. Before, an agent’s reach was whatever was compiled into it. Now it is whatever it can connect to, which means capability, policy and evidence have to be enforced at the connection rather than in a review that happened last quarter. The specifics, including what to require of a server before an agent in a regulated estate may call it, are on MCP governance.

It cuts the other way too, and that is the more interesting half. If an agent will consult external judgment at runtime, the provenance and honesty of what it consults becomes part of your governed surface. An interface that answers confidently about a regulation that does not exist is a supply-chain failure in stage 07. That argument, and a working demonstration of the alternative, is at machine-accessible AI expertise.

§ 04Adjacent work

Who else is building this, named.

Agent governance is a crowded and serious field, and a page that implied otherwise would not be worth citing. What follows is where the substantial work sits and what it does not cover, which is the only honest way to explain what is different here.

Microsoft’s Agent Governance Toolkit is runtime enforcement: policy, zero-trust identity, execution sandboxing. It is strong on stages 01, 04 and 05 and it is infrastructure, which means it enforces decisions someone else has to make. Credo AI’s GAIA puts a governance assistant inside a governance platform, which is a system of record for stages 02 and 11. The identity vendors are converging on agent identity as a third IAM domain, which is stage 01 done properly and stops there. The academic control-plane work, including the AGL-1 and CAGE-1 papers, is the most complete published treatment of the middle of the stack.

What none of them supplies is the part that is not software: which controls a given regulator will actually accept, what evidence satisfies them, and where the boundary sits between a decision an agent may take and one it may not in a specific institution under a specific supervisor. That is judgement, it is what this practice sells, and stages 03, 07, 10 and 11 are where it is spent. The tooling above and this work are complements, not competitors, and an institution running agents at scale will want both.

§ 05What binds it

The stack is not optional for much longer.

Dates rather than direction, because direction is what makes a board defer. EU AI Act Article 50 transparency duties have been in force since 2 August 2026, while the Annex III high-risk obligations moved to 2 December 2027. OSFI Guideline E-23 takes effect 1 May 2027 for Canadian federally regulated institutions and now explicitly covers AI and machine-learning models, insurers included. DIFC Regulation 10 has been in full enforcement since 1 January 2026, requiring an AI register, certification and an Autonomous Systems Officer for high-risk processing. The CBUAE issued its AI and machine-learning guidance note on 23 February 2026, supervisory rather than binding.

Be exact about the negatives too, because they are where confident summaries go wrong: no GCC state has a binding horizontal AI statute, Saudi Arabia governs through SDAIA guidance that is non-binding with the PDPL carrying enforcement, SAMA has issued no dedicated AI standard, and Canada has no federal AI act, AIDA having died on the Order Paper in January 2025. The full instrument-by-instrument position, dated and sourced, is at the regulatory tables and by jurisdiction at the GCC hub and Canada.

§ 06Stated limits

What this page does not claim.

Read this before you cite the page

  • The term "agent governance stack" is in general use and is used here descriptively. Nothing on this page claims it was coined here, and § 04 names the adjacent work so the comparison can be checked.
  • The eleven stages are this practice’s ordering, not a standard. They extend NIST AI RMF and ISO/IEC 42001 rather than replacing either, and an institution already running a different model should map to it rather than restart.
  • The mechanisms in the fourth column are published, but four of them (capability contracts, trust tiers, PEVG, PARA) are specified in books released through September 2026. Cite them as new work rather than as established references.
  • Regulatory dates are current as of the date in the title block and are maintained at the regulatory tables. Check there rather than quoting this page in six months.
§ 07Questions

What readers ask first.

What is an agent governance stack?

An agent governance stack is the ordered set of controls that decide what an AI agent is, what it may do, what constrains it, what it must record, and who is accountable when it acts. The eleven stages used here are identity, purpose, authority, capability, policy, memory, evidence, action, verification, escalation and audit. The order is the order in which things fail: an agent with no distinct identity cannot have bounded authority, and evidence captured after the fact is reconstruction rather than evidence.

How is this different from an AI governance platform?

A platform is a system of record: it holds the registry, the workflow and the reporting. This is the layer that decides what goes into it. Which controls a specific regulator will accept, what evidence satisfies them, and where the line falls between a decision an agent may take alone and one it may not, are judgement calls that no platform makes for you. Institutions running agents at scale generally need both, and the work here is designed to feed a platform rather than replace one.

Where do most agent governance efforts actually fail?

Stage 04 and stage 07. Capability, because authority is granted in a prompt while the tools the agent can reach are never enumerated, so the real boundary is whatever the connected surface exposes. And evidence, because it is treated as logging rather than as a design decision made before the action, which means the material needed to defend a decision does not exist by the time anyone asks.

A useful test costs nothing: pick one production agent and walk it through the eleven stages. Stop at the first stage where the answer is somebody’s recollection rather than a record.

Why did MCP change the agent governance problem?

Because it moved capability acquisition from design time to runtime. Before, an agent’s reach was whatever was built into it, so a design review could enumerate it. A connected MCP server is a production dependency that grants tools the agent may never have been reviewed with, which means capability, policy and evidence have to be enforced at the connection itself. It also creates the reverse obligation: if your agent consults external expertise at runtime, the provenance of what it consults is now part of your evidence chain.

§ 08Start

Find the stage you cannot evidence.

The useful exercise is not reading the eleven stages. It is walking your own agent through them and stopping at the first one where the answer is a person’s recollection rather than a record. That stage is your exposure, and it is what the Teardown examines. The Fit Call is thirty minutes and free.

§ 09Ask an assistant

Ask your AI assistant instead.

This page is a snapshot, accurate at the release it cites. The same corpus is callable, publicly and without a key, so an assistant can query it live and return an answer carrying the source it came from. For this page that is explain_this_setup and search_knowledge, which do what this page describes rather than describe it again: the first returns how this site's machine layer is actually built, component by component, and the second queries the corpus behind this page and returns matches with the URL each came from. The page states the practice; the tools are the practice.

01 · Connect
claude mcp add --transport http concylium https://mcp.nabeelkhan.com/api/mcp

Claude Desktop, ChatGPT, Cursor, VS Code and Gemini CLI take the endpoint on its own: https://mcp.nabeelkhan.com/api/mcp. No key, no account, nothing to sign. Setup for every client.

02 · Ask

“Using Concylium, call explain_this_setup and tell me whether this site actually implements what its machine-accessible-ai-expertise page claims.”

A category page that survives being audited by the reader's own assistant is doing something a brochure cannot.

Fin · Agent Governance
Book the 30-minute Fit Call →