AI Agent Governance Frameworks
AI agent governance frameworks are the standards organizations use to demonstrate that their AI systems are controlled: principally the EU AI Act, NIST AI RMF, ISO 42001, and the OWASP Agentic AI Top 10. None was written with autonomous agents in mind. All of them require evidence of what a system actually did, which is the requirement agents make hard.
Which frameworks actually apply to AI agents?
Four matter for most enterprises, and they do different jobs. One is law, one is a voluntary structuring device, one is a certifiable standard procurement teams ask for by name, and one is a threat taxonomy rather than a governance framework at all.
| Framework | Applies to | Status | What it asks of agent operators |
|---|---|---|---|
| EU AI Act | Organizations placing AI systems on the EU market or whose output is used in the EU, wherever they are based | Binding law, with obligations phased in and penalties attached | Classify each system by risk tier, then meet the obligations for that tier. At higher tiers, that includes record-keeping and the ability to explain how a system reached an outcome |
| NIST AI RMF | Any organization; US-origin but widely adopted outside it | Voluntary, non-certifiable | Organize AI risk work around four functions (govern, map, measure, manage) and evidence that each is being performed |
| ISO/IEC 42001 | Any organization; international standard | Voluntary but certifiable by an accredited body | Operate a documented AI management system, with defined roles, controls and internal audit, sustained over time rather than demonstrated once |
| OWASP Agentic AI Top 10 | Any organization building or running agents | Community guidance, advisory only | Test against a named set of agent-specific failure modes; the only one of the four written for agents rather than adapted to them |
The practical consequence of that last row is worth sitting with. The three frameworks carrying regulatory or contractual weight were designed for AI systems that produce outputs a human then acts on. Agents act directly. The frameworks still apply, and the evidence they ask for has to be assembled differently.
A note on what is deliberately absent from the table. Sector regulation (financial services rules, healthcare and medical device regimes, data protection law) frequently applies to agents and is not superseded by any of the above. Teams already operating under those obligations should treat the four frameworks here as additional structure rather than a replacement for what they already meet.
What do these frameworks have in common?
More than their differences suggest, and this is the useful observation for anyone mapping several at once. Each one, in its own vocabulary, asks an organization to show four things: which AI systems it operates, what those systems are permitted to do, what they actually did, and who is accountable for each.
That commonality is why framework-by-framework compliance projects tend to be wasteful. The evidence underneath is largely the same evidence. What differs is the classification scheme applied on top, the documentation format, and the enforcement mechanism. Build the evidence layer once and map it repeatedly, rather than running three parallel programs that each reconstruct the same underlying record.
What makes agents harder to evidence than earlier AI systems?
A model that produces a recommendation leaves a human decision in the record. Someone read the output and chose to act, and that choice is captured in whatever system the human was working in. The audit trail exists independently of the AI.
An agent that takes an action directly leaves no such trace. The system’s own record becomes the only account of what happened, which means the quality of that record determines whether the organization can evidence anything at all. This is the point at which governance stops being a documentation exercise and becomes an instrumentation problem. It is also why agent audit trails are a compliance dependency rather than an operational nicety.
Teams already operating under financial services regulation tend to describe the novelty precisely. Most of what the EU AI Act asks for maps onto obligations they already meet elsewhere. The genuinely new requirement is explaining what an agent did and why, and existing compliance programs have no equivalent of it.
Can you comply with a framework without knowing how many agents you have?
No, and this is the sequencing error that most agent governance programs make. Every framework above presumes an accurate inventory as its starting condition. An organization that cannot enumerate its agents is not partially compliant; it has no foundation on which any of the subsequent controls can be assessed.
The gap between assumed and actual inventory is typically large. Within a single proof of concept, Geordie identified 327% more agents at Owkin than existing inventories had captured. A risk classification exercise run against the original inventory would have been thorough, well-documented, and wrong about the majority of the estate.
Compliance is downstream of discovery. Any program that begins with policy drafting rather than with finding the agents is producing documentation rather than assurance.
What evidence should an agent governance program actually produce?
Five artifacts cover the substantive requirements across all four frameworks, whatever vocabulary each uses.
A current inventory of agents, continuously maintained rather than periodically refreshed, since agents are created by people entitled to create them and by other agents. A record of what each agent is permitted to reach, covering the full tool supply chain rather than only the connections that pass through a managed chokepoint. A behavioral record of what each agent actually did, in a form that survives the session. Named ownership for every agent, so that findings have an addressee. And a record of what was blocked or redirected, and on what basis, because demonstrating that controls exist is a different claim from demonstrating that they fired.
The last one is routinely missed. Organizations document their policies and then cannot show a single instance of a policy taking effect.
Does compliance work slow agent adoption down?
It does when it is run as a gate, and the evidence suggests it need not be. At Owkin, EU AI Act alignment was demonstrated rapidly during pharmaceutical partner procurement. Compliance functioned as a commercial unblocker rather than a brake on the program.
This is the more useful way to frame the investment. An organization that can evidence what its agents do can say yes to agent adoption faster, because the alternative to evidence is caution. The frameworks are not the obstacle. The absence of a record is.
Common questions
Is ISO 42001 certification necessary?
Necessary for nobody, valuable for organizations selling into enterprises that have started asking for it in procurement. Its practical advantage over NIST AI RMF is that it is certifiable, which makes it an answer to a supplier questionnaire rather than an internal structuring device.
Does the EU AI Act apply to agents used internally?
It applies based on the system's risk classification and the role the organization plays, not on whether the system is customer-facing. Internal agents touching employment decisions, creditworthiness, or other classified areas carry obligations. Internal agents writing code generally do not, which is a distinction worth establishing early rather than assuming either way across the whole estate.
Where does the OWASP Agentic AI Top 10 fit?
As a threat model feeding the measure and manage functions of the others, not as a compliance standard in its own right. It is the most agent-specific document of the four and the most useful for deciding what to test for.
Is a policy document enough to demonstrate governance?
No. Every framework above distinguishes between having a control and evidencing that it operated. A policy states intent; an audit trail states outcome. Programs that produce only the first tend to discover the gap during their first real audit.