A four-stage control set that proves every AI agent in your enterprise has an owner, least-privilege access, runtime oversight, and audit-ready evidence.
An AI agent compliance checklist runs in four stages, applied in order: Discover, Govern, Enforce, Evidence. You cannot govern what you have not found, and you cannot prove what you never logged. - Configuration snapshots are not evidence. Auditors and boards want runtime truth: what the agent actually did, on whose behalf, and against which data. - Non-human identities now outnumber people, and an agent moves roughly 16 times more data than a human does. A compliance program that only covers humans is already out of scope. - Runtime enforcement today means Claude and Microsoft Copilot. Agentforce, Bedrock, and n8n are discover and govern only, with enforcement on the roadmap. State that honestly in any regulator response. - Every control should map to an agent's effective access inside the connected app: which service account, which objects, which actions. Not the vendor's theoretical configuration page. - Maker mode, orphaned agents, and shadow connections are the first three gaps an auditor probes. Roughly nine in ten agents are over-permissioned when first assessed.
An AI agent compliance checklist is an operational control set, not a policy PDF. Each item has an owner, a data source, and a pass or fail state. This AI agent compliance checklist is built for CISOs, GRC leads, and security architects who have to answer a direct question from a regulator or a board: what are these agents doing, and can you prove it. It works in four stages, and the order matters.
Most teams inherit a program that already has hundreds or thousands of agents running with no clear ownership. Discover first, because you cannot govern what you cannot see. Govern next, because policy without an inventory is fiction. Enforce third, because detection without prevention is expensive logging. Evidence last, because that trail is what turns your work into something an auditor accepts.
Treat each stage as a gate. Do not advance to Govern until Discover is complete for your top three AI platforms. Run the whole set quarterly, because agent sprawl tends to double faster than an annual review cycle can track.
The design principle underneath every stage: correlate an agent's runtime activity to its effective access inside the third-party app, not to the config page the platform vendor shows you. A checkbox marked "read-only" in a console means little if the agent's service account can write to a revenue object. Runtime truth is the standard the rest of this checklist enforces.
Discovery is the foundation. An agent you have not found is a machine insider with credentials and no supervision.
Each of these is verifiable. You either have the inventory record or you do not. If you find agents you cannot attribute to a named business owner within 30 days, quarantine or disable them.
Governance turns the inventory into accountability. Every item here is a control you can check against a system of record.
Least privilege here is a statement about non-human identity. The agent is the identity, and its permissions are the ones you review, not the human who happened to build it.
Enforcement is where most programs quietly fail. Detection without prevention is documentation of failure, not compliance. Be precise about what your tooling can actually do today, because a regulator will hold you to it.
Truth in labeling on runtime enforcement. Autonomous runtime blocking today applies to Claude and Microsoft Copilot only. For Salesforce Agentforce, Amazon Bedrock, and n8n, current capability is discover and govern only. Runtime enforcement on those platforms is on the roadmap and must never be described as available today in a regulator response, an audit workpaper, or a board slide.
A note on how the runtime control works, because auditors ask. Obsidian does not sit inline as a gateway or proxy in the request path. It runs as an "Intel Inside" decision engine that the platform consults at runtime, so the enforcement decision is grounded in the agent's effective access rather than a static config snapshot.
Auditors do not accept "we have a policy." They want a record. The evidence stage produces the runtime truth that the other three stages exist to generate, so it should require almost no extra effort once the pipeline is running.
Configuration snapshots will not satisfy the oversight and record-keeping expectations these frameworks set out. A snapshot says the door could be locked. A runtime log says whether anyone walked through it.
Auditors will ask which control answers which obligation. Map at the level of documented themes, and confirm exact clause references with your GRC team against the current text of each framework rather than citing numbers from memory.
| Control area | EU AI Act theme | NIST AI RMF function | ISO 42001 theme |
|---|---|---|---|
| Agent inventory and record-keeping | Record-keeping and logging for high-risk systems | Map, Govern | Planning and operation of the AI management system |
| Owner accountability | Deployer obligations | Govern | Leadership and assigned roles |
| Least privilege and effective access | Risk management for high-risk systems | Manage | Operational controls |
| Runtime oversight | Human oversight | Measure | Performance evaluation |
| Logging and audit trail | Record-keeping and transparency | Measure | Monitoring and measurement |
| Incident response | Serious-incident reporting | Manage | Improvement |
The point of the map is not to win a framework trivia contest. It is to show that a single runtime control, correlating an agent to its effective access, satisfies several obligations at once. That is what lets one control plane replace a stack of point checks and cut the number of fire drills.
The same gaps repeat across enterprise reviews:
These are also the gaps behind real-world exposure. The campaign tracked as UNC6395, tied to compromised Salesloft Drift integrations, reached roughly 700 organizations by riding the effective access of a trusted non-human identity. That is the machine-insider pattern this checklist is built to catch before it becomes an incident.
It is an operational control set that proves every AI agent in your enterprise has a known owner, least-privilege access measured against effective access, runtime oversight where the platform supports it, and an audit-ready log. It differs from an AI policy because each item is a verifiable pass or fail state, not a statement of intent.
Autonomous runtime blocking is available for Claude and Microsoft Copilot. For Agentforce, Bedrock, and n8n, coverage today is discover and govern only, with enforcement on the roadmap. Do not represent runtime blocking on those platforms as present-tense capability in any regulator or audit response.
Maker mode means an agent runs with its creator's embedded credentials. Any user who invokes the agent inherits those privileges, which bypasses standard IAM and expands the agent's blast radius well beyond its stated purpose. It is the privilege-escalation pattern auditors ask about first.
You need logs that correlate the invoking user, the agent identity, the tool called, and the data accessed, tied back to the agent's effective access inside the connected app. Configuration snapshots describe what could happen. Runtime truth describes what did.
The four stages stay the same across platforms. What changes is the specific risk factors per platform and, critically, whether runtime enforcement is available or still on the roadmap. Keep the stages constant and vary the enforcement expectation by platform.
It is several risk factors on one agent that are moderate alone but critical together. A common example is an unmanaged shadow agent that also holds org-wide access to sensitive data, because the combination gives a machine insider both stealth and reach.