All ArticlesRuntime Truth
Guardrails
Buyer's Guide
AI Agent Security

AI Agent Compliance Checklist for Enterprises

A four-stage control set that proves every AI agent in your enterprise has an owner, least-privilege access, runtime oversight, and audit-ready evidence.

Obsidian Editorial Team
Security Research
·
Obsidian Security
·
July 30, 2026
Key Takeaways

An AI agent compliance checklist runs in four stages, applied in order: Discover, Govern, Enforce, Evidence. You cannot govern what you have not found, and you cannot prove what you never logged. - Configuration snapshots are not evidence. Auditors and boards want runtime truth: what the agent actually did, on whose behalf, and against which data. - Non-human identities now outnumber people, and an agent moves roughly 16 times more data than a human does. A compliance program that only covers humans is already out of scope. - Runtime enforcement today means Claude and Microsoft Copilot. Agentforce, Bedrock, and n8n are discover and govern only, with enforcement on the roadmap. State that honestly in any regulator response. - Every control should map to an agent's effective access inside the connected app: which service account, which objects, which actions. Not the vendor's theoretical configuration page. - Maker mode, orphaned agents, and shadow connections are the first three gaps an auditor probes. Roughly nine in ten agents are over-permissioned when first assessed.

How to Use This AI Agent Compliance Checklist

An AI agent compliance checklist is an operational control set, not a policy PDF. Each item has an owner, a data source, and a pass or fail state. This AI agent compliance checklist is built for CISOs, GRC leads, and security architects who have to answer a direct question from a regulator or a board: what are these agents doing, and can you prove it. It works in four stages, and the order matters.

Most teams inherit a program that already has hundreds or thousands of agents running with no clear ownership. Discover first, because you cannot govern what you cannot see. Govern next, because policy without an inventory is fiction. Enforce third, because detection without prevention is expensive logging. Evidence last, because that trail is what turns your work into something an auditor accepts.

Treat each stage as a gate. Do not advance to Govern until Discover is complete for your top three AI platforms. Run the whole set quarterly, because agent sprawl tends to double faster than an annual review cycle can track.

The design principle underneath every stage: correlate an agent's runtime activity to its effective access inside the third-party app, not to the config page the platform vendor shows you. A checkbox marked "read-only" in a console means little if the agent's service account can write to a revenue object. Runtime truth is the standard the rest of this checklist enforces.

Discover: Inventory Every Agent and Its Access

Discovery is the foundation. An agent you have not found is a machine insider with credentials and no supervision.

  • [ ] Inventory every AI agent across Copilot Studio, Agentforce, Bedrock, Vertex, Azure AI Foundry, n8n, ChatGPT Enterprise, and Claude, and record a single ID for each.
  • [ ] Attribute each agent to a creator and a current owner. Flag any agent whose owner account is disabled or has left the company.
  • [ ] Map every connection the agent uses, sanctioned and unsanctioned, including MCP server connections.
  • [ ] Detect shadow AI: personal ChatGPT or Claude accounts moving corporate data through unmanaged browsers.
  • [ ] Record each agent's effective access inside every connected app, which service account it uses and which objects, tables, and actions that account can touch. This is the difference between effective access and theoretical configuration.
  • [ ] Log any agent set to org-wide or anonymous public access.
  • [ ] Identify agents built in maker mode with embedded creator credentials.

Each of these is verifiable. You either have the inventory record or you do not. If you find agents you cannot attribute to a named business owner within 30 days, quarantine or disable them.

Govern: Policies, Ownership, and Least Privilege

Governance turns the inventory into accountability. Every item here is a control you can check against a system of record.

  • [ ] Every agent has a named business owner and a named security reviewer.
  • [ ] Every agent has a documented purpose and a data classification level.
  • [ ] Agents touching sensitive data (PII, PHI, financials, source code) are tagged and reviewed on a fixed cadence.
  • [ ] Least privilege is enforced against effective access. If an agent's job needs read-only, its service account does not hold write.
  • [ ] Toxic combinations are documented and rated. A shadow agent with org-wide access to sensitive data is a critical finding, because its blast radius is the whole tenant.
  • [ ] Maker mode is disallowed for agents touching regulated data, or the exception is justified in writing.
  • [ ] Agent-to-agent paths are mapped, especially any path where a low-privilege agent can invoke a higher-privilege one.
  • [ ] A published acceptable-use policy names sanctioned platforms and prohibited data types.

Least privilege here is a statement about non-human identity. The agent is the identity, and its permissions are the ones you review, not the human who happened to build it.

Enforce: Runtime Controls and Remediation

Enforcement is where most programs quietly fail. Detection without prevention is documentation of failure, not compliance. Be precise about what your tooling can actually do today, because a regulator will hold you to it.

Truth in labeling on runtime enforcement. Autonomous runtime blocking today applies to Claude and Microsoft Copilot only. For Salesforce Agentforce, Amazon Bedrock, and n8n, current capability is discover and govern only. Runtime enforcement on those platforms is on the roadmap and must never be described as available today in a regulator response, an audit workpaper, or a board slide.

A note on how the runtime control works, because auditors ask. Obsidian does not sit inline as a gateway or proxy in the request path. It runs as an "Intel Inside" decision engine that the platform consults at runtime, so the enforcement decision is grounded in the agent's effective access rather than a static config snapshot.

  • [ ] Runtime monitoring is active on Claude and Copilot agents, with deterministic guardrails that block privilege escalation and unauthorized tool calls as they happen.
  • [ ] For Agentforce, Bedrock, and n8n, continuous posture review runs on a fixed cadence with alerts on high-risk configuration drift. Enforcement here is roadmap, so remediation is human-driven today.
  • [ ] Orphaned agents are disabled within a defined SLA (a 7-day target is reasonable).
  • [ ] Confused-deputy cases, where a user invokes an agent to reach data they personally lack rights to, are detected and alerted.
  • [ ] Every high-severity alert opens a ticket with an owner and an SLA through your existing case system.
  • [ ] Incident response runbooks include agent scenarios: token compromise, maker-mode abuse, and data exfiltration through crafted prompts.

Evidence: Logging and Audit Trail

Auditors do not accept "we have a policy." They want a record. The evidence stage produces the runtime truth that the other three stages exist to generate, so it should require almost no extra effort once the pipeline is running.

  • [ ] Every agent action against a connected system is logged with four fields: the invoking identity, the agent identity, the tool called, and the data accessed.
  • [ ] Logs are retained to meet your regulatory obligation. Confirm the retention window with your GRC team rather than guessing.
  • [ ] Board reporting shows agent count, risk distribution, the top toxic combinations, and remediation SLAs against target.
  • [ ] Access reviews cover non-human identities on the same cadence as human ones, not as an afterthought.
  • [ ] Every control in this checklist maps to at least one framework obligation.
  • [ ] Change history is preserved: who modified an agent, when, and why.

Configuration snapshots will not satisfy the oversight and record-keeping expectations these frameworks set out. A snapshot says the door could be locked. A runtime log says whether anyone walked through it.

Framework Mapping and Common Gaps

Mapping Controls to the EU AI Act, NIST AI RMF, and ISO 42001

Auditors will ask which control answers which obligation. Map at the level of documented themes, and confirm exact clause references with your GRC team against the current text of each framework rather than citing numbers from memory.

Control area EU AI Act theme NIST AI RMF function ISO 42001 theme
Agent inventory and record-keeping Record-keeping and logging for high-risk systems Map, Govern Planning and operation of the AI management system
Owner accountability Deployer obligations Govern Leadership and assigned roles
Least privilege and effective access Risk management for high-risk systems Manage Operational controls
Runtime oversight Human oversight Measure Performance evaluation
Logging and audit trail Record-keeping and transparency Measure Monitoring and measurement
Incident response Serious-incident reporting Manage Improvement

The point of the map is not to win a framework trivia contest. It is to show that a single runtime control, correlating an agent to its effective access, satisfies several obligations at once. That is what lets one control plane replace a stack of point checks and cut the number of fire drills.

Common Gaps This AI Agent Compliance Checklist Catches

The same gaps repeat across enterprise reviews:

  • Maker-mode agents touching CRM or financial data with the creator's embedded admin credentials. Any invoker inherits those rights, and standard IAM is bypassed.
  • Orphaned agents whose creator has left, yet the agent still runs, still holds tokens, and still moves data.
  • Shadow connections wired into sanctioned agents but never approved by security, including unreviewed MCP servers.
  • Public agent URLs with anonymous access, most often on ChatGPT Enterprise and Copilot Studio.
  • Agent-to-agent chaining, where a limited agent calls a broader one to reach data neither should hold alone.
  • Personal AI accounts in managed browsers uploading corporate documents to services you do not control.

These are also the gaps behind real-world exposure. The campaign tracked as UNC6395, tied to compromised Salesloft Drift integrations, reached roughly 700 organizations by riding the effective access of a trusted non-human identity. That is the machine-insider pattern this checklist is built to catch before it becomes an incident.

Frequently Asked Questions

What is an AI agent compliance checklist?

It is an operational control set that proves every AI agent in your enterprise has a known owner, least-privilege access measured against effective access, runtime oversight where the platform supports it, and an audit-ready log. It differs from an AI policy because each item is a verifiable pass or fail state, not a statement of intent.

Which platforms support runtime enforcement today?

Autonomous runtime blocking is available for Claude and Microsoft Copilot. For Agentforce, Bedrock, and n8n, coverage today is discover and govern only, with enforcement on the roadmap. Do not represent runtime blocking on those platforms as present-tense capability in any regulator or audit response.

What is maker mode and why does it matter for compliance?

Maker mode means an agent runs with its creator's embedded credentials. Any user who invokes the agent inherits those privileges, which bypasses standard IAM and expands the agent's blast radius well beyond its stated purpose. It is the privilege-escalation pattern auditors ask about first.

How do we prove runtime behavior to an auditor?

You need logs that correlate the invoking user, the agent identity, the tool called, and the data accessed, tied back to the agent's effective access inside the connected app. Configuration snapshots describe what could happen. Runtime truth describes what did.

Can one checklist cover Copilot, Agentforce, and Bedrock at once?

The four stages stay the same across platforms. What changes is the specific risk factors per platform and, critically, whether runtime enforcement is available or still on the roadmap. Keep the stages constant and vary the enforcement expectation by platform.

What is a toxic combination in AI agent compliance?

It is several risk factors on one agent that are moderate alone but critical together. A common example is an unmanaged shadow agent that also holds org-wide access to sensitive data, because the combination gives a machine insider both stealth and reach.