Evaluating AI Agent Security in a SOC 2 Type II Report
A SOC 2 Type II report can provide partial assurance over the infrastructure hosting AI agents, but it was not built to test agent-specific risks such as dynamic permissions, autonomous tool calls, or non-human identity lifecycle management. Reviewers must check whether the system description explicitly includes AI agent components and whether control testing addresses runtime agent behavior, not just static IT controls.
What a SOC 2 Report Was Built to Test
Trust Services CriteriaSecurity, availability, processing integrity, confidentiality, and privacy: none written with autonomous agents in mind.
Common Criteria SeriesCC6 access controls and CC8 change management, mapped to organization-defined activities, not agent runtime behavior.
The Coverage GapNo AICPA-issued category addresses dynamic permissions, tool-call sequences, or non-human identity lifecycle.
Buyer Questions for Reviewing a SOC 2 Report
- Does the system description explicitly identify AI agents or tool-integration platforms within the audited boundary?
- What control activities test dynamic or session-based permission grants issued to agents, as opposed to static human roles?
- Are agent-initiated configuration or policy changes included in change management testing and logging?
- Does monitoring capture tool-call sequences or anomalous autonomous actions distinct from standard application logs?
- How is the lifecycle of non-human identities associated with agents provisioned, rotated, and deprovisioned?
- Does the report reference any risk framework beyond AICPA criteria, and how was it incorporated into testing?
Full Article
SOC 2 Type II reports are structured around the AICPA's Trust Services Criteria: security, availability, processing integrity, confidentiality, and privacy. These criteria, along with the underlying Common Criteria series (CC1 through CC9), predate the widespread deployment of autonomous AI agents. Auditors test whether an organization's self-defined controls operate effectively over a review period, typically six to twelve months. This means a SOC 2 report is only as relevant to AI agent risk as the control descriptions the organization chose to include in scope. If the system description does not mention AI agents, tool integrations, or autonomous decision-making, the audit almost certainly did not test for those risks, regardless of how comprehensive the report appears on its face. Compliance leaders reviewing a vendor's or an internal system's SOC 2 report need to treat this as a starting point for inquiry, not a conclusion about agent security posture.
Mapping Trust Services Criteria to Agent Runtime Behavior
Two Common Criteria families carry the most weight when evaluating AI agent coverage: CC6 (logical and physical access controls) and CC8 (change management). CC6 controls are typically written around human user provisioning, deprovisioning, and role-based access. They do not natively address permissions that are generated dynamically for an agent session, scoped to a single task, or expire after a tool call completes. CC8 controls address code and configuration changes, but auditors rarely extend that testing to changes in agent prompts, policy configurations, or tool permission scopes, even though these function as runtime configuration changes with security implications. CC7, the monitoring criterion, is also relevant: it generally tests for anomaly detection and incident response processes, but reviewers should not assume this extends to detecting anomalous tool-call sequences or unauthorized autonomous actions unless the control language says so explicitly.
What to Look for in Control Descriptions and Testing Procedures
When reading a SOC 2 Type II report for an organization that operates AI agents, look for specific language rather than general assurances. Does the system description name AI agents, autonomous workflows, or tool-integration platforms as part of the audited system boundary? Are there discrete control activities for issuing and revoking dynamic or session-based permissions to agents, separate from static human user roles? Does change management testing include agent-initiated configuration or policy changes, and are those changes logged and reviewed the same way code changes are? Reports that only describe generic identity and access management, without acknowledging non-human or agent identities as a distinct category, are unlikely to have tested for the specific risks that matter in agentic systems.
Where Standard Frameworks Leave Gaps
Supplementary frameworks help identify what a SOC 2 report is unlikely to cover. NIST's AI Risk Management Framework separates governance, mapping, measuring, and managing AI risk from traditional IT control testing, and NIST's generative AI profile guidance identifies risks such as unauthorized tool or plugin use that fall outside standard access and change management criteria. OWASP's Top 10 for LLM Applications similarly treats excessive agency and insecure plugin or tool design as distinct risk categories, separate from the access control weaknesses SOC 2's CC6 was designed to catch. None of these frameworks are audit standards, and none have been formally incorporated into AICPA guidance. They are useful as a checklist for gaps, not as a substitute for what a SOC 2 report actually tests.
Tool-Use Permissioning and the Model Context Protocol
As organizations adopt protocols like the Model Context Protocol (MCP) to connect AI models to external tools and data sources, a new permissioning layer emerges that sits outside traditional application access control boundaries. MCP's specification defines an authorization framework built on OAuth-based patterns for tool servers, which means tool-call permissioning has its own scope-enforcement mechanisms that a SOC 2 audit would need to explicitly test to provide assurance over it. Reviewers should ask whether MCP servers or equivalent tool-integration layers are named in the system description, and whether the organization can produce evidence of how tool access scopes are granted, enforced, and revoked. In the absence of that evidence, assume the audit did not evaluate this layer.
Treating SOC 2 as One Input, Not a Complete Answer
SOC 2 Type II provides assurance only over controls the organization itself defined and included in scope. Because no AICPA-issued supplemental guidance for AI agents currently exists, control language addressing agent identity, dynamic permissions, or tool-use governance varies significantly by auditor and by organization, which reduces comparability across reports. Compliance leaders should treat a SOC 2 report as one input among several when evaluating an agentic AI system, alongside architecture review, penetration testing, and an AI-specific risk assessment that directly examines runtime behavior. This is where runtime policy enforcement and agent identity controls become the practical complement to a compliance report: they generate the evidence, such as audit logs of tool approvals and permission grants, that a SOC 2 report is unlikely to produce on its own. Trussed AI supports this layer through runtime governance, agent identity and permissions management, and audit logging designed for AI agent environments, giving compliance teams evidence that maps directly to the questions a SOC 2 report leaves open.
Close the Gap Between SOC 2 and Agent Runtime Risk
See how runtime governance and audit logging generate the evidence traditional compliance reports were not built to capture.
Learn About AI Agent Security