AI Governance Evidence Pack for SOC 2 Audit
A SOC 2-ready AI governance evidence pack maps agent identity, permission scoping, tool-call logs, and policy enforcement records to the specific Trust Services Criteria they support, primarily Security and Processing Integrity, since no AI-specific criterion currently exists. Organizations must document this mapping explicitly and structure evidence by TSC category rather than by system, because auditors request evidence in that format and will not infer the connection between agent runtime behavior and control objectives on their own.
Field Checklist Before You Start Collecting Evidence
- Confirm with the auditor in advance which Trust Services Criteria will be applied to AI agent controls, since this is not standardized across engagements.
- Document agent identity provisioning and de-provisioning using the same structure applied to human user access control evidence.
- Capture policy enforcement decisions at runtime, not just policy configuration, to demonstrate least-privilege was actually applied.
- Define log retention periods and integrity safeguards before evidence collection begins, not during fieldwork.
- Separate AI governance framework documentation, such as NIST AI RMF artifacts, from SOC 2 evidence, and label each explicitly by purpose.
Why AI Agents Create a Structural Gap in SOC 2 Evidence
SOC 2 examinations are governed by the AICPA Trust Services Criteria, which include Security, Availability, Processing Integrity, Confidentiality, and Privacy. The Security criterion, often called the Common Criteria, is mandatory in every SOC 2 report and covers logical access controls, system operations, change management, and monitoring. No AICPA-issued Trust Services Criterion addresses AI agents specifically. AICPA has published non-authoritative guidance discussing how AI-related risks should be considered during SOC examinations, but this guidance does not create a distinct control domain. In practice, this means organizations deploying AI agents must justify how existing criteria apply to agentic behavior rather than rely on a prescriptive AI checklist. That justification is the foundation of a credible evidence pack, and its absence is the most common source of scope disputes during audits.
Mapping Runtime Controls to Trust Services Criteria
The Security criterion requires evidence of logical access security measures, including provisioning, authentication, and authorization over system components. For AI agents, this maps to agent identity issuance, credential lifecycle records, and permission scoping documentation. The Processing Integrity criterion addresses whether system processing is complete, valid, accurate, timely, and authorized, which is directly relevant to automated decision-making outputs produced by agents. AICPA guidance notes that entities using AI systems should document controls over data inputs, model outputs, and monitoring of AI system behavior as part of the broader internal control environment. Rather than creating a parallel AI control framework, evidence should be integrated into existing change management and monitoring documentation, with explicit notes on how each artifact supports a named criterion.
What Runtime Enforcement Logs Need to Contain
Audit log content expectations for automated system components are addressed in NIST SP 800-53's audit and accountability control family, which specifies requirements for log generation, content, retention, and review. Applied to AI agents, policy enforcement decisions should be logged with actor identity, the action attempted, the resource targeted, a timestamp, and the enforcement outcome with rationale for allow or deny decisions. Tool-call logs function as the automated-system analog to traditional user access logs and should be structured so they can be sampled and reviewed for anomalous or unauthorized activity, consistent with the monitoring expectations under the Security criterion. Evidence of least-privilege enforcement requires linking the static policy definition (permitted actions and resources) to the actual runtime enforcement outcome. A policy document alone does not demonstrate enforcement; the log showing the policy was applied at execution time does.
Structuring the Evidence Pack for Auditor Review
Evidence packs should be organized by Trust Services Criteria category rather than by internal system architecture, because auditors issue evidence requests structured around the TSC and will expect responses in that format. A practical structure separates identity and access evidence, change management evidence, and monitoring evidence into discrete sections, with each artifact labeled against the specific criterion it supports. Change management documentation for AI agent policy updates, including who approved a change, when it occurred, and what was modified, should be retained in the same manner as traditional software change control evidence. Log retention periods and integrity controls, including tamper-evidence and immutability where applicable, should be documented separately, since auditors evaluate the reliability of automated evidence sources as part of their overall assessment. Sampling methodology for agent activity logs should also be defined in advance, since SOC 2 auditors typically test a representative sample of logged events rather than reviewing the full record set.
Where AI Governance Frameworks Fit, and Where They Do Not
NIST's AI Risk Management Framework defines four core functions (Govern, Map, Measure, and Manage) intended to help organizations document AI system risk controls. Documentation of system design, data provenance, and monitoring or testing activities produced through these functions can support Processing Integrity evidence, since they demonstrate that system behavior was evaluated and tracked over time. However, NIST AI RMF is voluntary guidance, not an audit standard, so its artifacts support but do not substitute for SOC 2-specific evidence. Organizations should document the mapping between AI RMF activities and SOC 2 control evidence carefully, without implying that RMF alignment equals SOC 2 compliance. This distinction matters during scope negotiations, since conflating the two can create disputes over what has actually been attested.
Core Evidence Categories for AI Agent SOC 2 Readiness
Agent Identity Records
Provisioning and de-provisioning evidence for non-human agent credentials.
Permission Scoping
Documented least-privilege policy definitions tied to actual runtime enforcement.
Tool-Call Logs
Action-level records supporting monitoring and anomaly review requirements.
Policy Enforcement Decisions
Allow/deny outcomes with rationale, mapped to Processing Integrity evidence.
Assemble Audit-Ready Evidence from Runtime Controls
Trussed AI provides runtime governance and enforcement for AI agents, including agent identity, permission scoping, and audit logging that can support the evidence mapping described above.
Explore Runtime Governance