Model Risk Management for AI in Financial Services
Model risk management for AI extends the SR 11-7 and OCC 2011-12 framework, originally built for statistical and quantitative models, to machine learning, generative AI, and autonomous agent systems. It requires the same three pillars, development and use, validation, and governance, applied to systems that are harder to interpret, subject to continuous change, and (in the case of agents) capable of taking actions rather than only producing outputs.
AI Model Risk Management Framework
SR 11-7's structure maps onto AI systems across four practical areas of focus, from initial development through the continuous oversight that agentic systems require in production.
Development and Use
SR 11-7's expectations for robust design and appropriate use, extended to machine learning and generative AI systems.
Validation
Conceptual soundness review, monitoring, and outcomes analysis applied to less interpretable, non-deterministic models.
Governance and Controls
Board and management accountability, policy, and independent effective challenge over AI systems.
Runtime Monitoring
Continuous oversight of drift, output quality, and agent actions in production, beyond periodic review cycles.
Core Components of an AI-Specific MRM Framework
The following elements recur across supervisory expectations when existing model risk guidance is extended to AI, machine learning, and autonomous agent systems.
- Robust development, implementation, and appropriate use of AI/ML models, consistent with SR 11-7's original expectations for model design.
- Independent validation covering conceptual soundness, ongoing monitoring, and outcomes analysis, adapted for less interpretable and non-deterministic systems.
- Governance anchored by effective challenge: critical review by independent parties with the authority and expertise to identify limitations and require changes.
- Continuous, rather than periodic, drift monitoring for models that are retrained frequently or operate on live data streams.
- Oversight of vendor and third-party AI models, since external validation evidence alone does not satisfy institutional validation obligations.
- Runtime monitoring for autonomous agents that tracks executed actions, not just generated outputs, including anomaly detection and structured audit logging.
- Separation of output validation from action authorization, with distinct approval thresholds for recommendations versus executed actions.
- Documentation of the institutional interpretive basis used to extend existing guidance to generative AI and autonomous agents.
SR 11-7 as the Foundation for AI Model Risk
SR 11-7, issued by the Federal Reserve Board on April 4, 2011, defines a model broadly as a quantitative method, system, or approach that applies statistical, economic, or mathematical theories to process input data into quantitative estimates. OCC Bulletin 2011-12, issued the same day, adopts this guidance for national banks and federal savings associations. Under this broad definition, machine learning systems, generative AI applications, and autonomous agents built on large language models fall within scope, even though neither document was written with these architectures in mind.
SR 11-7 organizes model risk management around three components: robust development, implementation, and use; ongoing validation; and governance, policy, and controls, anchored by the principle of effective challenge. Effective challenge means critical review by parties independent of model development who have the authority and expertise to identify limitations, assumptions, and required changes. No supervisory guidance issued since 2011 has superseded this structure for AI systems. Institutions are expected to extend SR 11-7 through interpretation, not apply a separate AI-specific rule.
Where AI Systems Diverge from Traditional Model Assumptions
SR 11-7's validation triad, conceptual soundness review, ongoing monitoring, and outcomes analysis such as backtesting, assumes a relatively static and interpretable model. Complex machine learning architectures and large language models complicate conceptual soundness review because their internal logic is less transparent than traditional statistical models.
Generative AI outputs are also probabilistic and non-deterministic, which makes traditional backtesting and outcomes analysis harder to apply directly without adapted evaluation methods. Drift monitoring, already a core SR 11-7 expectation, requires continuous rather than periodic assessment when models are retrained frequently or operate on live data streams. Vendor and third-party AI models increase reliance on external validation evidence, something SR 11-7 already identifies as insufficient on its own for satisfying institutional validation obligations. Finally, autonomous agents that take actions rather than only generate outputs introduce behavioral risk that falls outside the output-focused validation scope originally described in SR 11-7.
Runtime Monitoring and Auditability for AI Agents
Autonomous agents that execute transactions or take other actions in production require monitoring pipelines that operate continuously rather than at fixed validation intervals. This includes detecting drift, flagging anomalous outputs, and tracking the specific actions an agent takes, not just the content it generates. Logging and audit trails need to be structured so that agent decisions and actions can support post-hoc effective challenge and regulatory examination, consistent with SR 11-7's independent review expectations.
Governance controls should also separate output validation from action authorization. An agent's recommendation and an agent's executed action carry different risk profiles and may warrant different approval thresholds. Runtime governance capabilities, including agent identity, permission enforcement, least privilege access, tool approval workflows, and structured audit logging, are the kinds of controls that support this distinction in production environments. Trussed AI provides runtime governance and security for enterprise AI agents built around these control categories, intended to help institutions maintain auditability for agentic systems operating under existing MRM obligations.
Governance and Accountability Considerations
Board and senior management oversight responsibilities under SR 11-7 extend to AI systems, requiring clear accountability for AI model approval and ongoing risk acceptance. The independent validation function must retain sufficient authority and expertise to challenge AI/ML models, consistent with SR 11-7's effective challenge principle, even as model complexity increases.
Institutions remain accountable for third-party or vendor AI model risk even when transparency into model internals is limited, per SR 11-7's existing vendor model guidance. No current US regulation formally replaces SR 11-7 or OCC 2011-12 for AI systems. Institutions should document the interpretive basis they use to extend existing guidance to generative AI and autonomous agents, since this remains a matter of institutional judgment rather than a settled regulatory mapping.
Implementation Considerations for AI-Specific MRM Programs
Extending SR 11-7 and OCC 2011-12 to AI systems is a matter of institutional judgment. The following considerations reflect the governance and accountability expectations already embedded in existing supervisory guidance.
- Establish clear board and senior management accountability for AI model approval and ongoing risk acceptance.
- Ensure the independent validation function retains sufficient authority and expertise to challenge AI/ML models under the effective challenge principle.
- Maintain accountability for third-party and vendor AI model risk, consistent with SR 11-7's existing vendor model guidance, even where internal transparency is limited.
- Document the institutional interpretive basis used to extend SR 11-7 and OCC 2011-12 to generative AI and autonomous agent systems.
Extend Model Risk Management to AI Agents in Production
Runtime governance for AI agents supports the monitoring, identity, permissioning, and audit logging needed to maintain SR 11-7-aligned oversight as AI systems move into production.
Explore Runtime Governance