See how Trussed maps to your regulation in minutes

    No generic demo, just the controls relevant to your program.

    Book a session

    Banking / Model Risk Management

    How to Write an AI Model Validation Report for Bank Examiners

    An AI model validation report for bank examiners should be built on the same three core elements SR 11-7 already requires: conceptual soundness, outcomes analysis, and ongoing monitoring. Each section must then be extended with AI-specific evidence, including explainability documentation, training data provenance, drift detection results, and logged monitoring activity, so examiners can assess the model using the same framework applied to traditional statistical models.

    An AI model validation report for bank examiners should be built on the same three core elements SR 11-7 already requires: conceptual soundness, outcomes analysis, and ongoing monitoring. Each section must then be extended with AI-specific evidence, including explainability documentation, training data provenance, drift detection results, and logged monitoring activity, so examiners can assess the model using the same framework applied to traditional statistical models.

    Validation Report Structure Under SR 11-7

    Examiners evaluate AI models through the established SR 11-7 pillars. Structure the report so each pillar includes the AI-specific evidence described below.

    Conceptual Soundness

    Model design, assumptions, and explainability methods with stated limitations.

    Outcomes Analysis

    Backtesting and benchmarking results across relevant data segments.

    Ongoing Monitoring

    Drift detection thresholds, performance tracking, and runtime evidence.

    Governance Accountability

    Independence, board oversight, and documented reporting lines.

    SR 11-7 Still Governs AI Model Validation

    There is no confirmed agency-issued guidance from the Federal Reserve, OCC, or FDIC that replaces SR 11-7 specifically for AI or machine learning models. SR 11-7, issued by the Federal Reserve in April 2011, along with the OCC's companion Bulletin 2011-12, defines a model broadly as any quantitative method, system, or approach that applies statistical, economic, financial, or mathematical techniques to produce quantitative estimates. That definition is broad enough that examiners routinely apply it to AI/ML systems without a separate framework. For validation teams, this means the starting point for an AI model validation report is not a new template but the existing SR 11-7 structure, extended to account for characteristics that traditional deterministic models do not exhibit, such as non-determinism, limited explainability, and dependence on training data composition.

    Mapping the Three Core Validation Elements to AI-Specific Risk

    SR 11-7 defines validation as three core elements: evaluation of conceptual soundness, ongoing monitoring, and outcomes analysis, including benchmarking and backtesting. Each element carries additional documentation requirements when applied to AI/ML models.

    Conceptual soundness review must address how the model was trained, what data shaped its behavior, and what explainability technique, such as feature importance or a surrogate model, was used to assess its logic, along with the known limitations of that technique. Outcomes analysis should include performance metrics broken out across data segments rather than a single aggregate result, since AI models can behave inconsistently across subpopulations in ways linear models typically do not. Ongoing monitoring must account for the fact that model inputs and underlying relationships can shift after deployment, requiring monitoring activity that is more continuous than the periodic review cycles common to traditional statistical models validated at a single point in time.

    Ongoing Monitoring Evidence and Runtime Governance

    SR 11-7 requires ongoing monitoring to confirm that a model continues to perform as intended and that prior validation conclusions remain valid, including outcomes analysis conducted at a regular frequency. For AI/ML models, satisfying this requirement typically means including dated monitoring evidence tied to specific production periods rather than a general description of a monitoring process.

    Drift detection results, performance degradation tracking, and logs showing that runtime policy enforcement remained active over the reporting period all serve as supporting evidence that the model's behavior in production has been observed and controlled, not just validated once at deployment. Where a bank has runtime governance controls in place, such as logging of model access, output review, and enforcement of usage policies, that log data can be referenced in the ongoing monitoring section as evidence of continued oversight, provided it is dated and tied to the specific model and production period under review.

    Practical drafting note

    Prefer dated artifacts (monitoring exports, drift alerts, access logs) over narrative process descriptions. Examiners assess whether conclusions remain valid for the period under review, not whether a process exists in policy alone.

    Practical Guidelines for Drafting the Report

    Keep the report aligned to the SR 11-7 outline examiners already know. Extend each section with AI-specific evidence rather than inventing a parallel template. State limitations of explainability methods explicitly. Pair every identified limitation with compensating controls. Distinguish retraining events from code or configuration changes in version records so validators and examiners can reconstruct what changed and when.

    Documentation Examiners Expect for AI-Specific Risk Areas

    Use the following checklist to confirm the AI-specific materials are present before submission.

    • Training, validation, and test data sources, including any transformation steps that affect model inputs
    • Explainability method used and its documented limitations, rather than an unqualified claim of interpretability
    • Drift detection metrics, defined thresholds, and escalation triggers used in ongoing monitoring
    • Version control records distinguishing model retraining events from code or configuration changes
    • Evidence that independent validators had access to training data and model logic sufficient to assess conceptual soundness
    • Identified model limitations paired with compensating controls, consistent with SR 11-7 documentation expectations

    Extending Validation Reports With Runtime Evidence

    Ongoing monitoring sections of an AI model validation report are only as strong as the operational evidence behind them. Runtime governance logs showing policy enforcement, access control, and monitored model behavior can support the evidence examiners expect for continued model soundness after deployment.

    Explore Runtime Governance