Model Cards and Datasheets: What Regulators Now Expect to See
Regulators now expect model cards and datasheets to go beyond static, single-release disclosures. Under the EU AI Act's Article 11 and Annex IV, and NIST's AI Risk Management Framework and Generative AI Profile (AI 600-1), documentation must cover intended purpose, design specifications, data provenance, and validation testing. It must also remain traceable to the specific model or agent version deployed in production, not just its original release.
From Academic Format to Regulatory Requirement
Model cards and datasheets originated as academic proposals. Model Cards for Model Reporting (Mitchell et al., 2019) described a structured format covering intended use, training and evaluation data, and performance across demographic groups. Datasheets for Datasets (Gebru et al., 2018 and 2021) proposed similar structure for dataset documentation, covering motivation, composition, and collection process.
These formats were designed for static, single-release disclosure, typically published alongside a model or dataset at the time it was made public. That design assumption is now the primary point of friction as regulators formalize documentation requirements for enterprise-deployed systems, which are updated, fine-tuned, and retrained continuously in production rather than released once.
The Two Anchor Frameworks Enterprises Are Measured Against
EU AI Act (Regulation (EU) 2024/1689)
The EU AI Act requires providers of high-risk AI systems to maintain technical documentation under Article 11, with required content specified in Annex IV. That content includes a description of the system's intended purpose, its design specifications, and information on training, validation, and testing datasets, including their provenance. Documentation obligations are tied to a system's risk classification under Annex III, meaning governance teams must first establish and record the risk tier before determining documentation scope.
NIST AI Risk Management Framework and AI 600-1
NIST's AI Risk Management Framework (AI 100-1, January 2023) organizes governance activity into four functions: Govern, Map, Measure, and Manage. It treats documentation as a mechanism for risk traceability across the AI lifecycle rather than a standalone deliverable. NIST's Generative AI Profile (AI 600-1, July 2024) adds a documentation category not present in earlier academic formats: tracking content provenance for generative outputs.
NIST guidance is voluntary in the United States context. EU AI Act obligations are legally mandated for in-scope systems. This distinction matters for organizations operating across both jurisdictions.
| Framework | Key Documentation Scope | Legal Status |
|---|---|---|
| EU AI Act Annex IV | Intended purpose, design specifications, data provenance, validation testing for high-risk systems before market placement | Legally mandated (EU) |
| NIST AI RMF (AI 100-1) | Lifecycle risk traceability across Govern, Map, Measure, and Manage functions | Voluntary (US) |
| NIST AI 600-1 (2024) | Extends AI RMF to generative AI; adds content and data provenance tracking as a documentation category | Voluntary (US) |
| Academic origins | Model Cards (2019) and Datasheets for Datasets (2018/2021) established the structural basis regulators now build on | Non-binding reference |
Where Traditional Model Cards Fall Short
Academic-style model cards and datasheets remain structurally useful, but three gaps consistently appear when they are mapped against current regulatory content categories.
- Data lineage: Annex IV and Datasheets for Datasets both call for traceability back to original data sources and collection processes, which many existing templates document only at a summary level.
- Validation testing detail: Annex IV expects documentation of testing evidence, not just a description of intended use.
- Version specificity: A model card written for a base model does not necessarily describe the behavior of a fine-tuned or policy-adjusted version running in production.
Multi-agent or chained-model deployments compound this gap. A single model card format was not built to capture composite system behavior across multiple models or agents interacting in a pipeline, which is increasingly how enterprise AI systems are architected.
Static Documentation Versus Ongoing Operational Evidence
Both frameworks treat documentation as part of an ongoing risk management process rather than a one-time compliance artifact. This is a meaningful shift from the original academic intent. A static model card produced at release describes the system as it existed at one point in time. Regulators reviewing high-risk systems, and NIST's lifecycle-oriented guidance, expect evidence that documentation reflects the system's current, deployed state.
This creates a practical requirement: documentation update processes should be triggered by model retraining, fine-tuning, or prompt and policy changes, not only at initial release. Runtime logging of policy enforcement and model behavior can serve as supplementary evidence alongside static documentation, helping demonstrate how a system actually behaved in production over time.
Operational runtime evidence is distinct from the documentation artifact itself and does not replace it. It addresses the currency gap that static formats cannot close on their own.
Building an Audit-Ready Documentation Process
Moving from a static model card practice to one that satisfies regulatory scrutiny requires connecting documentation workflows to the AI system lifecycle. Key process considerations include:
- Mapping each deployed system to the appropriate EU AI Act Annex III risk tier before scoping documentation requirements.
- Maintaining version-specific records for each fine-tuned or policy-updated model variant, not only the base model release.
- Defining explicit triggers that initiate a documentation update, such as retraining events, serving policy changes, or dataset updates.
- Capturing runtime logs and policy enforcement records that can supplement static documentation during a regulatory audit.
- Assigning clear documentation ownership across data science, legal, and governance functions, with accountability enforced operationally rather than assumed.
Questions Governance Leaders Should Be Able to Answer
- Does our model card or datasheet template cover every content category specified in EU AI Act Annex IV, including data provenance and validation testing?
- Can we produce documentation reflecting the specific deployed version of a model or agent at a given point in time, not just its base version?
- What process triggers a documentation update when a model is fine-tuned, retrained, or its serving policy changes?
- Do we have operational evidence, such as logs, that can supplement static documentation during an audit?
- Who owns documentation accuracy across data science, legal, and governance functions, and how is that ownership enforced day to day?
Frequently Asked Questions
Are model cards legally required under the EU AI Act?
The EU AI Act does not mandate the use of the specific "model card" format as originally proposed academically. What Article 11 and Annex IV require is that providers of high-risk AI systems maintain technical documentation covering specific content categories, including intended purpose, design specifications, and data provenance. Organizations may use any documentation format that satisfies those content requirements.
Does NIST AI 600-1 apply to all generative AI systems?
NIST AI 600-1 is guidance, not a legal mandate in the United States. It extends the AI RMF to address risks specific to generative AI, including content provenance. Organizations that have adopted the AI RMF or are subject to sector-specific regulations that reference NIST standards should evaluate how AI 600-1 applies to their generative AI deployments.
How often should model cards be updated?
Regulatory expectations tied to lifecycle risk management imply that documentation should reflect the system as currently deployed, not as originally released. In practice, documentation update triggers should include model retraining, fine-tuning, significant prompt or policy changes, and dataset updates. The specific cadence will depend on the system's risk tier and rate of change.
How do multi-agent systems affect documentation requirements?
Multi-agent or chained-model deployments present a documentation challenge that single model card formats were not designed to address. Documentation for composite systems should capture the behavior of the overall pipeline, not only individual components. This may require supplementary architecture documentation alongside per-component model cards.
What is the difference between a model card and a datasheet for datasets?
Model cards focus on a trained model: its intended use, performance characteristics across groups, and evaluation methodology. Datasheets for Datasets focus on the dataset used: its motivation, composition, collection process, and preprocessing steps. Regulatory frameworks such as EU AI Act Annex IV draw on both formats, requiring documentation of both the model and the data used to train and evaluate it.
Keep Model Documentation Aligned With What Is Actually Deployed
Trussed AI provides runtime governance and audit logging that help enterprise teams maintain evidence of how AI models and agents actually behave in production, supporting documentation that stays current as systems change.
Explore Runtime Governance