What Frontier AI Safety Frameworks Actually Define
A frontier AI safety framework is a published document in which an AI lab defines capability thresholds it monitors in its own models, the evaluations used to detect approach to those thresholds, and the mitigations it commits to apply before or at the point a threshold is reached.
Anthropic's Responsible Scaling Policy (RSP) defines AI Safety Levels from ASL-1 through ASL-4, with ASL-5 held as a conceptual future tier. These levels are tied to catastrophic-risk domains such as CBRN weapons assistance and autonomous AI research and development capability. Google DeepMind's Frontier Safety Framework, introduced in May 2024, uses a structurally similar concept called Critical Capability Levels (CCLs) across CBRN, cyber-offense, ML R&D, and deceptive alignment, paired with early-warning evaluations conducted before a level is reached. OpenAI's Preparedness Framework tracks the same general risk categories, including cybersecurity, CBRN, persuasion, and model autonomy, using low, medium, high, and critical designations that gate further deployment or scaling.
All three rely on internal capability evaluations run during training and deployment rather than solely on post-deployment monitoring. Critically, none currently offers an enterprise-facing certification or third-party attestation mechanism. For governance leaders, this distinction matters: these are vendor risk-management commitments, not audited compliance artifacts.
How the Major Frameworks Compare
The three leading frameworks share a common structure, capability thresholds tied to risk domains, evaluation requirements, and mitigation commitments, but they differ in terminology, scope, and specificity.
| Framework | Published by | Tier structure | Risk domains covered | Evaluation timing |
|---|---|---|---|---|
| Responsible Scaling Policy (RSP) | Anthropic | ASL-1 to ASL-4 (ASL-5 conceptual) | CBRN, autonomous AI R&D capability | During training and deployment; pre-threshold |
| Frontier Safety Framework (FSF) | Google DeepMind | Critical Capability Levels (CCLs) | CBRN, cyber-offense, ML R&D, deceptive alignment | Early-warning evaluations before a level is reached |
| Preparedness Framework | OpenAI | Low / medium / high / critical | Cybersecurity, CBRN, persuasion, model autonomy | During training; gates further deployment or scaling |
| Agentic AI Standards | Emerging, no single body | No unified tier structure | Tool-call permissions, agent identity, scope | Not defined at cross-industry level |
The structural similarities make cross-vendor comparison possible at a high level, but the absence of shared evaluation methodologies or common thresholds means that an ASL-3 designation from Anthropic and a high CCL designation from Google DeepMind cannot be treated as equivalent in a procurement or audit context.
Where Agentic AI Standards Fit, and Where They Do Not
None of the reviewed frontier frameworks were built primarily for autonomous, tool-using agents. They were designed around model capability thresholds assessed at training time. Agentic AI standards, covering how an agent is authenticated, what tools it can call, and under what permission scope, remain fragmented and without a single ratified cross-industry specification.
Anthropic's Model Context Protocol (MCP), introduced in November 2024, standardizes how AI assistants connect to external tools and data sources, which is directly relevant to enterprise runtime enforcement. However, MCP addresses connection standardization rather than authorization or audit at a governance-policy level.
OpenAI's Preparedness Framework references a "model autonomy" risk category, which is a useful signal that self-directed task execution is a recognized risk class across labs, but it does not translate into a specific technical control an enterprise can deploy.
Agent authorization, scoping, and audit logging remain the enterprise's responsibility. These controls must be layered on top of whatever protocol-level standardization a vendor provides, because no published frontier safety framework supplies them out of the box.
The Governance Gap: Self-Disclosure, Not Certification
A recurring limitation across RSP, Frontier Safety Framework, and Preparedness Framework documents is that none currently provides a formal, enterprise-facing certification or compliance-attestation mechanism. Commitments are self-reported by the publishing lab, based on internal capability evaluations that are not independently audited or standardized across labs.
At the May 2024 Seoul AI Safety Summit, sixteen AI companies signed the Frontier AI Safety Commitments, agreeing to publish safety frameworks defining unacceptable risk thresholds, describe mitigations, and refrain from deployment if thresholds cannot be adequately mitigated. This is a useful directional signal for procurement conversations, but it is a voluntary multilateral commitment, not a binding regulation or an enforceable compliance standard.
The Frontier Model Forum, founded in 2023 by Anthropic, Google, Microsoft, and OpenAI, has published technical work on threshold-setting referenced by member labs, which gives enterprises another reference point for assessing vendor safety maturity. Governance leaders should treat all of this as vendor risk disclosure to be mapped into internal control requirements, not as evidence that an outside party has verified the vendor's practices.
Operational Steps to Align Internal Governance with Published Frameworks
Because frontier safety frameworks are vendor disclosures rather than compliance certifications, enterprises must perform their own translation work. The following steps provide a starting structure.
- Inventory AI systems by vendor framework. Map each deployed or evaluated model to the applicable framework (RSP, FSF, Preparedness Framework, or none) and record the current tier or risk level designation disclosed by the vendor.
- Identify relevant risk domains. Determine which risk categories in each framework (CBRN, cyber-offense, autonomy, persuasion) are material to your use cases. Most enterprise deployments operate well below catastrophic-risk thresholds, but documenting this assessment creates an auditable record.
- Map vendor commitments to internal controls. For each vendor commitment, such as "we will not deploy a model above ASL-3 without additional safeguards," identify whether your internal controls (access management, output review, rate limiting) provide an independent check or rely entirely on the vendor.
- Address the agentic gap separately. For any agentic deployment, define agent identity, tool-call permission scopes, approval workflows for multi-step actions, and audit log requirements independently. Do not assume a vendor's published framework covers these controls.
- Establish a monitoring cadence. Frontier safety frameworks are living documents; labs update tier definitions and evaluation criteria as model capabilities advance. Assign ownership for tracking framework revisions and assessing whether updates require changes to internal controls.
- Include framework alignment in procurement reviews. Add questions about framework tier, recent evaluation results, and third-party audit availability to vendor security questionnaires. Note gaps where vendors cannot supply attestation artifacts.