OpenAI evaluates all frontier models with repeatable evaluations (formerly "scorecards"), including at every 2x increase in effective training compute. Capabilities are classified within tracked categories and results are mapped to capability thresholds. Reaching the "High capability" threshold requires safeguards that sufficiently minimize risk before deployment; the "Critical capability" threshold requires safeguards during development as well. The Safety Advisory Group reviews whether safeguards are sufficient and makes recommendations to leadership, which makes the final decisions. The framework also relies on scalable automated evaluations and allows requirements to be adjusted in response to shifts in the competitive landscape.
Frontier-model development can produce capabilities able to cause severe harm (bioweapons, cyberattacks, autonomous AI self-improvement) before adequate safeguards exist. The framework addresses the lack of a systematic, published process for measuring these capabilities and tying development and deployment decisions to explicit risk thresholds.
Established high-risk capability areas. In the 2025 version: Biological and Chemical, Cybersecurity, and AI Self-improvement. In the Beta (2023): cybersecurity, CBRN, persuasion and model autonomy.
Introduced in 2025: areas that could pose severe risk but do not yet meet the criteria for Tracked Categories โ Long-range Autonomy, Sandbagging, Autonomous Replication and Adaptation, Undermining Safeguards, and Nuclear and Radiological.
Since 2025, two thresholds: "High capability" (could amplify existing pathways to severe harm โ requires safeguards before deployment) and "Critical capability" (could introduce unprecedented new pathways to severe harm โ requires safeguards during development too). The Beta used a four-level scale: low, medium, high, critical.
Technical and operational measures that must "sufficiently minimize" risk before a model reaching a given capability threshold is deployed or developed further.
A cross-functional team of internal safety leaders that reviews whether safeguards sufficiently minimize severe risk and makes recommendations โ from approving deployment to requesting further evaluation or stronger protections โ to OpenAI Leadership.
Repeatable evaluations updated at, among other points, every 2x increase in effective training compute; in 2025 extended with scalable automated evaluations complemented by expert-led "deep dives".
First version with four tracked categories (cybersecurity, CBRN, persuasion, model autonomy) and a four-level risk scale.
Update of 15 April 2025: narrowed tracked categories (Biological and Chemical, Cybersecurity, AI Self-improvement), new Research Categories, two thresholds (High, Critical) and the Safety Advisory Group; persuasion moved out of the framework.