A frontier model is not an architecture but a classification of models based on their position at the capability frontier. In practice frontier status is determined by: (1) training scale โ cumulative compute in FLOP (e.g. the 10^25 FLOP threshold in the EU AI Act) and parameter count; (2) dangerous-capability evaluations and red-teaming before deployment; (3) governance mechanisms โ registration, reporting, post-deployment monitoring, and safety standards coordinated by bodies such as the Frontier Model Forum and AI safety institutes. Because dangerous capabilities can be emergent, assessment combines quantitative measures (compute, parameters, benchmarks) with qualitative risk analysis.
Earlier notions (foundation model, LLM) did not distinguish models that, by virtue of scale and capability, may pose public-safety risks from everyday-use models. The term provides an operational category โ often grounded in compute thresholds โ that lets regulation, risk assessment and oversight focus on the highest-risk models.
Anderljung et al. define frontier models as highly capable foundation models with potentially dangerous capabilities.
On 26 July 2023 Anthropic, Google, Microsoft and OpenAI found an industry body for frontier model safety.
First international summit on frontier model risks; the Bletchley Declaration.
The EU AI Act ties systemic risk of general-purpose models to a 10^25 FLOP training-compute threshold.
Cumulative compute used for training; the key regulatory threshold (EU AI Act: 10^25 FLOP).
Model size; correlated with capability, though not the sole determinant of frontier status.
Coverage of text, image, audio and video; frontier models are increasingly multimodal.
Training frontier models requires large-scale GPU clusters with tensor cores (tens of thousands of accelerators).
TPUs (e.g. Google) are used to train and serve frontier models such as Gemini.