An evaluation defines threats and risk thresholds, then applies tests: automated benchmarks, structured red-teaming, capability assessments under controlled conditions and behavioral experiments. Results are compared against safety thresholds (e.g. Responsible Scaling Policies) that determine mitigations or a deployment hold.
Increasingly capable models can pose serious risks; without rigorous, standardized evaluation it is impossible to reliably determine whether a model is safe to deploy.