The model answers closed-ended biology questions, designed so that some concern routine research work and some are potentially dangerous queries. The closed format allows fully automatic grading with no model-based judge. Results are aggregated as endpoint pass rates โ mean pass rates computed at the level of a single evaluation and averaged across multiple run configurations, which limits the influence of random variation. The competence dimension and the refusal dimension are reported separately, so it is visible whether a low risk-category score comes from contextual understanding or from the model's general ignorance.
Assessing models for biological risk is often reduced to the refusal rate alone, which rewards incompetence โ a model that does not understand biology will not help with misuse either, but is useless for research. The benchmark separates competence from safety and lets both be seen at once.
Tests performance on routine research tasks, including detecting and fixing errors in molecular-biology protocols.
Measures whether the model refuses dangerous dual-use queries.
The closed format avoids a model-based judge; results are reported as endpoint pass rates.
Closed questions give objective grading but do not test how the model behaves in a long, multi-step conversation about a protocol.
The 62.4% figure is attributed in sources sometimes to Grok 4.6 and sometimes to Grok 4.7 โ any citation must name the specific source.
The benchmark covers models from xAI, Anthropic, OpenAI and Google; reported endpoint pass rates span 14% to 50%, with a leading figure of 62.4%.