The starting point is real clinician conversations, from which physicians selected 525 tasks. Each task gets a rubric: a set of criteria describing what the answer should contain and what it must not. The rubrics are not one person's work — at least three physicians write and adjudicate them across three phases, with difficulty explicitly marked on a 1-to-7 Likert scale. The model's answer is scored against the rubric on a continuous 0-to-1 scale, so partial correctness is visible. Tasks are spread across three clinical use cases — care consultation, writing and documentation, and medical research — and roughly a third come from red teaming, meaning they were constructed to make the model fail.
Models evaluated on medical exam questions perform excellently, which says nothing about how they will fare in a conversation with a patient or when producing documentation. HealthBench Professional moves evaluation onto material drawn from real practice and onto criteria physicians deemed material.
Tasks selected by physicians, spread across care consultation, writing and documentation, and medical research.
Written and adjudicated by at least three physicians across three phases, with difficulty on a 1–7 Likert scale.
About one third of the examples were constructed to induce model error.
26 specialties, 52 languages used professionally; every example went through a three-phase review.
A high benchmark score does not mean the model is fit for autonomous clinical use or that it meets medical device regulatory requirements.
One third of the set comes from red teaming, so absolute scores are lower than on benchmarks built from typical material.
In April 2026 OpenAI released a set of 525 clinical tasks with physician rubrics, built by 190 physicians from 50 countries.