HealthBench
HealthBench measures the quality and safety of open-ended clinical conversations rather than performance on medical multiple-choice exams. Responses are graded against rubrics written by practising physicians. It became the 2026 reference for medical AI because it tests the thing that actually matters clinically: what the model says to a person, not whether it can pass a test.
| What it measures | Open-ended clinical conversation quality and safety, graded against rubrics written by practising physicians. |
|---|---|
| Built by | OpenAI (Arora, Wei, Soskin Hicks et al.), 2025 |
| Format | 5,000 multi-turn conversations with users and health professionals, scored against 48,562 rubric criteria written by 262 physicians |
| Scoring metric | Rubric score, graded by a model grader against physician-written criteria |
| Status | Active |
| Representative top score | Not independently confirmed; see the leaderboard |
How HealthBench works
The benchmark contains 5,000 multi-turn conversations between a model and either an individual user or a healthcare professional. Each conversation has its own rubric, and across the set there are 48,562 unique criteria written by 262 physicians, spanning contexts such as emergencies, transforming clinical data and global health, and behavioural dimensions such as accuracy, instruction following and communication. A model grader scores responses against those criteria.
History and current status
OpenAI released HealthBench in 2025. It arrived against a background of medical benchmarks that had gone stale: MedQA, built from board-exam questions, was saturated and heavily contaminated, and "AI passes the medical licensing exam" headlines had stopped being informative. HealthBench replaced the exam format with rubric-graded conversation, and Stanford MedHELM provides the independent academic counterpart.
What the score does not tell you
Two structural caveats. It was built and is run by a model vendor, which is the same independence problem that dogs vendor-run benchmarks generally. And it is graded by a model against physician-written criteria, so the grader’s own reliability is part of the measurement. The physician-written rubrics are a genuine strength; the vendor-built, model-graded pipeline around them is where to apply scepticism.
This is the general case, not a quirk of one benchmark. For the four failure modes that decide whether any benchmark number means anything, read are AI benchmarks reliable, and for this row in particular, how benchmark contamination inflates scores.
Who reports HealthBench, and how to read it
OpenAI reports it for its own models, and it now appears widely in medical AI coverage. For a question about fitness for a specific clinical workflow rather than general conversation quality, Stanford MedHELM is the better citation because it is independent and organised around a clinician-validated task taxonomy.
Benchmarks to read alongside this one
MedHELM
Clinical ability across the breadth of real medical work, on a clinician-validated taxonomy rather than exam questions.
MedQA
Medical knowledge, using real questions from professional medical board examinations.
LegalBench
Legal reasoning across the specific skills lawyers actually use, as defined by legal professionals.
HealthBench: frequently asked questions
- What is HealthBench?
- HealthBench is a 2025 OpenAI benchmark of 5,000 multi-turn health conversations, scored against 48,562 rubric criteria written by 262 physicians. Unlike multiple-choice medical benchmarks it evaluates open-ended clinical conversation quality and safety.
- How is HealthBench different from MedQA?
- MedQA is multiple-choice board-exam questions, now saturated and contaminated. HealthBench is open-ended conversation graded against physician-written rubrics. Passing an exam and safely handling a patient conversation are different skills, and HealthBench measures the second.
- Is HealthBench independent?
- No. It was built and is run by OpenAI, and responses are graded by a model rather than by the physicians who wrote the rubrics. Stanford MedHELM is the independent academic alternative, organised around a clinician-validated taxonomy of real clinical tasks.
Get each breakdown before it makes the rounds
You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.
No spam. Unsubscribe anytime.