HealthBench Hard

The 1,000 hardest conversations from OpenAI's HealthBench, graded against physician-written rubrics of what a good health-related response should do.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorydomain
Subcategoryclinical and health conversations, hardest subset
Page statusactive
Metricrubric criteria met (mean score, bootstrap-aggregated)
Directionhigher_is_better
Unit%
Dataset size1000
Dataset licenceMIT
PublisherOpenAI

What it measures

HealthBench Hard tests how a model handles realistic, multi-turn healthcare conversations, standing in for a patient, caregiver or clinician. Each conversation was written or adapted with input from physicians to probe a specific behaviour: giving correct information, communicating clearly, being appropriately cautious, or handling an emergency or context correctly. HealthBench Hard is a fixed 1,000-conversation slice of the full 5,000-conversation HealthBench set, chosen as the conversations that scored lowest, on average, across a panel of model providers when the benchmark was built -- the cases current frontier models handled worst.

Task format

Multi-turn healthcare conversation; the model's final response is graded by an LLM judge against a per-conversation, physician-written rubric of scoring criteria.

Models reporting this benchmark

These figures come from the model cards, which carry one collection date per card and no per-score attribution. They are shown as reported, not as verified evidence.
ModelProviderScoreCard as of
Muse SparkMeta42.82026-04

Data

This page as JSON · Edit on GitHub