The 1,000 hardest conversations from OpenAI's HealthBench, graded against physician-written rubrics of what a good health-related response should do.
unassessed
| Category | domain |
|---|---|
| Subcategory | clinical and health conversations, hardest subset |
| Page status | active |
| Metric | rubric criteria met (mean score, bootstrap-aggregated) |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 1000 |
| Dataset licence | MIT |
| Publisher | OpenAI |
HealthBench Hard tests how a model handles realistic, multi-turn healthcare conversations, standing in for a patient, caregiver or clinician. Each conversation was written or adapted with input from physicians to probe a specific behaviour: giving correct information, communicating clearly, being appropriately cautious, or handling an emergency or context correctly. HealthBench Hard is a fixed 1,000-conversation slice of the full 5,000-conversation HealthBench set, chosen as the conversations that scored lowest, on average, across a panel of model providers when the benchmark was built -- the cases current frontier models handled worst.
Multi-turn healthcare conversation; the model's final response is graded by an LLM judge against a per-conversation, physician-written rubric of scoring criteria.
| Model | Provider | Score | Card as of |
|---|---|---|---|
| Muse Spark | Meta | 42.8 | 2026-04 |