MMLU: Professional Medicine

MMLU subject subset: USMLE-style clinical vignettes on diagnosis, mechanism and management across medical specialties.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryknowledge
Subcategoryhealth
Page statusactive
Metricaccuracy
Directionhigher_is_better
Unit%
Dataset size272
Dataset licenceMIT
PublisherUC Berkeley (original); Center for AI Safety (current host)

What it measures

USMLE-style clinical vignettes: a patient history and findings followed by a question on diagnosis, mechanism or next step in management. Questions are four-option multiple-choice, drawn from the MMLU test set's "health" subcategory within the benchmark's "other" top-level group, and are graded on the single correct labelled option.

Task format

Four-option multiple-choice questions, graded on the single correct labelled option; commonly evaluated 5-shot, consistent with the rest of MMLU.

Models reporting this benchmark

These figures come from the model cards, which carry one collection date per card and no per-score attribution. They are shown as reported, not as verified evidence.
ModelProviderScoreCard as of
Meta Llama 3 70BMeta89.02024-07
Meta Llama 3 70B InstructMeta89.02026-04
Meta Llama 3 70B InstructNous Research89.02026-04
Mixtral 8x22B Instruct v0.1Mistral AI88.62026-04
Yi 1.5 34B01.AI84.92026-04
Nous Hermes 2 Yi 34BNous Research83.12024-07
Yi 1.5 34B 32K01.AI82.72024-07
Mixtral 8x7B v0.1Mistral AI81.22026-04
Yi 1.5 34B Chat01.AI81.22026-04
Yi 1.5 34B Chat 16K01.AI81.22024-07
Yi 34B 200K01.AI80.92024-07
Mixtral 8x7B Instruct v0.1Mistral AI79.42026-04
Nous Hermes 2 Mixtral 8x7B DPONous Research78.72024-07
Yi 34B Chat01.AI78.32024-07
Phi 3 mini 4K instructMicrosoft76.12024-07
Nous Hermes 2 SOLAR 10.7BNous Research75.72024-07
Phi 3 mini 128K instructMicrosoft72.42024-07
Meta Llama 3 8BMeta72.12024-07
Meta Llama 3 8BNous Research72.12024-07
Meta Llama 3 8B InstructMeta71.72024-07
Meta Llama 3 8B InstructNous Research71.72024-07
Yi 9B01.AI71.02024-07
Yi 1.5 9B01.AI70.22024-07
Hermes 2 Pro Llama 3 8BNous Research69.52024-07
Hermes 2 Theta Llama 3 8BNous Research69.52024-07
Yi 1.5 9B 32K01.AI69.52024-07
Mistral 7B v0.3Mistral AI68.82024-07
mistral 7B v0.3 bnb 4bitUnsloth68.82024-07
Yi 1.5 9B Chat 16K01.AI68.02024-07
Yi 1.5 9B Chat01.AI67.62024-07
Yi 6B01.AI67.32024-07
Yi 6B Chat01.AI67.32024-07
Yi 1.5 6B01.AI65.82024-07
gemma 7B itGoogle DeepMind63.22024-07
Mistral 7B Instruct v0.2Mistral AI61.82024-07
falcon 40BTII61.02024-07
Yi 1.5 6B Chat01.AI53.72024-07
Qwen2 1.5B InstructAlibaba / Qwen Team50.42024-07
phi 2Microsoft48.22024-07
deepseek llm 7B baseDeepSeek46.72024-07
deepseek llm 7B chatDeepSeek46.72024-07
deepseek coder 6.7B baseDeepSeek44.92024-07
Qwen2 0.5B InstructAlibaba / Qwen Team44.92024-07
deepseek coder 1.3B baseDeepSeek41.52024-07
deepseek coder 1.3B instructDeepSeek41.52024-07
OLMo 1B hfAllen AI40.42024-07
chatglm2 6BZhipu AI35.32024-07
gemma 2BGoogle DeepMind35.32024-07
deepseek coder 6.7B instructDeepSeek34.92024-07
gemma 2B itGoogle DeepMind20.22024-07

Data

This page as JSON · Edit on GitHub