MMLU: Clinical Knowledge

MMLU subject subset: General clinical medicine facts and patient-care knowledge, at a level below the licensing-exam-style Professional Medicine subject.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryknowledge
Subcategoryhealth
Page statusactive
Metricaccuracy
Directionhigher_is_better
Unit%
Dataset size265
Dataset licenceMIT
PublisherUC Berkeley (original); Center for AI Safety (current host)

What it measures

General clinical medicine facts and patient-care knowledge, at a level below the licensing-exam-style Professional Medicine subject. Questions are four-option multiple-choice, drawn from the MMLU test set's "health" subcategory within the benchmark's "other" top-level group, and are graded on the single correct labelled option.

Task format

Four-option multiple-choice questions, graded on the single correct labelled option; commonly evaluated 5-shot, consistent with the rest of MMLU.

Models reporting this benchmark

These figures come from the model cards, which carry one collection date per card and no per-score attribution. They are shown as reported, not as verified evidence.
ModelProviderScoreCard as of
Claude Opus 4Anthropic87.52026-04
Claude Opus 4.6Anthropic87.52026-04
Meta Llama 3 70BMeta86.82024-07
Meta Llama 3 70B InstructMeta86.82026-04
Meta Llama 3 70B InstructNous Research86.82026-04
Gemini 2.5 ProGoogle DeepMind86.52026-04
GPT-4.1OpenAI86.22026-04
GPT-4oOpenAI85.52026-04
GPT-4o (2024-05-13)OpenAI85.52026-04
GPT-4o (2024-08-06)OpenAI85.52026-04
GPT-4o (2024-11-20)OpenAI85.52026-04
GPT-4o miniOpenAI85.52026-04
Claude Sonnet 4Anthropic84.22026-04
Claude Sonnet 4.5Anthropic84.22026-04
Claude Sonnet 4.5 (latest)Anthropic84.22026-04
DeepSeek R1DeepSeek83.52026-04
DeepSeek R1 0528DeepSeek83.52026-04
DeepSeek R1 0528 NVFP4 v2NVIDIA83.52026-04
DeepSeek R1 0528 Qwen3 8BDeepSeek83.52026-04
DeepSeek R1 Distill Llama 70BDeepSeek83.52026-04
DeepSeek R1 Distill Llama 8BDeepSeek83.52026-04
DeepSeek R1 Distill Qwen 1.5BDeepSeek83.52026-04
DeepSeek R1 Distill Qwen 14BDeepSeek83.52026-04
DeepSeek R1 Distill Qwen 32BDeepSeek83.52026-04
DeepSeek R1 Distill Qwen 7BDeepSeek83.52026-04
DeepSeek ReasonerDeepSeek83.52026-04
Yi 1.5 34B01.AI83.02026-04
Yi 1.5 34B 32K01.AI82.62024-07
Mixtral 8x22B Instruct v0.1Mistral AI82.32026-04
Qwen 3 235B InstructCerebras82.22026-04
Qwen3 235B-A22BAlibaba / Qwen Team82.22026-04
Yi 34B 200K01.AI81.12024-07
Mistral Large (latest)Mistral AI80.52026-04
Mistral Large 2.1Mistral AI80.52026-04
Mistral Large 3Mistral AI80.52026-04
Gemma 4 31BGoogle DeepMind80.22026-04
gemma 4 31B itGoogle DeepMind80.22026-04
gemma 4 31B it GGUFUnsloth80.22026-04
Gemma 4 31B IT NVFP4NVIDIA80.22026-04
Nous Hermes 2 Yi 34BNous Research80.02024-07
Gemma 4 26BGoogle DeepMind79.22026-04
Nous Hermes 2 Mixtral 8x7B DPONous Research79.22024-07
Yi 34B Chat01.AI79.22024-07
Llama 3.3 70B Instruct NVFP4NVIDIA78.52026-04
Llama-3.3-70B-InstructMeta78.52026-04
Mixtral 8x7B v0.1Mistral AI78.52026-04
Mixtral 8x7B Instruct v0.1Mistral AI77.72026-04
Yi 1.5 34B Chat01.AI77.42026-04
Yi 1.5 34B Chat 16K01.AI77.42024-07
Llama 3.1 70BMeta77.22026-04
Llama 3.1 70B InstructMeta77.22026-04
phi 4Microsoft76.22026-04
Phi 4 mini instructMicrosoft76.22026-04
Phi 4 multimodal instructMicrosoft76.22026-04
Meta Llama 3 8BMeta75.52024-07
Meta Llama 3 8BNous Research75.52024-07
Meta Llama 3 8B InstructMeta74.72024-07
Meta Llama 3 8B InstructNous Research74.72024-07
Hermes 2 Theta Llama 3 8BNous Research74.32024-07
Phi 3 mini 4K instructMicrosoft74.32024-07
Phi 3 mini 128K instructMicrosoft74.02024-07
Hermes 2 Pro Llama 3 8BNous Research73.62024-07
Yi 9B01.AI73.22024-07
Yi 1.5 9B Chat01.AI72.82024-07
Yi 1.5 9B Chat 16K01.AI72.52024-07
Yi 1.5 9B 32K01.AI71.72024-07
Nous Hermes 2 SOLAR 10.7BNous Research69.12024-07
gemma 7B itGoogle DeepMind68.72024-07
Mistral 7B v0.3Mistral AI68.32024-07
mistral 7B v0.3 bnb 4bitUnsloth68.32024-07
Yi 1.5 9B01.AI68.32024-07
Yi 1.5 6B Chat01.AI67.92024-07
Mistral 7B Instruct v0.2Mistral AI67.22024-07
Yi 6B01.AI66.82024-07
Yi 6B Chat01.AI66.82024-07
Yi 1.5 6B01.AI65.72024-07
falcon 40BTII60.82024-07
phi 2Microsoft60.42024-07
Qwen2 1.5B InstructAlibaba / Qwen Team60.02024-07
deepseek llm 7B baseDeepSeek53.22024-07
deepseek llm 7B chatDeepSeek53.22024-07
Qwen2 0.5B InstructAlibaba / Qwen Team50.22024-07
chatglm2 6BZhipu AI48.72024-07
gemma 2BGoogle DeepMind46.82024-07
gemma 2B itGoogle DeepMind42.32024-07
deepseek coder 6.7B baseDeepSeek41.12024-07
deepseek coder 6.7B instructDeepSeek41.12024-07
deepseek coder 1.3B baseDeepSeek30.22024-07
deepseek coder 1.3B instructDeepSeek30.22024-07
OLMo 1B hfAllen AI19.62024-07

Data

This page as JSON · Edit on GitHub