MedQA

Four-option USMLE-style clinical multiple-choice questions, the most widely reported medical exam benchmark for LLMs.

Also known as: MedQA-USMLE, MedQA-USMLE-4-options

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorydomain
Subcategorymedical licensing exam question answering
Page statusactive
Metricaccuracy
Directionhigher_is_better
Unit%
Dataset size12723
Dataset licenceNot stated by the paper; the authors' GitHub repository is MIT-licensed; the accompanying textbook corpus is released under a research-use-only agreement
PublisherMIT Computer Science and Artificial Intelligence Laboratory (CSAIL)

What it measures

MedQA tests whether a model can pick the correct answer to a clinical multiple-choice question written in the style of the United States Medical Licensing Examination. Most questions present a short patient vignette (age, presenting symptoms, exam findings, sometimes lab values) and ask for a diagnosis, a next step in management, or an underlying mechanism, then offer several candidate answers. It is a single-turn, English-language, text-only task built from real practice-exam question banks rather than written for the benchmark, so it leans on applied clinical reasoning more than isolated fact recall.

Task format

Four-option multiple-choice clinical vignette question; the model returns a single letter answer (A-D), usually zero-shot or few-shot.

Models reporting this benchmark

These figures come from the model cards, which carry one collection date per card and no per-score attribution. They are shown as reported, not as verified evidence.
ModelProviderScoreCard as of
o1OpenAI96.52026-04
Gemini 3.1 Pro PreviewGoogle DeepMind96.42026-04
GPT-5.1OpenAI96.42026-04
o3OpenAI96.12026-04
GPT-5.2OpenAI95.82026-04
o4-miniOpenAI95.22026-04
GPT-5OpenAI93.02026-04
GPT-5.4OpenAI93.02026-04
DeepSeek R1DeepSeek92.12026-04
DeepSeek R1 0528DeepSeek92.12026-04
o3-miniOpenAI91.42026-04
GPT-4.1OpenAI89.72026-04
medgemma 27B itGoogle DeepMind88.52026-04
Claude Sonnet 3.7Anthropic87.62026-04
Grok 3xAI86.12026-04
Qwen3 MaxAlibaba / Qwen Team85.52026-04
Qwen3 30B-A3BAlibaba / Qwen Team85.32026-04
Qwen3 235B-A22BAlibaba / Qwen Team84.82026-04
Gemini 2.0 FlashGoogle DeepMind83.22026-04
Llama 3.1 405B InstructMeta82.92026-04
Claude Opus 4Anthropic82.12026-04
Claude Opus 4.6Anthropic82.12026-04
Mistral Large 2.1Mistral AI81.52026-04
DeepSeek V3DeepSeek80.32026-04
Gemini 2.5 ProGoogle DeepMind80.22026-04
Gemini 2.5 Pro Preview 05-06Google DeepMind80.22026-04
Gemini 2.5 Pro Preview 06-05Google DeepMind80.22026-04
Gemini 2.5 Pro Preview TTSGoogle DeepMind80.22026-04
Mistral Medium (latest)Mistral AI79.12026-04
Mistral Medium 3Mistral AI79.12026-04
Mistral Medium 3.1Mistral AI79.12026-04
GPT-4oOpenAI78.52026-04
GPT-4o (2024-05-13)OpenAI78.52026-04
GPT-4o (2024-08-06)OpenAI78.52026-04
GPT-4o (2024-11-20)OpenAI78.52026-04
GPT-4o miniOpenAI78.52026-04
Llama 4 Maverick 17B 128E InstructMeta78.42026-04
Pixtral Large (latest)Mistral AI78.32026-04
Claude Haiku 3.5Anthropic77.82026-04
Claude Haiku 3.5 (latest)Anthropic77.82026-04
Claude Haiku 4.5Anthropic77.82026-04
Claude Haiku 4.5 (latest)Anthropic77.82026-04
phi 4Microsoft77.82026-04
Gemma 3 27BGoogle DeepMind74.92026-04
Command A VisionCohere73.32026-04
medgemma 4B itGoogle DeepMind72.12026-04
Llama 3.1 70BMeta65.22026-04
Llama 3.1 70B InstructMeta65.22026-04
medgemma 1.5 4B itGoogle DeepMind64.42026-04
Llama 3.2 3B InstructMeta52.62026-04
Llama 4 Scout 17B 16E InstructMeta52.02026-04

Data

This page as JSON · Edit on GitHub