Four-option USMLE-style clinical multiple-choice questions, the most widely reported medical exam benchmark for LLMs.
unassessed
| Category | domain |
|---|---|
| Subcategory | medical licensing exam question answering |
| Page status | active |
| Metric | accuracy |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 12723 |
| Dataset licence | Not stated by the paper; the authors' GitHub repository is MIT-licensed; the accompanying textbook corpus is released under a research-use-only agreement |
| Publisher | MIT Computer Science and Artificial Intelligence Laboratory (CSAIL) |
MedQA tests whether a model can pick the correct answer to a clinical multiple-choice question written in the style of the United States Medical Licensing Examination. Most questions present a short patient vignette (age, presenting symptoms, exam findings, sometimes lab values) and ask for a diagnosis, a next step in management, or an underlying mechanism, then offer several candidate answers. It is a single-turn, English-language, text-only task built from real practice-exam question banks rather than written for the benchmark, so it leans on applied clinical reasoning more than isolated fact recall.
Four-option multiple-choice clinical vignette question; the model returns a single letter answer (A-D), usually zero-shot or few-shot.
| Model | Provider | Score | Card as of |
|---|---|---|---|
| o1 | OpenAI | 96.5 | 2026-04 |
| Gemini 3.1 Pro Preview | Google DeepMind | 96.4 | 2026-04 |
| GPT-5.1 | OpenAI | 96.4 | 2026-04 |
| o3 | OpenAI | 96.1 | 2026-04 |
| GPT-5.2 | OpenAI | 95.8 | 2026-04 |
| o4-mini | OpenAI | 95.2 | 2026-04 |
| GPT-5 | OpenAI | 93.0 | 2026-04 |
| GPT-5.4 | OpenAI | 93.0 | 2026-04 |
| DeepSeek R1 | DeepSeek | 92.1 | 2026-04 |
| DeepSeek R1 0528 | DeepSeek | 92.1 | 2026-04 |
| o3-mini | OpenAI | 91.4 | 2026-04 |
| GPT-4.1 | OpenAI | 89.7 | 2026-04 |
| medgemma 27B it | Google DeepMind | 88.5 | 2026-04 |
| Claude Sonnet 3.7 | Anthropic | 87.6 | 2026-04 |
| Grok 3 | xAI | 86.1 | 2026-04 |
| Qwen3 Max | Alibaba / Qwen Team | 85.5 | 2026-04 |
| Qwen3 30B-A3B | Alibaba / Qwen Team | 85.3 | 2026-04 |
| Qwen3 235B-A22B | Alibaba / Qwen Team | 84.8 | 2026-04 |
| Gemini 2.0 Flash | Google DeepMind | 83.2 | 2026-04 |
| Llama 3.1 405B Instruct | Meta | 82.9 | 2026-04 |
| Claude Opus 4 | Anthropic | 82.1 | 2026-04 |
| Claude Opus 4.6 | Anthropic | 82.1 | 2026-04 |
| Mistral Large 2.1 | Mistral AI | 81.5 | 2026-04 |
| DeepSeek V3 | DeepSeek | 80.3 | 2026-04 |
| Gemini 2.5 Pro | Google DeepMind | 80.2 | 2026-04 |
| Gemini 2.5 Pro Preview 05-06 | Google DeepMind | 80.2 | 2026-04 |
| Gemini 2.5 Pro Preview 06-05 | Google DeepMind | 80.2 | 2026-04 |
| Gemini 2.5 Pro Preview TTS | Google DeepMind | 80.2 | 2026-04 |
| Mistral Medium (latest) | Mistral AI | 79.1 | 2026-04 |
| Mistral Medium 3 | Mistral AI | 79.1 | 2026-04 |
| Mistral Medium 3.1 | Mistral AI | 79.1 | 2026-04 |
| GPT-4o | OpenAI | 78.5 | 2026-04 |
| GPT-4o (2024-05-13) | OpenAI | 78.5 | 2026-04 |
| GPT-4o (2024-08-06) | OpenAI | 78.5 | 2026-04 |
| GPT-4o (2024-11-20) | OpenAI | 78.5 | 2026-04 |
| GPT-4o mini | OpenAI | 78.5 | 2026-04 |
| Llama 4 Maverick 17B 128E Instruct | Meta | 78.4 | 2026-04 |
| Pixtral Large (latest) | Mistral AI | 78.3 | 2026-04 |
| Claude Haiku 3.5 | Anthropic | 77.8 | 2026-04 |
| Claude Haiku 3.5 (latest) | Anthropic | 77.8 | 2026-04 |
| Claude Haiku 4.5 | Anthropic | 77.8 | 2026-04 |
| Claude Haiku 4.5 (latest) | Anthropic | 77.8 | 2026-04 |
| phi 4 | Microsoft | 77.8 | 2026-04 |
| Gemma 3 27B | Google DeepMind | 74.9 | 2026-04 |
| Command A Vision | Cohere | 73.3 | 2026-04 |
| medgemma 4B it | Google DeepMind | 72.1 | 2026-04 |
| Llama 3.1 70B | Meta | 65.2 | 2026-04 |
| Llama 3.1 70B Instruct | Meta | 65.2 | 2026-04 |
| medgemma 1.5 4B it | Google DeepMind | 64.4 | 2026-04 |
| Llama 3.2 3B Instruct | Meta | 52.6 | 2026-04 |
| Llama 4 Scout 17B 16E Instruct | Meta | 52.0 | 2026-04 |