The multimodal split of MedXpertQA, pairing expert-level clinical questions with medical images and up to ten answer options.
unassessed
| Category | domain |
|---|---|
| Subcategory | expert-level multimodal clinical reasoning |
| Page status | active |
| Metric | accuracy |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 2005 |
| Dataset licence | MIT |
| Publisher | Tsinghua University |
This page covers MedXpertQA MM, the multimodal split of the MedXpertQA benchmark. Each item pairs an expert-level clinical question, drawn from and modelled on specialty board exam material across 17 medical specialties and 11 body systems, with one or more medical images (for example a radiograph, histopathology slide, or ECG trace) and supporting clinical documentation such as history and exam findings. The model must integrate the image and the text to answer. Example questions on the project's own site show as many as ten lettered answer options, noticeably more than the four or five typical of MedQA or MedMCQA. Questions are also labelled by type, "Understanding" (applying known medical knowledge) or "Reasoning" (multi-step clinical inference). MedXpertQA's other half, a text-only split called MedXpertQA Text, does not have a separate page in this repository; be careful to check which split a reported "MedXpertQA" score refers to.
Multiple-choice clinical question paired with one or more medical images and supporting exam findings; the model returns a single letter from as many as ten options.
| Model | Provider | Score | Card as of |
|---|---|---|---|
| Muse Spark | Meta | 78.4 | 2026-04 |