MedXpertQA MM

The multimodal split of MedXpertQA, pairing expert-level clinical questions with medical images and up to ten answer options.

Also known as: MedXpertQA Multimodal, MedXpertQA-MM

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorydomain
Subcategoryexpert-level multimodal clinical reasoning
Page statusactive
Metricaccuracy
Directionhigher_is_better
Unit%
Dataset size2005
Dataset licenceMIT
PublisherTsinghua University

What it measures

This page covers MedXpertQA MM, the multimodal split of the MedXpertQA benchmark. Each item pairs an expert-level clinical question, drawn from and modelled on specialty board exam material across 17 medical specialties and 11 body systems, with one or more medical images (for example a radiograph, histopathology slide, or ECG trace) and supporting clinical documentation such as history and exam findings. The model must integrate the image and the text to answer. Example questions on the project's own site show as many as ten lettered answer options, noticeably more than the four or five typical of MedQA or MedMCQA. Questions are also labelled by type, "Understanding" (applying known medical knowledge) or "Reasoning" (multi-step clinical inference). MedXpertQA's other half, a text-only split called MedXpertQA Text, does not have a separate page in this repository; be careful to check which split a reported "MedXpertQA" score refers to.

Task format

Multiple-choice clinical question paired with one or more medical images and supporting exam findings; the model returns a single letter from as many as ten options.

Models reporting this benchmark

These figures come from the model cards, which carry one collection date per card and no per-score attribution. They are shown as reported, not as verified evidence.
ModelProviderScoreCard as of
Muse SparkMeta78.42026-04

Data

This page as JSON · Edit on GitHub