MedMCQA

Over 194,000 multiple-choice questions from India's AIIMS and NEET PG medical entrance exams, across 21 subjects.

Also known as: Med-MCQA

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorydomain
Subcategorymedical entrance exam question answering
Page statusactive
Metricaccuracy
Directionhigher_is_better
Unit%
Dataset size193155
Dataset licenceMIT
PublisherSaama AI Research

What it measures

MedMCQA tests whether a model can answer multiple-choice medical questions drawn from two of India's largest postgraduate medical entrance examinations, AIIMS and NEET PG. Each question covers one of 21 medical subjects (anatomy, pharmacology, surgery, obstetrics, and so on) and asks for the single best answer among several options, mirroring the format used to screen doctors applying for postgraduate specialty training in India. It is a single-turn, English-language, text-only task; the authors report it requires more than ten distinct types of reasoning across the question set, from single-fact recall to multi-hop clinical reasoning. Because the exams it draws from are specific to the Indian medical curriculum, its subject mix and phrasing differ somewhat from the US-focused MedQA, even though both are "medical multiple-choice" benchmarks.

Task format

Four-option multiple-choice medical exam question; the model returns a single letter answer (A-D).

Models reporting this benchmark

These figures come from the model cards, which carry one collection date per card and no per-score attribution. They are shown as reported, not as verified evidence.
ModelProviderScoreCard as of
medgemma 27B itGoogle DeepMind74.22026-04
Gemma 3 27BGoogle DeepMind62.62026-04
medgemma 1.5 4B itGoogle DeepMind55.72026-04
medgemma 4B itGoogle DeepMind55.72026-04

Data

This page as JSON · Edit on GitHub