PubMedQA

Yes/no/maybe research questions answered from their source PubMed abstract, testing biomedical reading comprehension.

Also known as: PQA-L

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorydomain
Subcategorybiomedical literature question answering
Page statusactive
Metricaccuracy
Directionhigher_is_better
Unit%
Dataset size1000
Dataset licenceMIT
PublisherUniversity of Pittsburgh

What it measures

PubMedQA tests whether a model can answer a yes/no/maybe research question using the PubMed abstract that the question was derived from. Each item pairs a question phrased from a paper's own title or conclusion (for example, "Do preoperative statins reduce atrial fibrillation after coronary artery bypass grafting?") with that paper's abstract, and asks the model to decide whether the abstract's evidence supports a yes, no, or maybe/inconclusive answer. It is a single-turn, English-language, text-only reading-comprehension task grounded in biomedical research literature rather than clinical practice or exam questions, which sets it apart from MedQA and MedMCQA.

Task format

Three-way yes/no/maybe classification given a research question and its source PubMed abstract.

Models reporting this benchmark

These figures come from the model cards, which carry one collection date per card and no per-score attribution. They are shown as reported, not as verified evidence.
ModelProviderScoreCard as of
medgemma 27B itGoogle DeepMind76.82026-04
Gemma 3 27BGoogle DeepMind73.42026-04
medgemma 1.5 4B itGoogle DeepMind73.42026-04
medgemma 4B itGoogle DeepMind73.42026-04

Data

This page as JSON · Edit on GitHub