Yes/no/maybe research questions answered from their source PubMed abstract, testing biomedical reading comprehension.
unassessed
| Category | domain |
|---|---|
| Subcategory | biomedical literature question answering |
| Page status | active |
| Metric | accuracy |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 1000 |
| Dataset licence | MIT |
| Publisher | University of Pittsburgh |
PubMedQA tests whether a model can answer a yes/no/maybe research question using the PubMed abstract that the question was derived from. Each item pairs a question phrased from a paper's own title or conclusion (for example, "Do preoperative statins reduce atrial fibrillation after coronary artery bypass grafting?") with that paper's abstract, and asks the model to decide whether the abstract's evidence supports a yes, no, or maybe/inconclusive answer. It is a single-turn, English-language, text-only reading-comprehension task grounded in biomedical research literature rather than clinical practice or exam questions, which sets it apart from MedQA and MedMCQA.
Three-way yes/no/maybe classification given a research question and its source PubMed abstract.
| Model | Provider | Score | Card as of |
|---|---|---|---|
| medgemma 27B it | Google DeepMind | 76.8 | 2026-04 |
| Gemma 3 27B | Google DeepMind | 73.4 | 2026-04 |
| medgemma 1.5 4B it | Google DeepMind | 73.4 | 2026-04 |
| medgemma 4B it | Google DeepMind | 73.4 | 2026-04 |