Open-ended consumer questions about medications paired with trusted reference answers drawn from DailyMed, MedlinePlus and similar sources.
unassessed
| Category | domain |
|---|---|
| Subcategory | consumer medication question answering |
| Page status | active |
| Metric | medication_qa_accuracy (HELM LLM-jury average of accuracy, completeness and clarity, each 1-5) |
| Direction | higher_is_better |
| Unit | points |
| Dataset size | 674 |
| Dataset licence | CC-BY-4.0 on the GitHub README for the dataset; the MEDINFO 2019 article itself is CC BY-NC 4.0 |
| Publisher | Lister Hill National Center for Biomedical Communications, U.S. National Library of Medicine; HELM scenario by Stanford CRFM |
MedicationQA tests whether a model can answer a real consumer question about a drug in plain English. Questions come from MedlinePlus users and always have a drug as the focus. Types include information, dose, usage, side effects, indication and interaction, among 25 labels in the gold standard. Each item pairs the question with one expert-retrieved reference answer and its source URL. HELM treats the whole set as a zero-shot generation task. It is not MedQA, not MEDIQA 2019 ranking, and not a multiple-choice exam.
Zero-shot generation. HELM's instruction is "Please answer the following consumer health question." The model writes free text (max 512 tokens in the HELM run spec).
No model card in ModelSpec reports this benchmark yet.