Medbullets

308 USMLE Step 2/3-style clinical vignette questions with expert explanations, built specifically to be harder than MedQA and to test explanation quality, not just the answer letter.

Also known as: MedBullets, MedBullet

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorydomain
SubcategoryUSMLE Step 2/3-style clinical multiple-choice questions with expert explanations
Page statusactive
Metricaccuracy (plus separate explanation-quality evaluation)
Directionhigher_is_better
Unit%
Dataset size308
Dataset licenceNot stated: the GitHub repository carries no licence file or badge, and the underlying question bank is republished from the Medbullets USMLE-preparation website rather than authored by the paper's authors.
PublisherRice University and Johns Hopkins University

What it measures

Medbullets presents a clinical vignette in the style of USMLE Step 2 and Step 3 (later-stage, more applied medical-licensing exam content than Step 1) and asks the model to choose the correct diagnosis or management step from four or five answer options. Each item also carries an expert-written explanation, because the paper that introduced Medbullets was built specifically to let researchers evaluate the quality of a model's reasoning, not only whether it picked the right letter. The underlying questions are drawn from the existing Medbullets USMLE-preparation question bank rather than written for the paper, so they reflect realistic, simulated clinical scenarios rather than deliberately adversarial ones.

Task format

Four-option ("op4") or five-option ("op5") multiple-choice clinical vignette question; the model returns a single letter answer, typically zero- or few-shot, sometimes paired with a free-text explanation that is separately evaluated.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub