MedCalc-Bench

Tests whether a model can compute clinical values (dosages, risk scores, dates) from a patient note the way a bedside medical calculator would, graded by exact match or numeric tolerance.

Also known as: MedCalc_Bench

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorydomain
Subcategoryclinical/medical calculation
Page statusactive
Metricaccuracy (exact match for rule-based and date calculators; within 5% tolerance for equation-based lab/physical/dosage calculators)
Directionhigher_is_better
Unit%
Dataset size11643
Dataset licenceCC-BY-SA-4.0

What it measures

MedCalc-Bench gives a model a patient note and a question naming a specific clinical calculation — a creatinine clearance, a Glasgow Coma Score, a gestational age, a drug dosage — and asks it to extract the relevant values from the note, apply the correct formula or clinical rule, and produce the answer a bedside calculator would give. This targets a narrower, more mechanical skill than open-ended medical question answering: correctly identifying which numbers in a note matter, and computing with them correctly, rather than recalling medical facts in prose.

Task format

Given a patient note (drawn from case reports or clinical vignettes) and a question naming one of 55 calculators, the model must extract relevant entities and output a numeric or categorical answer with its computation shown.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub