MIMIC-BHC (MedHELM)

MedHELM's gated wrap of MIMIC-IV-BHC: write a Brief Hospital Course from a discharge note, scored by an LLM jury plus overlap metrics.

Also known as: MIMIC-IV-BHC, MIMIC-IV-Ext-BHC, Brief Hospital Course

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorydomain
Subcategorydischarge-note summarization into a Brief Hospital Course
Page statusactive
Metricmimic_bhc_accuracy (HELM LLM-jury average of accuracy, completeness, clarity, each 1-5)
Directionhigher_is_better
Unitpoints
Dataset size270033
Dataset licencePhysioNet Credentialed Health Data License 1.5.0 (DUA 1.5.0; CITI training required)
PublisherStanford University (MIMI / CRFM MedHELM packaging)

What it measures

This id is HELM's mimic_bhc scenario, not the authors' original BLEU and BERTScore study by itself. MIMIC-IV-BHC pairs a preprocessed MIMIC-IV discharge note with the Brief Hospital Course section of that stay. HELM prompts the model to summarize the note into a BHC (zero-shot, max 1,024 tokens) and grades the text with an LLM jury on accuracy, completeness and clarity (1–5) as mimic_bhc_accuracy, while also logging summarization overlap metrics. Inputs and outputs are English clinical text.

Task format

Zero-shot generation. HELM instructions: "Summarize the clinical note into a brief hospital course." Input noun Clinical Note, output noun Brief Hospital Course. max_train_instances=0, max_tokens=1024.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub