MedHELM yes/no matching of n2c2 2018 patients to one inclusion criterion from de-identified notes, scored with exact match.
unassessed
| Category | domain |
|---|---|
| Subcategory | HELM MedHELM wrap of n2c2 2018 clinical-trial cohort selection |
| Page status | active |
| Metric | exact_match (HELM schema); n2c2 literature uses micro/macro F1 on met vs not-met |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 288 |
| Dataset licence | n2c2 data-use agreement (notes are not public); HELM code Apache-2.0 |
| Publisher | n2c2 / Harvard DBMI (dataset); Stanford CRFM (HELM MedHELM scenario) |
n2c2_ct_matching is HELM's MedHELM scenario for the 2018 n2c2 Track 1 cohort-selection task. The model reads 2–5 de-identified English notes for one patient and must say whether that patient meets a single named inclusion criterion (for example ADVANCED-CAD or HBA1C). Labels are patient-level met / not met from Stubbs et al. 2019. HELM's prompt text follows Wornow et al. 2024, with expanded criterion definitions. Official gated run entries only instantiate three of the thirteen criteria. This is note-based eligibility, not de-identification and not the SHC privacy scenarios.
Zero-shot joint multiple choice. Adapter instructions "Answer A for yes, B for no." max_train_instances 0. Scenario references are the strings "yes" and "no". Run spec name n2c2_ct_matching:subject={CRITERION}. data_path must point at local train/ and test/ XML folders. get_instances currently loads only the test split.
No model card in ModelSpec reports this benchmark yet.