N2C2-CT Matching (HELM)

MedHELM yes/no matching of n2c2 2018 patients to one inclusion criterion from de-identified notes, scored with exact match.

Also known as: N2C2-CT, n2c2 2018 Track 1, n2c2 cohort selection

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorydomain
SubcategoryHELM MedHELM wrap of n2c2 2018 clinical-trial cohort selection
Page statusactive
Metricexact_match (HELM schema); n2c2 literature uses micro/macro F1 on met vs not-met
Directionhigher_is_better
Unit%
Dataset size288
Dataset licencen2c2 data-use agreement (notes are not public); HELM code Apache-2.0
Publishern2c2 / Harvard DBMI (dataset); Stanford CRFM (HELM MedHELM scenario)

What it measures

n2c2_ct_matching is HELM's MedHELM scenario for the 2018 n2c2 Track 1 cohort-selection task. The model reads 2–5 de-identified English notes for one patient and must say whether that patient meets a single named inclusion criterion (for example ADVANCED-CAD or HBA1C). Labels are patient-level met / not met from Stubbs et al. 2019. HELM's prompt text follows Wornow et al. 2024, with expanded criterion definitions. Official gated run entries only instantiate three of the thirteen criteria. This is note-based eligibility, not de-identification and not the SHC privacy scenarios.

Task format

Zero-shot joint multiple choice. Adapter instructions "Answer A for yes, B for no." max_train_instances 0. Scenario references are the strings "yes" and "no". Run spec name n2c2_ct_matching:subject={CRITERION}. data_path must point at local train/ and test/ XML folders. get_instances currently loads only the test split.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub