MedHELM wrap of MTSamples surgical notes: generate a plan or findings from an operative transcription, scored by an LLM jury plus overlap metrics.
unassessed
| Category | domain |
|---|---|
| Subcategory | operative-note plan, summary or findings generation from a surgical transcription |
| Page status | active |
| Metric | mtsamples_procedures_accuracy (HELM LLM-jury average of accuracy, completeness, clarity, each 1-5) |
| Direction | higher_is_better |
| Unit | points |
| Publisher | Stanford CRFM / MedHELM; source notes from MTSamples.com, packaged by raulista1997/benchmarkdata |
This id is HELM's mtsamples_procedures scenario in MedHELM, not a standalone shared-task paper and not the mixed-specialty mtsamples_replicate scenario. Each item is an English transcribed operative note from MTSamples.com, copied into raulista1997/benchmarkdata. HELM strips PLAN, SUMMARY and FINDINGS from the prompt and asks the model for a treatment plan. The reference is the first of those three sections that exists. Nature Medicine lists the task under clinical note generation / recording procedures.
Zero-shot generation. HELM instructions: "Here are information about a patient, return a reasonable treatment plan for the patient." Input noun Patient Notes, output noun Answer. max_train_instances=0, max_tokens=512.
No model card in ModelSpec reports this benchmark yet.