MedHELM wrap of mixed-specialty MTSamples notes: generate a treatment plan from a clinical transcription, scored by an LLM jury plus overlap metrics.
unassessed
| Category | domain |
|---|---|
| Subcategory | treatment-plan generation from mixed-specialty transcribed clinical reports |
| Page status | active |
| Metric | mtsamples_replicate_accuracy (HELM LLM-jury average of accuracy, completeness, clarity, each 1-5) |
| Direction | higher_is_better |
| Unit | points |
| Publisher | Stanford CRFM / MedHELM; source notes from MTSamples.com, packaged by raulista1997/benchmarkdata |
This id is HELM's mtsamples_replicate scenario in MedHELM, not the surgical-only mtsamples_procedures wrap and not a third-party MTSamples leaderboard. Each item is an English transcribed report from MTSamples.com, stored as mtsamples_processed on raulista1997/benchmarkdata. HELM prefers a PLAN section as the reference, else SUMMARY, else FINDINGS, and removes only PLAN from the prompt. MedHELM places the task in clinical decision support / planning treatments. Display name on the schema is MTSamples.
Zero-shot generation. HELM instructions: "Given various information about a patient, return a reasonable treatment plan for the patient." No input noun. Output noun Answer. max_train_instances=0, max_tokens=512.
No model card in ModelSpec reports this benchmark yet.