MTSamples Replicate (MedHELM)

MedHELM wrap of mixed-specialty MTSamples notes: generate a treatment plan from a clinical transcription, scored by an LLM jury plus overlap metrics.

Also known as: MTSamples, mtsamples_processed

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorydomain
Subcategorytreatment-plan generation from mixed-specialty transcribed clinical reports
Page statusactive
Metricmtsamples_replicate_accuracy (HELM LLM-jury average of accuracy, completeness, clarity, each 1-5)
Directionhigher_is_better
Unitpoints
PublisherStanford CRFM / MedHELM; source notes from MTSamples.com, packaged by raulista1997/benchmarkdata

What it measures

This id is HELM's mtsamples_replicate scenario in MedHELM, not the surgical-only mtsamples_procedures wrap and not a third-party MTSamples leaderboard. Each item is an English transcribed report from MTSamples.com, stored as mtsamples_processed on raulista1997/benchmarkdata. HELM prefers a PLAN section as the reference, else SUMMARY, else FINDINGS, and removes only PLAN from the prompt. MedHELM places the task in clinical decision support / planning treatments. Display name on the schema is MTSamples.

Task format

Zero-shot generation. HELM instructions: "Given various information about a patient, return a reasonable treatment plan for the patient." No input noun. Output noun Answer. max_train_instances=0, max_tokens=512.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub