MINT shards 1,035 medical cases into multi-turn evidence to test whether models diagnose too early, self-correct, or get lured by lab results.
unassessed
| Category | domain |
|---|---|
| Subcategory | multi-turn medical diagnosis (hold, lure, self-correction) |
| Page status | active |
| Metric | diagnostic accuracy (multiple-choice) |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 1035 |
| Publisher | The University of Texas at Austin; New York University; UT Southwestern; UNC Chapel Hill; University of Illinois Urbana-Champaign |
MINT (Medical Incremental N-Turn Benchmark) tests diagnostic multiple-choice accuracy when clinical evidence arrives in sequence instead of as one vignette. Each of 1,035 cases is split into labeled shards (history, exam, labs, imaging, and related categories). The model may hold, answer, or revise as shards appear. The authors isolate three behaviours: premature commitment (hold), later correction of a wrong answer (self-correction), and early answering triggered by salient labs (lure). English clinical text; no images in the released description.
Multi-turn chat over evidence shards, with a multiple-choice diagnostic question and answer options. Main contrast: Ask-Question-First (question shown up front; model may wait or answer at any turn) versus Ask-Question-Last (question withheld until the final turn). Also FULL versus CONCAT single-turn controls. Turn-count variants fix 4, 8, 12, or 16 turns.
No model card in ModelSpec reports this benchmark yet.