MINT (Medical Incremental N-Turn Benchmark)

MINT shards 1,035 medical cases into multi-turn evidence to test whether models diagnose too early, self-correct, or get lured by lab results.

Also known as: MINT, Medical Incremental N-Turn Benchmark, Benchmarking Multi-turn Medical Diagnosis

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorydomain
Subcategorymulti-turn medical diagnosis (hold, lure, self-correction)
Page statusactive
Metricdiagnostic accuracy (multiple-choice)
Directionhigher_is_better
Unit%
Dataset size1035
PublisherThe University of Texas at Austin; New York University; UT Southwestern; UNC Chapel Hill; University of Illinois Urbana-Champaign

What it measures

MINT (Medical Incremental N-Turn Benchmark) tests diagnostic multiple-choice accuracy when clinical evidence arrives in sequence instead of as one vignette. Each of 1,035 cases is split into labeled shards (history, exam, labs, imaging, and related categories). The model may hold, answer, or revise as shards appear. The authors isolate three behaviours: premature commitment (hold), later correction of a wrong answer (self-correction), and early answering triggered by salient labs (lure). English clinical text; no images in the released description.

Task format

Multi-turn chat over evidence shards, with a multiple-choice diagnostic question and answer options. Main contrast: Ask-Question-First (question shown up front; model may wait or answer at any turn) versus Ask-Question-Last (question withheld until the final turn). Also FULL versus CONCAT single-turn controls. Turn-count variants fix 4, 8, 12, or 16 turns.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub