MTR-Bench evaluates multi-turn interactive reasoning across 40 tasks and 3,600 instances in four task classes.
unassessed
| Category | reasoning |
|---|---|
| Metric | task success |
| Direction | higher_is_better |
| Unit | score |
| Dataset size | 3600 |
| Publisher | MTR-Bench authors |
MTR-Bench tests reasoning models in interactive environments rather than isolated single-turn prompts. It covers four classes, 40 tasks, and 3,600 instances with fine-grained difficulty and multi-turn interaction.
Multi-turn text interaction with task environments.
No model card in ModelSpec reports this benchmark yet.