MTR-Bench

MTR-Bench evaluates multi-turn interactive reasoning across 40 tasks and 3,600 instances in four task classes.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryreasoning
Metrictask success
Directionhigher_is_better
Unitscore
Dataset size3600
PublisherMTR-Bench authors

What it measures

MTR-Bench tests reasoning models in interactive environments rather than isolated single-turn prompts. It covers four classes, 40 tasks, and 3,600 instances with fine-grained difficulty and multi-turn interaction.

Task format

Multi-turn text interaction with task environments.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub