MTR-Suite

MTR-Suite audits and synthesizes conversational retrieval benchmarks and introduces a production-style MTR-Bench.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorycomposite
Metricretrieval quality
Directionhigher_is_better
Unitscore
PublisherMTR-Suite authors

What it measures

MTR-Suite evaluates conversational retrieval through an LLM auditor, a multi-agent dialogue synthesis pipeline, and a general-domain benchmark. It targets hard topic switches, verbosity, and alignment gaps in existing retrieval tests.

Task format

Conversational retrieval with synthetic and audited multi-turn dialogues.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub