MTR-Suite audits and synthesizes conversational retrieval benchmarks and introduces a production-style MTR-Bench.
unassessed
| Category | composite |
|---|---|
| Metric | retrieval quality |
| Direction | higher_is_better |
| Unit | score |
| Publisher | MTR-Suite authors |
MTR-Suite evaluates conversational retrieval through an LLM auditor, a multi-agent dialogue synthesis pipeline, and a general-domain benchmark. It targets hard topic switches, verbosity, and alignment gaps in existing retrieval tests.
Conversational retrieval with synthetic and audited multi-turn dialogues.
No model card in ModelSpec reports this benchmark yet.