MULTI-Bench evaluates spoken dialogue models on multi-turn emotional intelligence through basic understanding and advanced support tracks.
unassessed
| Category | human-preference |
|---|---|
| Subcategory | spoken emotional intelligence |
| Metric | task performance |
| Direction | higher_is_better |
| Unit | score |
| Dataset size | 3200 |
| Publisher | MULTI-Bench authors |
MULTI-Bench tests whether spoken dialogue models sustain interactive conversations with emotional awareness. Its basic track covers emotion understanding and reasoning; its advanced track covers emotion support and application.
Multi-turn spoken dialogue across five tasks and eight subsets.
No model card in ModelSpec reports this benchmark yet.