MULTI-Bench

MULTI-Bench evaluates spoken dialogue models on multi-turn emotional intelligence through basic understanding and advanced support tracks.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryhuman-preference
Subcategoryspoken emotional intelligence
Metrictask performance
Directionhigher_is_better
Unitscore
Dataset size3200
PublisherMULTI-Bench authors

What it measures

MULTI-Bench tests whether spoken dialogue models sustain interactive conversations with emotional awareness. Its basic track covers emotion understanding and reasoning; its advanced track covers emotion support and application.

Task format

Multi-turn spoken dialogue across five tasks and eight subsets.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub