M3-DuplexBench

M3-DuplexBench evaluates multilingual full-duplex spoken dialogue with turn-taking, backchannels, and user interruptions.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorygeneration
Subcategoryspoken dialogue
Metricdialogue performance
Directionhigher_is_better
Unitscore
PublisherM3-DuplexBench authors

What it measures

M3-DuplexBench tests spoken dialogue systems that listen while speaking. It covers English and Japanese, casual conversation and multi-turn question answering, and multiple context conditions for studying dialogue history.

Task format

Multi-turn audio dialogue with full-duplex interaction and context variants.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub