Full-Duplex-Bench evaluates spoken dialogue models on pause handling, backchanneling, turn-taking and interruption management.
unassessed
| Category | multimodal |
|---|---|
| Subcategory | spoken dialogue interaction |
| Page status | active |
| Metric | task success rate |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 214 |
| Publisher | Dr. Bench authors |
Full-Duplex-Bench evaluates spoken dialogue models on pause handling, backchanneling, turn-taking and interruption management.
Interactive benchmark tasks evaluated with the released protocol.
No model card in ModelSpec reports this benchmark yet.