Full-Duplex-Bench-v3 evaluates spoken language models under naturalistic human audio, disfluencies and chained tool use.
unassessed
| Category | multimodal |
|---|---|
| Subcategory | naturalistic spoken agents |
| Page status | active |
| Metric | task success rate |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 214 |
| Publisher | Dr. Bench authors |
Full-Duplex-Bench-v3 evaluates spoken language models under naturalistic human audio, disfluencies and chained tool use.
Interactive benchmark tasks evaluated with the released protocol.
No model card in ModelSpec reports this benchmark yet.