Full-Duplex-Bench-v3

Full-Duplex-Bench-v3 evaluates spoken language models under naturalistic human audio, disfluencies and chained tool use.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorymultimodal
Subcategorynaturalistic spoken agents
Page statusactive
Metrictask success rate
Directionhigher_is_better
Unit%
Dataset size214
PublisherDr. Bench authors

What it measures

Full-Duplex-Bench-v3 evaluates spoken language models under naturalistic human audio, disfluencies and chained tool use.

Task format

Interactive benchmark tasks evaluated with the released protocol.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub