Joint diarization and speaker-attributed ASR across all 22 scheduled Indian languages on about 108 hours of overlapping multi-speaker audio.
unassessed
| Category | multimodal |
|---|---|
| Subcategory | joint speaker diarization and speaker-attributed ASR for 22 scheduled Indian languages |
| Page status | active |
| Metric | DER (no forgiveness collar, overlap included); cpWER and WDER reported alongside |
| Direction | lower_is_better |
| Unit | % |
| Dataset size | 1164 |
| Dataset licence | CC-BY-4.0 |
| Publisher | Sarvam AI; AI4Bharat, IIT Madras |
Indic DiarBench tests whether a speech system can transcribe Indian conversational audio and assign each word to the right speaker at the same time. Items are natural multi-speaker recordings, not read prompts. They include English code-mixing, dialectal variation, and overlapping talk. The suite spans all 22 scheduled Indian languages under three acoustic conditions: near-field virtual meetings, far-field distant microphones, and in-the-wild YouTube conversations. It measures speaker-attributed ASR, not isolated single-speaker recognition and not diarization without transcripts.
Audio in (16 kHz mono WAV). The system must emit time-aligned, speaker-labelled transcripts. All systems in the paper were given the same single-channel mixed audio. Diarization-only models such as Pyannote are excluded because they do not produce joint ASR output.
No model card in ModelSpec reports this benchmark yet.