Indic DiarBench

Joint diarization and speaker-attributed ASR across all 22 scheduled Indian languages on about 108 hours of overlapping multi-speaker audio.

Also known as: IndicDiarBench, sarvamai/indic-diarbench

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorymultimodal
Subcategoryjoint speaker diarization and speaker-attributed ASR for 22 scheduled Indian languages
Page statusactive
MetricDER (no forgiveness collar, overlap included); cpWER and WDER reported alongside
Directionlower_is_better
Unit%
Dataset size1164
Dataset licenceCC-BY-4.0
PublisherSarvam AI; AI4Bharat, IIT Madras

What it measures

Indic DiarBench tests whether a speech system can transcribe Indian conversational audio and assign each word to the right speaker at the same time. Items are natural multi-speaker recordings, not read prompts. They include English code-mixing, dialectal variation, and overlapping talk. The suite spans all 22 scheduled Indian languages under three acoustic conditions: near-field virtual meetings, far-field distant microphones, and in-the-wild YouTube conversations. It measures speaker-attributed ASR, not isolated single-speaker recognition and not diarization without transcripts.

Task format

Audio in (16 kHz mono WAV). The system must emit time-aligned, speaker-labelled transcripts. All systems in the paper were given the same single-channel mixed audio. Diarization-only models such as Pyannote are excluded because they do not produce joint ASR output.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub