RO-N3WS

A 126-hour Romanian ASR set that trains on broadcast news and tests on news plus audiobook, film, story, and podcast speech, scored by word error rate.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorydomain
SubcategoryRomanian automatic speech recognition
Page statusproposed
Metricword error rate (WER)
Directionlower_is_better
Unit%
Dataset size74134
PublisherDepartment of Computer Science, University of Bucharest

What it measures

RO-N3WS measures Romanian speech-to-text under domain shift. The in-domain half is studio and field news from ProTV and Antena 1 (Observator). The out-of-distribution half is literary audiobooks, Romanian film dialogue, children's stories, and podcasts. Clips are short, usually under ten seconds. Transcripts restore diacritics, expand spoken numbers, and keep named entities as pronounced. The skill is transcription robustness, not dialogue or translation.

Task format

A Romanian audio clip in; the system returns a word sequence. WER is computed against a manually corrected transcript, with extra references for number and punctuation variants on commercial APIs.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub