A 126-hour Romanian ASR set that trains on broadcast news and tests on news plus audiobook, film, story, and podcast speech, scored by word error rate.
unassessed
| Category | domain |
|---|---|
| Subcategory | Romanian automatic speech recognition |
| Page status | proposed |
| Metric | word error rate (WER) |
| Direction | lower_is_better |
| Unit | % |
| Dataset size | 74134 |
| Publisher | Department of Computer Science, University of Bucharest |
RO-N3WS measures Romanian speech-to-text under domain shift. The in-domain half is studio and field news from ProTV and Antena 1 (Observator). The out-of-distribution half is literary audiobooks, Romanian film dialogue, children's stories, and podcasts. Clips are short, usually under ten seconds. Transcripts restore diacritics, expand spoken numbers, and keep named entities as pronounced. The skill is transcription robustness, not dialogue or translation.
A Romanian audio clip in; the system returns a word sequence. WER is computed against a manually corrected transcript, with extra references for number and punctuation variants on commercial APIs.
No model card in ModelSpec reports this benchmark yet.