SE-Eval

Human-rated speech-editing set of 9,151 clips with 1-5 MOS and consistency labels, released as ground truth for scoring speech editors while the ICME 2026 paper stays anonymous.

Also known as: SE-Eval

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorymultimodal
Subcategoryspeech editing quality: MOS, boundary naturalness, and contextual consistency
Page statusproposed
Metricmean opinion score and consistency scores on a 1-5 scale
Directionhigher_is_better
Unitpoints (1-5)
Dataset size9151
Dataset licenceCC-BY-4.0

What it measures

SE-Eval scores speech editors on local edits, not on full-utterance TTS quality. Each item pairs a source clip with a model-edited clip and the source and target transcripts. Human raters mark overall quality, join-point naturalness, and whether environment, prosody, or emotion stayed consistent after the edit. Audio is English speech. The Hub card presents the labels as ground truth for automatic evaluators, not as a live leaderboard of new systems.

Task format

A rater hears original and edited speech with the intended transcript and assigns 1-5 scores on the dimensions that apply to that sub-domain. RealEdit and LongHard items carry overall MOS and boundary MOS. Environment, Prosody, and Emotion items carry the matching consistency score. Protocol details sit in the still-anonymous ICME 2026 paper and are not restated on the dataset card.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub