Human-rated speech-editing set of 9,151 clips with 1-5 MOS and consistency labels, released as ground truth for scoring speech editors while the ICME 2026 paper stays anonymous.
unassessed
| Category | multimodal |
|---|---|
| Subcategory | speech editing quality: MOS, boundary naturalness, and contextual consistency |
| Page status | proposed |
| Metric | mean opinion score and consistency scores on a 1-5 scale |
| Direction | higher_is_better |
| Unit | points (1-5) |
| Dataset size | 9151 |
| Dataset licence | CC-BY-4.0 |
SE-Eval scores speech editors on local edits, not on full-utterance TTS quality. Each item pairs a source clip with a model-edited clip and the source and target transcripts. Human raters mark overall quality, join-point naturalness, and whether environment, prosody, or emotion stayed consistent after the edit. Audio is English speech. The Hub card presents the labels as ground truth for automatic evaluators, not as a live leaderboard of new systems.
A rater hears original and edited speech with the intended transcript and assigns 1-5 scores on the dimensions that apply to that sub-domain. RealEdit and LongHard items carry overall MOS and boundary MOS. Environment, Prosody, and Emotion items carry the matching consistency score. Protocol details sit in the still-anonymous ICME 2026 paper and are not restated on the dataset card.
No model card in ModelSpec reports this benchmark yet.