WMT 14

HELM's WMT_14 scenario scores machine translation on five WMT14 English language pairs using sentence-level BLEU-4.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorytranslation
Subcategorymachine translation
Page statusactive
MetricBLEU-4
Directionhigher_is_better
Unitpoints

What it measures

HELM's WMT_14 scenario evaluates machine translation across five English-paired language directions from the 2014 Workshop on Statistical Machine Translation shared task.

Task format

Text generation; the model is given a source-language sentence and must generate the translation in the target language, for each of five language pairs.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub