Translation Tasks

A family of translation tasks configured in lm-evaluation-harness.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorytranslation
Subcategorymachine translation
Page statusactive
MetricBLEU, TER, and chrF (reported together for the WMT14/WMT16 groups; other groups may differ)
Directionhigher_is_better
PublisherEleutherAI

What it measures

The group covers translation evaluations including WMT14, WMT16, WMT20, and IWSLT2017 task groups.

Task format

Generate a translation for a source sentence.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub