BLEU score for English-to-Chinese translation on the FLORES-200 devtest set, as reported in this repository's model cards.
unassessed
| Category | translation |
|---|---|
| Subcategory | en-zh |
| Page status | active |
| Metric | BLEU |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 1012 |
| Dataset licence | CC BY-SA 4.0 |
| Publisher | Meta AI (FAIR), NLLB Team; now maintained by the Open Language Data Initiative (OLDI) |
flores_en_zh scores how well a model translates the FLORES-200 devtest sentences from English into Chinese. FLORES-200 tracks Simplified and Traditional Chinese as distinct language codes (zho_Hans and zho_Hant); this repository's card data does not state which variant its flores_en_zh figures use.
Translate the 1012 FLORES-200 devtest sentences from English into Chinese; score against the human Chinese reference.
| Model | Provider | Score | Card as of |
|---|---|---|---|
| GPT-4 | OpenAI | 59.7 | 2026-04 |
| GPT-4 Turbo | OpenAI | 59.7 | 2026-04 |
| GPT-4o | OpenAI | 59.7 | 2026-04 |
| GPT-4o (2024-05-13) | OpenAI | 59.7 | 2026-04 |
| GPT-4o (2024-08-06) | OpenAI | 59.7 | 2026-04 |
| GPT-4o (2024-11-20) | OpenAI | 59.7 | 2026-04 |
| GPT-4o mini | OpenAI | 59.7 | 2026-04 |
| Claude Sonnet 3.5 | Anthropic | 59.2 | 2026-04 |
| Claude Sonnet 3.5 v2 | Anthropic | 59.2 | 2026-04 |
| Gemini 1.5 Flash | Google DeepMind | 56.8 | 2026-04 |
| Gemini 1.5 Flash-8B | Google DeepMind | 56.8 | 2026-04 |
| Gemini 1.5 Pro | Google DeepMind | 56.8 | 2026-04 |
| Gemini 2.0 Flash | Google DeepMind | 56.8 | 2026-04 |
| Gemini 2.5 Flash | Google DeepMind | 56.8 | 2026-04 |
| Gemini 2.5 Pro | Google DeepMind | 56.8 | 2026-04 |