Parallel four-way reading-comprehension set: 900 questions in each of 122 language variants, 109,800 items, passages from FLORES-200.
unassessed
| Category | knowledge |
|---|---|
| Subcategory | massively multilingual multiple-choice reading comprehension |
| Page status | active |
| Metric | accuracy |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 109800 |
| Dataset licence | CC-BY-SA-4.0 |
| Publisher | Meta (FAIR / facebookresearch) |
Belebele tests whether a model can read a short passage and pick the correct answer among four options. Every question is written so that it is answerable from the passage, then translated into 122 language variants (115 distinct languages, 29 scripts, 27 families). Passages come from FLORES-200, so the same 488 passages and 900 questions are aligned across languages. English alone is intended to be hard enough to separate models; the parallel design is meant to make accuracy comparable across resource levels without changing the underlying questions.
Four-way multiple-choice reading comprehension. lm-eval uses log-likelihood over A/B/C/D with English instructions and the template P/Q/A/B/C/D/Answer, zero-shot or few-shot. The paper also reports finetuning, translate-train, and cross-lingual settings that the harness group does not implement.
No model card in ModelSpec reports this benchmark yet.