Belebele

Parallel four-way reading-comprehension set: 900 questions in each of 122 language variants, 109,800 items, passages from FLORES-200.

Also known as: The Belebele Benchmark

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryknowledge
Subcategorymassively multilingual multiple-choice reading comprehension
Page statusactive
Metricaccuracy
Directionhigher_is_better
Unit%
Dataset size109800
Dataset licenceCC-BY-SA-4.0
PublisherMeta (FAIR / facebookresearch)

What it measures

Belebele tests whether a model can read a short passage and pick the correct answer among four options. Every question is written so that it is answerable from the passage, then translated into 122 language variants (115 distinct languages, 29 scripts, 27 families). Passages come from FLORES-200, so the same 488 passages and 900 questions are aligned across languages. English alone is intended to be hard enough to separate models; the parallel design is meant to make accuracy comparable across resource levels without changing the underlying questions.

Task format

Four-way multiple-choice reading comprehension. lm-eval uses log-likelihood over A/B/C/D with English instructions and the template P/Q/A/B/C/D/Answer, zero-shot or few-shot. The paper also reports finetuning, translate-train, and cross-lingual settings that the harness group does not implement.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub