BasqueBench

18 Basque-language tasks -- reused Basque NLU/QA sets plus six built for this suite -- bundled into one lm-evaluation-harness group to score base LLMs on Basque, part of the wider IberoBench project.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorycomposite
SubcategoryBasque-language multitask suite spanning question answering, NLI, paraphrase, commonsense, math and translation
Page statusactive
Metrictask-dependent: accuracy for most multiple-choice and NLI tasks; BLEU/ChrF-family scores for FLORES-eu translation directions; task-specific scoring for EusReading
Directionhigher_is_better
Unit%
Dataset licenceNot a single licence. Checking the Hugging Face cards of the suite's own component datasets directly found at least five different terms in use: CC BY-SA 4.0 (ARC-eu, MGSM-eu, Belebele's Basque config), CC BY 4.0 (XCOPA-eu), CC BY-NC 4.0 (XNLIeu), AFL 3.0 (PIQA-eu), and "other" with no further detail given (PAWS-eu); EusTrivia and BasqueGLUE's cards set no licence tag at all. Anyone reusing BasqueBench as a whole needs to check the licence of whichever specific sub-tasks they use rather than assume one licence covers the suite.
PublisherBarcelona Supercomputing Center (BSC-CNS), with Universitat Pompeu Fabra (UPF), the Centro Singular de Investigacion en Tecnoloxias Intelixentes (CiTIUS, Universidade de Santiago de Compostela), and the HiTZ Center - IXA, University of the Basque Country (UPV/EHU)

What it measures

BasqueBench is the Basque-language slice of IberoBench, a multilingual, multi-task benchmark for the official languages of the Iberian peninsula (Basque, Catalan, Galician, European Spanish and European Portuguese), built on EleutherAI's lm-evaluation-harness. Each of the five languages gets its own same-shaped suite -- BasqueBench, CatalanBench, GalicianBench, PortugueseBench and SpanishBench -- covering that language's own tasks rather than one shared multilingual task set. BasqueBench bundles 18 task groups: six built specifically for this project (ARC-eu, MGSM-eu, PAWS-eu, PIQA-eu, WNLI-eu and XCOPA-eu, translations or adaptations of established English benchmarks into Basque) alongside reused, previously published Basque resources -- the Latxa evaluation suite (EusExams, EusProficiency, EusReading, EusTrivia), Belebele's Basque config, FLORES translation pairs, BasqueGLUE's QNLIeu, XNLIeu, and XStoryCloze's Basque config. It measures broad base-model competence in Basque -- reading comprehension, commonsense inference, natural-language inference, paraphrase detection, grade-school math, and bidirectional translation -- as one aggregate suite rather than any single skill.

Task format

Mixed by sub-task: multiple-choice question answering and commonsense reasoning (ARC-eu, PIQA-eu, XCOPA-eu, Belebele-eu, EusExams, EusProficiency, EusTrivia); extractive reading comprehension (EusReading); sentence-pair classification for natural-language inference and paraphrase detection (XNLIeu, WNLI-eu, PAWS-eu, QNLIeu); free-form grade-school math word problems, in both a direct-answer and a native chain-of-thought variant (MGSM-eu); and bidirectional machine translation between Basque and eight other languages (FLORES-eu). The IberoBench paper evaluates all of these under 0-shot and 5-shot prompting.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub