CatalanBench

An EleutherAI lm-evaluation-harness suite aggregating roughly two dozen Catalan- and Valencian-language tasks -- QA, NLI, commonsense, paraphrase, summarisation and translation -- built and funded by BSC's Projecte AINA.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorycomposite
Subcategorynative and adapted Catalan/Valencian language suite: QA, NLI, commonsense, paraphrase, summarisation and translation
Page statusactive
Metricvaries by sub-task: accuracy or F1 for most classification/QA tasks, ROUGE for summarisation, spBLEU for translation
Directionhigher_is_better
Unit%
Dataset licenceVaries by sub-task (for example CC-BY-SA-4.0 for arc_ca, per its own Hugging Face card); this page did not check all roughly two dozen constituent datasets individually, so one suite-wide licence is not established.
PublisherBarcelona Supercomputing Center (BSC-CNS); most individual constituent datasets are curated by BSC's Language Technologies Unit and funded by Projecte AINA. The umbrella IberoBench survey adds three further institutions: Universitat Pompeu Fabra (UPF), the Centro Singular de Investigacion en Tecnoloxias Intelixentes (CiTIUS-USC), and the HiTZ Center - IXA, University of the Basque Country (UPV/EHU).

What it measures

CatalanBench aggregates existing and purpose-built datasets to test Catalan, and for a handful of tasks Valencian, language understanding and generation across many different skills rather than one. Implemented in EleutherAI's lm-evaluation-harness as the `catalan_bench` task group, it currently spans roughly two dozen sub-tasks: multiple-choice science and commonsense reasoning (ARC_ca, OpenBookQA_ca, PIQA_ca, SIQA_ca, COPA-ca, XStoryCloze_ca), extractive and generative question answering (CatalanQA, CoQCat, XQuAD-ca, TerretaQA), natural language inference (TE-ca, WNLI-ca, XNLI-ca, XNLI-va), paraphrase and linguistic-acceptability judgement (Parafraseja, PAWS-ca, CatCoLA), reading comprehension (Belebele_ca), summarisation (caBREU), truthfulness (TruthfulQA_va, VeritasQA_ca), Catalan-translated grade-school math word problems (MGSM_ca), a Valencian language-proficiency exam (CieaCOVA) and Catalan-to/from-seven-other-language machine translation (FLORES_ca). Some constituent datasets were purpose-built for this suite by Projecte AINA; others are pre-existing multilingual datasets (Belebele, FLORES, XNLI among them) with a Catalan or Valencian configuration added.

Task format

Varies by sub-task: four- or five-option multiple-choice for the QA and commonsense tasks, binary or three-way classification for the NLI and acceptability tasks, span extraction or short free-text generation for the remaining QA and summarisation tasks, and sentence-level machine translation for the FLORES_ca directions. lm-evaluation-harness runs each sub-task independently and additionally exposes a `catalan_bench` group covering all of them together.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub