The IberoBench suite for Galician: 11 lm-evaluation-harness task groups mixing three natively built Galician resources with eight translated from English or multilingual sources.
unassessed
| Category | composite |
|---|---|
| Subcategory | Galician-language multitask suite: linguistic acceptability, QA, NLI, paraphrase, summarisation, commonsense reasoning, math and translation |
| Page status | active |
| Metric | task-dependent: accuracy for most multiple-choice, NLI and classification tasks; ROUGE for summarisation; BLEU/ChrF-family scores for flores_gl |
| Direction | higher_is_better |
| Unit | % |
| Dataset licence | Mostly CC BY 4.0 (galcola, summarization_gl, parafrases_gl, PAWS-gl, openbookqa_gl, mgsm_gl, xstorycloze_gl and belebele_gl, confirmed individually via the Hugging Face API), with two exceptions: truthfulqa_gl is Apache-2.0 and xnli_gl is CC BY-NC 4.0. This page did not check every component individually, so treat any single suite-wide licence claim with caution. |
| Publisher | Proxecto Nos (Galician-language AI initiative run through the Xunta de Galicia and the CiTIUS research centre at the Universidade de Santiago de Compostela), within the IberoBench project led by the Barcelona Supercomputing Center (BSC-CNS) |
GalicianBench is the Galician-language slice of IberoBench, the same project behind this repository's CatalanBench and BasqueBench pages, covering the official languages of the Iberian peninsula. It bundles 11 top-level lm-evaluation-harness task groups (one of which, flores_gl, itself expands into 16 directional translation subtasks). Unlike some sibling suites, provenance splits cleanly and was confirmed dataset by dataset: GalCoLA (linguistic acceptability, 17,088 sentences), summarization_gl (80,829 native news-article/summary pairs from three Galician outlets) and parafrases_gl (2,032 sentence pairs, sourced from Galician Wikipedia, novels and parliamentary sessions, with paraphrase variants generated by term replacement and back-translation then manually reviewed by two linguists, per the IberoBench paper's own dataset description) were built directly in Galician; openbookqa_gl, mgsm_direct_gl, xstorycloze_gl, truthfulqa_gl, xnli_gl and paws_gl are explicit translations of their English originals (OpenBookQA, MGSM, StoryCloze, TruthfulQA, XNLI, PAWS); and belebele_glg_Latn and flores_gl are Galician configurations of professionally translated multilingual suites (Belebele, FLORES).
Mixed by sub-task: four-option multiple-choice for openbookqa_gl; binary acceptability classification for galcola; three-way paraphrase classification for parafrases_gl and binary for paws_gl; three-way natural-language-inference classification for xnli_gl; two-ending narrative completion for xstorycloze_gl; free-form generation for truthfulqa_gl's generation split and summarization_gl; free-form grade-school math word problems for mgsm_direct_gl; multiple-choice reading comprehension for belebele_glg_Latn; and bidirectional machine translation between Galician and eight other languages for flores_gl.
No model card in ModelSpec reports this benchmark yet.