Fifteen-task European Spanish suite spanning commonsense, QA, NLI, translation, math, summarisation and more, run as one lm-evaluation-harness group tag.
unassessed
| Category | composite |
|---|---|
| Subcategory | fifteen-task European Spanish NLU/NLG suite (part of IberoBench) |
| Page status | active |
| Metric | per-task metric (accuracy, F1, BLEU/chrF, ROUGE, exact match; task-dependent) |
| Direction | higher_is_better |
| Unit | % |
| Publisher | Barcelona Supercomputing Center (Language Technologies Unit) |
SpanishBench is the European-Spanish slice of IberoBench, a multi-task benchmark covering five Iberian languages (Basque, Catalan, Galician, European Spanish and European Portuguese) built on lm-evaluation-harness. It combines two newly created tasks (COPA-es for commonsense causal reasoning, OpenBookQA_es for open-book QA) with adaptations of established datasets covering linguistic acceptability (EsCoLA), reading comprehension (Belebele_es), translation (FLORES_es), math word problems (MGSM_es), paraphrase identification (PAWS-X_es), natural language inference (WNLI-es, XNLI_es), summarisation (XL-Sum_es), story-ending commonsense (XStoryCloze_es), extractive QA (XQuAD_es), and a few additional tasks (Cocoteros_es, EQ-Bench_es, phrases_es) added to the harness after the original paper's publication.
Mixed by constituent task: multiple choice (COPA-es, OpenBookQA_es, XStoryCloze_es), sentence classification (EsCoLA, PAWS-X_es, WNLI-es, XNLI_es), span extraction (XQuAD_es), free-text generation scored against references (FLORES_es translation, XL-Sum_es summarisation, MGSM_es math), and other task-specific formats. lm-eval's `spanish_bench` tag runs all constituent tasks under one group; it does not fine-tune, only zero/few-shot prompt.
No model card in ModelSpec reports this benchmark yet.