SpanishBench

Fifteen-task European Spanish suite spanning commonsense, QA, NLI, translation, math, summarisation and more, run as one lm-evaluation-harness group tag.

Also known as: spanish_bench

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorycomposite
Subcategoryfifteen-task European Spanish NLU/NLG suite (part of IberoBench)
Page statusactive
Metricper-task metric (accuracy, F1, BLEU/chrF, ROUGE, exact match; task-dependent)
Directionhigher_is_better
Unit%
PublisherBarcelona Supercomputing Center (Language Technologies Unit)

What it measures

SpanishBench is the European-Spanish slice of IberoBench, a multi-task benchmark covering five Iberian languages (Basque, Catalan, Galician, European Spanish and European Portuguese) built on lm-evaluation-harness. It combines two newly created tasks (COPA-es for commonsense causal reasoning, OpenBookQA_es for open-book QA) with adaptations of established datasets covering linguistic acceptability (EsCoLA), reading comprehension (Belebele_es), translation (FLORES_es), math word problems (MGSM_es), paraphrase identification (PAWS-X_es), natural language inference (WNLI-es, XNLI_es), summarisation (XL-Sum_es), story-ending commonsense (XStoryCloze_es), extractive QA (XQuAD_es), and a few additional tasks (Cocoteros_es, EQ-Bench_es, phrases_es) added to the harness after the original paper's publication.

Task format

Mixed by constituent task: multiple choice (COPA-es, OpenBookQA_es, XStoryCloze_es), sentence classification (EsCoLA, PAWS-X_es, WNLI-es, XNLI_es), span extraction (XQuAD_es), free-text generation scored against references (FLORES_es translation, XL-Sum_es summarisation, MGSM_es math), and other task-specific formats. lm-eval's `spanish_bench` tag runs all constituent tasks under one group; it does not fine-tune, only zero/few-shot prompt.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub