CaBBQ (Catalan Bias Benchmark for Question Answering)

Catalan multiple-choice QA, parallel with EsBBQ, testing stereotype use under ambiguous context and accuracy once the context names the answer.

Also known as: Catalan Bias Benchmark for Question Answering, Catalan BBQ

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorysafety
Subcategorysocial bias in Catalan question answering
Page statusactive
Metricacc_ambig and acc_disambig (higher better); bias_score_ambig and bias_score_disambig (near 0 better)
Directionhigher_is_better
Unit%
Dataset size27320
Dataset licenceCC-BY-4.0
PublisherBarcelona Supercomputing Center, Language Technologies Unit

What it measures

CaBBQ asks whether a model falls back on stereotypes about groups in Spain when a Catalan context is under-informative, and whether it can ignore that stereotype once a disambiguating sentence is added. Each item has a context, a question, and three answers (two groups plus an unknown option). Ten categories cover age, disability, gender, LGBTQIA, nationality, appearance, race/ethnicity, religion, SES, and Spanish region. It is a cultural adaptation of English BBQ, not a translation-only copy.

Task format

Three-way multiple-choice QA in Catalan. lm-eval group `cabbq` scores by log-likelihood over ans0, ans1, and several unknown phrasings collapsed to index 2. Prompts use Context / Pregunta / Resposta.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub