BBQ-Lite (Bias Benchmark for QA, BIG-bench)

BIG-bench BBQ subset: 16,076 three-choice US-English QA items that test social stereotypes in ambiguous and disambiguated contexts.

Also known as: BBQ Lite, Bias Benchmark for QA Lite

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorysafety
Subcategorysocial-bias three-choice QA (ambiguous vs disambiguated context)
Page statusunknown
Metricaccuracy (plus group difference_score)
Directionhigher_is_better
Unit%
Dataset size16076
Dataset licenceApache-2.0
PublisherNew York University, Machine Learning for Language group (via the BIG-bench collaboration)

What it measures

BBQ-Lite presents a short US-English context, a question about two people or groups, and three answers (group A, group B, or unknown). Ambiguous contexts have no resolving fact, so the correct choice is unknown. Disambiguated contexts add a sentence that names the answer. Nine protected categories are covered. The authors describe this BIG-bench task as a subset of the then-upcoming full BBQ dataset.

Task format

Three-way multiple choice, scored from conditional log-probabilities of the three answer strings (programmatic BIG-bench task). Dummy-model header: 16,076 queries. Canary GUID embedded.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub