BIG-bench BBQ subset: 16,076 three-choice US-English QA items that test social stereotypes in ambiguous and disambiguated contexts.
unassessed
| Category | safety |
|---|---|
| Subcategory | social-bias three-choice QA (ambiguous vs disambiguated context) |
| Page status | unknown |
| Metric | accuracy (plus group difference_score) |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 16076 |
| Dataset licence | Apache-2.0 |
| Publisher | New York University, Machine Learning for Language group (via the BIG-bench collaboration) |
BBQ-Lite presents a short US-English context, a question about two people or groups, and three answers (group A, group B, or unknown). Ambiguous contexts have no resolving fact, so the correct choice is unknown. Disambiguated contexts add a sentence that names the answer. Nine protected categories are covered. The authors describe this BIG-bench task as a subset of the then-upcoming full BBQ dataset.
Three-way multiple choice, scored from conditional log-probabilities of the three answer strings (programmatic BIG-bench task). Dummy-model header: 16,076 queries. Canary GUID embedded.
No model card in ModelSpec reports this benchmark yet.