UnQover probes gender, nationality, ethnicity and religion stereotypes with underspecified span-question-answering templates.
unassessed
| Category | safety |
|---|---|
| Subcategory | stereotyping bias in underspecified QA |
| Page status | active |
| Metric | fairness (1 minus bias intensity); consistency alongside it |
| Direction | higher_is_better |
| Unit | score |
| Dataset size | 10552928 |
| Publisher | BIG-bench collaboration |
A paragraph names two candidates and asks an underspecified question. The model scores both spans even though context does not justify choosing either. Perturbations isolate stereotype bias from positional and question-attribute errors.
Generated span-based QA templates covering gender-occupation, nationality, ethnicity and religion, with candidate-order and polarity perturbations.
No model card in ModelSpec reports this benchmark yet.