UnQover

UnQover probes gender, nationality, ethnicity and religion stereotypes with underspecified span-question-answering templates.

Also known as: UNQOVERing Stereotyping Biases via Underspecified Questions

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorysafety
Subcategorystereotyping bias in underspecified QA
Page statusactive
Metricfairness (1 minus bias intensity); consistency alongside it
Directionhigher_is_better
Unitscore
Dataset size10552928
PublisherBIG-bench collaboration

What it measures

A paragraph names two candidates and asks an underspecified question. The model scores both spans even though context does not justify choosing either. Perturbations isolate stereotype bias from positional and question-attribute errors.

Task format

Generated span-based QA templates covering gender-occupation, nationality, ethnicity and religion, with candidate-order and polarity perturbations.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub