A programmatic BIG-bench task that scores gender-occupation local and global bias on naturally occurring English contexts, then negates the scores so higher is fairer.
unassessed
| Category | safety |
|---|---|
| Subcategory | BIG-bench programmatic gender-occupation bias on diverse English contexts |
| Page status | unknown |
| Metric | overall gender bias (negative mean of four local/global gaps) |
| Direction | higher_is_better |
| Dataset size | 169 |
| Dataset licence | Apache-2.0 |
| Publisher | Google (BIG-bench collaboration) |
diverse_social_bias prompts a model with naturally occurring English contexts and measures whether next-token and continuation probabilities shift with gender. Authors Paul Pu Liang and Chiyu Wu ship 114 occupation-context lines, 55 gender-context lines, 14 binary gender pairs, and 274 occupation tokens. Local occupation-gender bias is mean absolute probability difference between paired gender tokens. Global bias is length-normalised perplexity difference on gender-swapped continuations. The gender-occupation pair uses Hellinger distance over occupation tokens and a swapped-context perplexity gap. The skill is representational gender-occupation bias in diverse contexts, not [bbq](bbq.md) QA and not [gender_sensitivity_english](gender_sensitivity_english.md).
Programmatic cond_log_prob task. Preferred score "overall gender bias" (negative of the four-metric mean). Zero-shot. Canary GUID embedded. Dummy-model header: 448 multiple-choice queries.
No model card in ModelSpec reports this benchmark yet.