SuperGLUE Winogender diagnostic: 356 English premise–hypothesis pairs that test whether pronoun gender flips an entailment decision.
unassessed
| Category | safety |
|---|---|
| Subcategory | Winogender-as-NLI gender-parity diagnostic (SuperGLUE) |
| Page status | unknown |
| Metric | accuracy and gender parity score (GPS); OpenCompass reports accuracy only |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 356 |
| Dataset licence | other |
| Publisher | New York University (SuperGLUE packaging); Winogender from Rudinger et al.; DNC recast from Poliak et al. |
AX-g recasts Winogender (Rudinger et al., 2018) as two-way textual entailment using the Diverse Natural Language Inference Collection (Poliak et al., 2018). Each item is an English premise with a male or female pronoun and a hypothesis that names one possible antecedent (occupation or participant). Minimal pairs differ only in pronoun gender. The model must say entailment or not_entailment. SuperGLUE reports accuracy and a gender parity score: the share of pairs with the same prediction after the gender swap. High parity with chance accuracy is the trivial solution of always guessing one class. The set is a diagnostic, not one of the eight SuperGLUE score tasks.
Binary premise–hypothesis classification. Official metrics are accuracy and GPS, scaled by 100 in Table 3. OpenCompass generation asks A/B on premise then hypothesis; perplexity configs compare Yes/No continuations. OpenCompass scores accuracy only.
No model card in ModelSpec reports this benchmark yet.