SuperGLUE AX-g (Winogender Schema Diagnostics)

SuperGLUE Winogender diagnostic: 356 English premise–hypothesis pairs that test whether pronoun gender flips an entailment decision.

Also known as: AX-g, AXg, SuperGLUE_AX_g, AX_g, Winogender (SuperGLUE)

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorysafety
SubcategoryWinogender-as-NLI gender-parity diagnostic (SuperGLUE)
Page statusunknown
Metricaccuracy and gender parity score (GPS); OpenCompass reports accuracy only
Directionhigher_is_better
Unit%
Dataset size356
Dataset licenceother
PublisherNew York University (SuperGLUE packaging); Winogender from Rudinger et al.; DNC recast from Poliak et al.

What it measures

AX-g recasts Winogender (Rudinger et al., 2018) as two-way textual entailment using the Diverse Natural Language Inference Collection (Poliak et al., 2018). Each item is an English premise with a male or female pronoun and a hypothesis that names one possible antecedent (occupation or participant). Minimal pairs differ only in pronoun gender. The model must say entailment or not_entailment. SuperGLUE reports accuracy and a gender parity score: the share of pairs with the same prediction after the gender swap. High parity with chance accuracy is the trivial solution of always guessing one class. The set is a diagnostic, not one of the eight SuperGLUE score tasks.

Task format

Binary premise–hypothesis classification. Official metrics are accuracy and GPS, scaled by 100 in Table 3. OpenCompass generation asks A/B on premise then hypothesis; perplexity configs compare Yes/No continuations. OpenCompass scores accuracy only.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub