SuperGLUE AX-b (Broad Coverage Diagnostics)

SuperGLUE diagnostic: 1,104 English sentence pairs recast from the GLUE diagnostic as two-way entailment, scored with Matthews correlation.

Also known as: AX-b, AXb, SuperGLUE_AX_b, AX_b

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryreasoning
Subcategorytwo-way English diagnostic NLI recast from the GLUE broad-coverage diagnostic
Page statusunknown
MetricMatthews correlation (MCC); OpenCompass reports accuracy instead
Directionhigher_is_better
Unit%
Dataset size1104
Dataset licenceother
PublisherNew York University (SuperGLUE packaging); diagnostic items from the GLUE authors

What it measures

AX-b asks whether sentence2 is entailed by sentence1, as two-class English textual entailment (entailment versus not_entailment). SuperGLUE keeps the GLUE expert diagnostic set but collapses contradiction and neutral into not_entailment, because MultiNLI is not a SuperGLUE task. Submissions are asked to run the RTE model on this set. Items are tagged with logic phenomena (negation, monotone, conjunction, and others) for analysis. The set is a diagnostic, not one of the eight SuperGLUE score tasks. English text, sentence pairs.

Task format

Binary sentence-pair classification. Official SuperGLUE scoring is Matthews correlation (MCC) on the 1,104 labelled pairs, scaled by 100 in Table 3. OpenCompass generation asks A/B ("Is the sentence below entailed by the sentence above?") and scores accuracy; perplexity configs compare Yes/No or entailment/not_entailment continuations.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub