SuperGLUE diagnostic: 1,104 English sentence pairs recast from the GLUE diagnostic as two-way entailment, scored with Matthews correlation.
unassessed
| Category | reasoning |
|---|---|
| Subcategory | two-way English diagnostic NLI recast from the GLUE broad-coverage diagnostic |
| Page status | unknown |
| Metric | Matthews correlation (MCC); OpenCompass reports accuracy instead |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 1104 |
| Dataset licence | other |
| Publisher | New York University (SuperGLUE packaging); diagnostic items from the GLUE authors |
AX-b asks whether sentence2 is entailed by sentence1, as two-class English textual entailment (entailment versus not_entailment). SuperGLUE keeps the GLUE expert diagnostic set but collapses contradiction and neutral into not_entailment, because MultiNLI is not a SuperGLUE task. Submissions are asked to run the RTE model on this set. Items are tagged with logic phenomena (negation, monotone, conjunction, and others) for analysis. The set is a diagnostic, not one of the eight SuperGLUE score tasks. English text, sentence pairs.
Binary sentence-pair classification. Official SuperGLUE scoring is Matthews correlation (MCC) on the 1,104 labelled pairs, scaled by 100 in Table 3. OpenCompass generation asks A/B ("Is the sentence below entailed by the sentence above?") and scores accuracy; perplexity configs compare Yes/No or entailment/not_entailment continuations.
No model card in ModelSpec reports this benchmark yet.