SuperGPQA

Graduate-level multiple-choice questions across 13 disciplines, 72 fields and 285 subfields, filtered with a human-LLM pipeline to drop trivial and ambiguous items.

Also known as: Super-GPQA, Super GPQA

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryknowledge
Subcategorygraduate-level multidisciplinary multiple-choice QA
Page statusactive
Metricaccuracy
Directionhigher_is_better
Unit%
Dataset size26529
Dataset licenceODC-By
PublisherM-A-P Team, with ByteDance Seed and 2077.AI

What it measures

SuperGPQA asks a model to answer a graduate-level question in English by choosing among labelled options. Items span far beyond the three sciences in GPQA: the authors' taxonomy has 13 disciplines, 72 fields and 285 subfields, with Science, Engineering and Medicine holding most of the mass. Annotators rewrite source material into multiple-choice form, add distractors, and drop items that experts or models mark as trivial or ambiguous. A high score is meant to show graduate knowledge and reasoning in long-tail fields, not only in math, physics and CS.

Task format

Single-turn English multiple-choice generation. OpenCompass formats the stem plus lettered options (A, B, C, ...) and scores the extracted answer letter against `answer_letter`. Default config is zero-shot; a five-shot template also ships.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub