AGIEval

8,062 real questions from 20 human standardized exams -- Gaokao, SAT, LSAT, a Chinese bar exam, civil-service logic tests and math competitions -- scored mostly by multiple-choice accuracy.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorycomposite
Subcategoryhuman standardized-exam questions: college entrance, law, civil service and math competitions (bilingual English/Chinese)
Page statusactive
Metricaccuracy (MCQ tasks); Exact Match and F1 (cloze tasks)
Directionhigher_is_better
Unit%
Dataset size8062
Dataset licenceCode: MIT (Microsoft). The question data itself is not covered by a single licence; the publisher states usage of the data should follow the licence of each original source exam/dataset, which differ across the 20 tasks.
PublisherMicrosoft

What it measures

AGIEval tests a model with real, previously administered questions from human standardized exams rather than purpose-written benchmark items: the Chinese college entrance exam (Gaokao) across eight subjects, the American SAT (English and Math), the Law School Admission Test (LSAT, covering analytical reasoning, logical reasoning and reading comprehension), a Chinese lawyer qualification exam (JEC-QA), a civil-service-style logical-reasoning test (LogiQA, in both English and Chinese), and two existing math datasets folded in to represent GRE-style word problems (AQuA-RAT) and competition mathematics (MATH). The authors frame this as probing four capability dimensions -- understanding, knowledge, reasoning and calculation -- and report that GPT-4 already exceeded average human test-taker performance on several individual exams (SAT Math, LSAT, math competitions) at release, while lagging on tasks needing deeper domain reasoning.

Task format

Mostly four- or five-option multiple-choice questions (18 of 20 tasks); two are fill-in-the-blank cloze tasks (Gaokao-Math-Cloze and MATH). Several tasks include a reading passage (LSAT, SAT, Gaokao Chinese/English, LogiQA). The original paper evaluates under zero-shot, few-shot and chain-of-thought prompting; harnesses commonly report zero-shot or few-shot accuracy.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub