8,062 real questions from 20 human standardized exams -- Gaokao, SAT, LSAT, a Chinese bar exam, civil-service logic tests and math competitions -- scored mostly by multiple-choice accuracy.
unassessed
| Category | composite |
|---|---|
| Subcategory | human standardized-exam questions: college entrance, law, civil service and math competitions (bilingual English/Chinese) |
| Page status | active |
| Metric | accuracy (MCQ tasks); Exact Match and F1 (cloze tasks) |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 8062 |
| Dataset licence | Code: MIT (Microsoft). The question data itself is not covered by a single licence; the publisher states usage of the data should follow the licence of each original source exam/dataset, which differ across the 20 tasks. |
| Publisher | Microsoft |
AGIEval tests a model with real, previously administered questions from human standardized exams rather than purpose-written benchmark items: the Chinese college entrance exam (Gaokao) across eight subjects, the American SAT (English and Math), the Law School Admission Test (LSAT, covering analytical reasoning, logical reasoning and reading comprehension), a Chinese lawyer qualification exam (JEC-QA), a civil-service-style logical-reasoning test (LogiQA, in both English and Chinese), and two existing math datasets folded in to represent GRE-style word problems (AQuA-RAT) and competition mathematics (MATH). The authors frame this as probing four capability dimensions -- understanding, knowledge, reasoning and calculation -- and report that GPT-4 already exceeded average human test-taker performance on several individual exams (SAT Math, LSAT, math competitions) at release, while lagging on tasks needing deeper domain reasoning.
Mostly four- or five-option multiple-choice questions (18 of 20 tasks); two are fill-in-the-blank cloze tasks (Gaokao-Math-Cloze and MATH). Several tasks include a reading passage (LSAT, SAT, Gaokao Chinese/English, LogiQA). The original paper evaluates under zero-shot, few-shot and chain-of-thought prompting; harnesses commonly report zero-shot or few-shot accuracy.
No model card in ModelSpec reports this benchmark yet.