A 57-subject, four-choice knowledge test from elementary to professional difficulty; the standard reference for broad model knowledge since 2020, now saturated at the frontier.
unassessed
| Category | knowledge |
|---|---|
| Subcategory | multitask academic and professional knowledge |
| Page status | active |
| Metric | accuracy |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 14042 |
| Dataset licence | MIT |
| Publisher | UC Berkeley (original); Center for AI Safety (current host) |
Each question gives the model a short prompt and four labelled answer options drawn from one of 57 academic and professional subjects, from abstract algebra to professional law. The model must pick the single correct option. This mostly exercises declarative and procedural knowledge acquired during pretraining rather than multi-step reasoning, and the original paper found lopsided results across subjects rather than uniform competence.
Four-option multiple-choice question answering, graded on the single labelled option (A-D) the model selects. Evaluated zero-shot or few-shot; the original paper's dev split provides up to 5 worked examples per subject for the few-shot case, and 5-shot has become the common reporting condition. Scored either by comparing the log-likelihood the model assigns to each answer letter as a continuation, or by having the model generate free text and parsing out a letter.
No model card in ModelSpec reports this benchmark yet.