MMLU (Massive Multitask Language Understanding)

A 57-subject, four-choice knowledge test from elementary to professional difficulty; the standard reference for broad model knowledge since 2020, now saturated at the frontier.

Also known as: Massive Multitask Language Understanding, Hendrycks Test

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryknowledge
Subcategorymultitask academic and professional knowledge
Page statusactive
Metricaccuracy
Directionhigher_is_better
Unit%
Dataset size14042
Dataset licenceMIT
PublisherUC Berkeley (original); Center for AI Safety (current host)

What it measures

Each question gives the model a short prompt and four labelled answer options drawn from one of 57 academic and professional subjects, from abstract algebra to professional law. The model must pick the single correct option. This mostly exercises declarative and procedural knowledge acquired during pretraining rather than multi-step reasoning, and the original paper found lopsided results across subjects rather than uniform competence.

Task format

Four-option multiple-choice question answering, graded on the single labelled option (A-D) the model selects. Evaluated zero-shot or few-shot; the original paper's dev split provides up to 5 worked examples per subject for the few-shot case, and 5-shot has become the common reporting condition. Scored either by comparing the log-likelihood the model assigns to each answer letter as a continuation, or by having the model generate free text and parsing out a letter.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub