ARC-Easy

The easier, larger half of the AI2 Reasoning Challenge; questions that 2018-era retrieval or word-overlap baselines could already answer, now scored near ceiling by current models.

Also known as: ARC (Easy Set), AI2 Reasoning Challenge (Easy)

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryreasoning
Subcategorygrade-school science multiple-choice QA (Easy split)
Page statussuperseded
Metricaccuracy (often reported as acc_norm, length-normalised)
Directionhigher_is_better
Unit%
Dataset size5197
Dataset licenceCC-BY-SA-4.0
PublisherAllen Institute for AI (AI2)

What it measures

ARC-Easy is the larger of the two AI2 Reasoning Challenge splits: 5,197 grade-school-level natural science exam questions that landed here specifically because at least one of two 2018-era baseline solvers (an information-retrieval solver or a word-co-occurrence/PMI solver) answered them correctly. Every question that both baselines failed went to the harder ARC-Challenge split instead. Format, source and authorship are identical to Challenge; only the difficulty filter differs.

Task format

Multiple-choice science question, typically 4 answer options, single correct answer.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub