ARC (AI2 Reasoning Challenge)

A grade-school science multiple-choice question set split into Easy and Challenge halves, whose scores are reported interchangeably far too often despite very different difficulty.

Also known as: AI2 Reasoning Challenge

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryreasoning
Subcategorygrade-school science multiple-choice QA (Easy and Challenge splits)
Page statussuperseded
Metricaccuracy (often reported as acc_norm, length-normalised)
Directionhigher_is_better
Unit%
Dataset size7787
Dataset licenceCC-BY-SA-4.0
PublisherAllen Institute for AI (AI2)

What it measures

The AI2 Reasoning Challenge (ARC) is a set of 7,787 natural, grade-school-level science exam questions, split by the original authors into two pools of different difficulty: ARC-Easy (5,197 questions) and ARC-Challenge (2,590 questions). A question lands in the Challenge set only if two baseline solvers of the time -- an information-retrieval solver and a word-co-occurrence (PMI) solver -- both answered it incorrectly; anything either solver could get right went into Easy. That single filtering rule is the entire distinction between the two splits: they share the same format, the same source, and the same authors, and differ only in whether simple lexical-matching methods could solve them. Because both splits are commonly reported under the bare name "ARC" without specifying which one, and a model's score on Easy can be 20-30 points higher than on Challenge, conflating the two is one of the more common reporting errors in this space.

Task format

Multiple-choice science question, typically 4 answer options, single correct answer; identical format for both splits.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub