ARC-Challenge

A multiple-choice grade-school science question set built so that simple retrieval and word-overlap methods fail, isolating genuine multi-hop reasoning.

Also known as: AI2 Reasoning Challenge, ARC (Challenge Set)

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryreasoning
Subcategoryscience multiple-choice QA
Page statussuperseded
Metricaccuracy (often reported as acc_norm, length-normalised)
Directionhigher_is_better
Unit%
Dataset size2590
Dataset licenceCC BY-SA 4.0
PublisherAllen Institute for AI (AI2)

What it measures

ARC-Challenge gives a model a grade-school-level natural science question with several answer choices, drawn from real school exam materials, and asks it to pick the correct one. Every question in the set was specifically selected because neither a retrieval-based algorithm nor a simple word-co-occurrence algorithm could answer it correctly, so it requires more than surface-level lexical matching between the question and background text. It tests a mix of scientific fact recall and the multi-step reasoning needed to combine that fact with the question.

Task format

Multiple-choice science question, typically 4 answer options, single correct answer.

Models reporting this benchmark

These figures come from the model cards, which carry one collection date per card and no per-score attribution. They are shown as reported, not as verified evidence.
ModelProviderScoreCard as of
Mixtral 8x22B Instruct v0.1Mistral AI72.72026-04
Meta Llama 3 70BMeta71.42024-07
Meta Llama 3 70B InstructMeta71.42026-04
Meta Llama 3 70B InstructNous Research71.42026-04
Nous Hermes 2 Mixtral 8x7B DPONous Research71.12024-07
Yi 1.5 34B Chat01.AI70.52026-04
Yi 1.5 34B Chat 16K01.AI70.52024-07
Mixtral 8x7B Instruct v0.1Mistral AI70.12026-04
Nous Hermes 2 Yi 34BNous Research66.92024-07
Nous Hermes 2 SOLAR 10.7BNous Research66.72024-07
Mixtral 8x7B v0.1Mistral AI66.42026-04
Yi 1.5 34B01.AI65.82026-04
Yi 34B 200K01.AI65.82024-07
Yi 34B Chat01.AI65.42024-07
Yi 1.5 34B 32K01.AI64.42024-07
Yi 1.5 9B Chat 16K01.AI64.32024-07
Yi 1.5 9B Chat01.AI63.72024-07
Hermes 2 Pro Llama 3 8BNous Research63.52024-07
Hermes 2 Theta Llama 3 8BNous Research63.12024-07
Mistral 7B Instruct v0.2Mistral AI63.12024-07
Phi 3 mini 128K instructMicrosoft63.12024-07
Phi 3 mini 4K instructMicrosoft63.02024-07
falcon 40BTII61.92024-07
Yi 1.5 9B01.AI61.92024-07
Yi 9B01.AI61.22024-07
gemma 7B itGoogle DeepMind61.12024-07
Yi 1.5 9B 32K01.AI61.12024-07
phi 2Microsoft61.02024-07
Meta Llama 3 8B InstructMeta60.82024-07
Meta Llama 3 8B InstructNous Research60.82024-07
Yi 1.5 6B Chat01.AI60.72024-07
Mistral 7B v0.3Mistral AI60.52024-07
mistral 7B v0.3 bnb 4bitUnsloth60.52024-07
Meta Llama 3 8BMeta60.22024-07
Meta Llama 3 8BNous Research60.22024-07
Yi 1.5 6B01.AI57.32024-07
deepseek llm 7B baseDeepSeek55.72024-07
deepseek llm 7B chatDeepSeek55.72024-07
Yi 6B01.AI55.52024-07
Yi 6B Chat01.AI55.52024-07
gemma 2BGoogle DeepMind48.42024-07
Qwen2 1.5B InstructAlibaba / Qwen Team44.32024-07
gemma 2B itGoogle DeepMind43.92024-07
chatglm2 6BZhipu AI38.82024-07
deepseek coder 6.7B instructDeepSeek38.12024-07
deepseek coder 6.7B baseDeepSeek37.02024-07
OLMo 1B hfAllen AI34.62024-07
Qwen2 0.5B InstructAlibaba / Qwen Team31.92024-07
deepseek coder 1.3B baseDeepSeek28.62024-07
deepseek coder 1.3B instructDeepSeek28.62024-07

Data

This page as JSON · Edit on GitHub