A multiple-choice grade-school science question set built so that simple retrieval and word-overlap methods fail, isolating genuine multi-hop reasoning.
unassessed
| Category | reasoning |
|---|---|
| Subcategory | science multiple-choice QA |
| Page status | superseded |
| Metric | accuracy (often reported as acc_norm, length-normalised) |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 2590 |
| Dataset licence | CC BY-SA 4.0 |
| Publisher | Allen Institute for AI (AI2) |
ARC-Challenge gives a model a grade-school-level natural science question with several answer choices, drawn from real school exam materials, and asks it to pick the correct one. Every question in the set was specifically selected because neither a retrieval-based algorithm nor a simple word-co-occurrence algorithm could answer it correctly, so it requires more than surface-level lexical matching between the question and background text. It tests a mix of scientific fact recall and the multi-step reasoning needed to combine that fact with the question.
Multiple-choice science question, typically 4 answer options, single correct answer.
| Model | Provider | Score | Card as of |
|---|---|---|---|
| Mixtral 8x22B Instruct v0.1 | Mistral AI | 72.7 | 2026-04 |
| Meta Llama 3 70B | Meta | 71.4 | 2024-07 |
| Meta Llama 3 70B Instruct | Meta | 71.4 | 2026-04 |
| Meta Llama 3 70B Instruct | Nous Research | 71.4 | 2026-04 |
| Nous Hermes 2 Mixtral 8x7B DPO | Nous Research | 71.1 | 2024-07 |
| Yi 1.5 34B Chat | 01.AI | 70.5 | 2026-04 |
| Yi 1.5 34B Chat 16K | 01.AI | 70.5 | 2024-07 |
| Mixtral 8x7B Instruct v0.1 | Mistral AI | 70.1 | 2026-04 |
| Nous Hermes 2 Yi 34B | Nous Research | 66.9 | 2024-07 |
| Nous Hermes 2 SOLAR 10.7B | Nous Research | 66.7 | 2024-07 |
| Mixtral 8x7B v0.1 | Mistral AI | 66.4 | 2026-04 |
| Yi 1.5 34B | 01.AI | 65.8 | 2026-04 |
| Yi 34B 200K | 01.AI | 65.8 | 2024-07 |
| Yi 34B Chat | 01.AI | 65.4 | 2024-07 |
| Yi 1.5 34B 32K | 01.AI | 64.4 | 2024-07 |
| Yi 1.5 9B Chat 16K | 01.AI | 64.3 | 2024-07 |
| Yi 1.5 9B Chat | 01.AI | 63.7 | 2024-07 |
| Hermes 2 Pro Llama 3 8B | Nous Research | 63.5 | 2024-07 |
| Hermes 2 Theta Llama 3 8B | Nous Research | 63.1 | 2024-07 |
| Mistral 7B Instruct v0.2 | Mistral AI | 63.1 | 2024-07 |
| Phi 3 mini 128K instruct | Microsoft | 63.1 | 2024-07 |
| Phi 3 mini 4K instruct | Microsoft | 63.0 | 2024-07 |
| falcon 40B | TII | 61.9 | 2024-07 |
| Yi 1.5 9B | 01.AI | 61.9 | 2024-07 |
| Yi 9B | 01.AI | 61.2 | 2024-07 |
| gemma 7B it | Google DeepMind | 61.1 | 2024-07 |
| Yi 1.5 9B 32K | 01.AI | 61.1 | 2024-07 |
| phi 2 | Microsoft | 61.0 | 2024-07 |
| Meta Llama 3 8B Instruct | Meta | 60.8 | 2024-07 |
| Meta Llama 3 8B Instruct | Nous Research | 60.8 | 2024-07 |
| Yi 1.5 6B Chat | 01.AI | 60.7 | 2024-07 |
| Mistral 7B v0.3 | Mistral AI | 60.5 | 2024-07 |
| mistral 7B v0.3 bnb 4bit | Unsloth | 60.5 | 2024-07 |
| Meta Llama 3 8B | Meta | 60.2 | 2024-07 |
| Meta Llama 3 8B | Nous Research | 60.2 | 2024-07 |
| Yi 1.5 6B | 01.AI | 57.3 | 2024-07 |
| deepseek llm 7B base | DeepSeek | 55.7 | 2024-07 |
| deepseek llm 7B chat | DeepSeek | 55.7 | 2024-07 |
| Yi 6B | 01.AI | 55.5 | 2024-07 |
| Yi 6B Chat | 01.AI | 55.5 | 2024-07 |
| gemma 2B | Google DeepMind | 48.4 | 2024-07 |
| Qwen2 1.5B Instruct | Alibaba / Qwen Team | 44.3 | 2024-07 |
| gemma 2B it | Google DeepMind | 43.9 | 2024-07 |
| chatglm2 6B | Zhipu AI | 38.8 | 2024-07 |
| deepseek coder 6.7B instruct | DeepSeek | 38.1 | 2024-07 |
| deepseek coder 6.7B base | DeepSeek | 37.0 | 2024-07 |
| OLMo 1B hf | Allen AI | 34.6 | 2024-07 |
| Qwen2 0.5B Instruct | Alibaba / Qwen Team | 31.9 | 2024-07 |
| deepseek coder 1.3B base | DeepSeek | 28.6 | 2024-07 |
| deepseek coder 1.3B instruct | DeepSeek | 28.6 | 2024-07 |