A grid-puzzle benchmark of novel abstract-reasoning tasks, each solvable by most humans but built to resist memorisation and brute-force search by AI systems.
unassessed
| Category | reasoning |
|---|---|
| Subcategory | abstract visual reasoning |
| Page status | active |
| Metric | % of tasks solved (exact grid match within 2 attempts) |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 1000 |
| Dataset licence | Apache-2.0 |
| Publisher | ARC Prize Foundation |
ARC-AGI-2 shows a model a handful of input-output grid examples that share a hidden transformation rule, then asks it to apply that rule to a new input grid. Grids are small matrices of coloured cells, and every task is novel and hand-designed so that no amount of exposure to similar puzzles substitutes for actually inferring the rule. Version 2 specifically adds tasks that require symbolic interpretation, applying several interacting rules at once, and adapting a rule to context, to separate genuine generalisation from the pattern-matching and search strategies that had started to do well on the original ARC-AGI.
A small number of paired example grids (input and output) plus one or more test input grids; the model must output the exact matching grid, with up to two attempts per test input.
| Model | Provider | Score | Card as of |
|---|---|---|---|
| Muse Spark | Meta | 42.5 | 2026-04 |