ARC-AGI-2

A grid-puzzle benchmark of novel abstract-reasoning tasks, each solvable by most humans but built to resist memorisation and brute-force search by AI systems.

Also known as: ARC-AGI 2, Abstraction and Reasoning Corpus for AGI, version 2

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryreasoning
Subcategoryabstract visual reasoning
Page statusactive
Metric% of tasks solved (exact grid match within 2 attempts)
Directionhigher_is_better
Unit%
Dataset size1000
Dataset licenceApache-2.0
PublisherARC Prize Foundation

What it measures

ARC-AGI-2 shows a model a handful of input-output grid examples that share a hidden transformation rule, then asks it to apply that rule to a new input grid. Grids are small matrices of coloured cells, and every task is novel and hand-designed so that no amount of exposure to similar puzzles substitutes for actually inferring the rule. Version 2 specifically adds tasks that require symbolic interpretation, applying several interacting rules at once, and adapting a rule to context, to separate genuine generalisation from the pattern-matching and search strategies that had started to do well on the original ARC-AGI.

Task format

A small number of paired example grids (input and output) plus one or more test input grids; the model must output the exact matching grid, with up to two attempts per test input.

Models reporting this benchmark

These figures come from the model cards, which carry one collection date per card and no per-score attribution. They are shown as reported, not as verified evidence.
ModelProviderScoreCard as of
Muse SparkMeta42.52026-04

Data

This page as JSON · Edit on GitHub