ZeroBench

100 hand-made visual reasoning questions filtered so no evaluated frontier model answered any correctly at release, plus 334 subquestions to track partial progress.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorymultimodal
Subcategoryvisual reasoning, adversarially filtered to be unsolved at release
Page statusactive
Metricaccuracy (pass@1 / pass@5 / pass^5)
Directionhigher_is_better
Unit%
Dataset size100
PublisherUniversity of Cambridge

What it measures

ZeroBench pairs one or more images with a question requiring multi-step visual reasoning -- careful counting, spatial relations, fine detail extraction, cross-referencing several parts of an image -- built by a pool of more than 20 human question creators and then adversarially filtered: any candidate question that a baseline model answered correctly was discarded. The surviving 100 "main" questions are ones no model in the authors' 20-model baseline set solved at release. Because a set of yes/no answers this hard would still be impossible to fully differentiate, each main question also has, on average, 3.3 companion "subquestions" (334 in total) covering the intermediate reasoning steps, which let a model that fails the main question still show partial progress. Images are mostly natural photographs (70 of 100) rather than synthetic compositions (30 of 100), and mostly single-image questions (93 of 100, with 7 requiring more than one image).

Task format

Open-ended, free-text answers, required in the format "{final_answer}". Grading is exact string match against the reference answer, with no partial credit on the main questions, since the authors found no distance metric could fairly cover the diversity of correct-answer formats; questions with binary, multiple-choice or small-integer (under 10) answers were deliberately excluded during curation to keep guessing from working.

Models reporting this benchmark

These figures come from the model cards, which carry one collection date per card and no per-score attribution. They are shown as reported, not as verified evidence.
ModelProviderScoreCard as of
Muse SparkMeta33.02026-04

Data

This page as JSON · Edit on GitHub