100 hand-made visual reasoning questions filtered so no evaluated frontier model answered any correctly at release, plus 334 subquestions to track partial progress.
unassessed
| Category | multimodal |
|---|---|
| Subcategory | visual reasoning, adversarially filtered to be unsolved at release |
| Page status | active |
| Metric | accuracy (pass@1 / pass@5 / pass^5) |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 100 |
| Publisher | University of Cambridge |
ZeroBench pairs one or more images with a question requiring multi-step visual reasoning -- careful counting, spatial relations, fine detail extraction, cross-referencing several parts of an image -- built by a pool of more than 20 human question creators and then adversarially filtered: any candidate question that a baseline model answered correctly was discarded. The surviving 100 "main" questions are ones no model in the authors' 20-model baseline set solved at release. Because a set of yes/no answers this hard would still be impossible to fully differentiate, each main question also has, on average, 3.3 companion "subquestions" (334 in total) covering the intermediate reasoning steps, which let a model that fails the main question still show partial progress. Images are mostly natural photographs (70 of 100) rather than synthetic compositions (30 of 100), and mostly single-image questions (93 of 100, with 7 requiring more than one image).
Open-ended, free-text answers, required in the format "{final_answer}". Grading is exact string match against the reference answer, with no partial credit on the main questions, since the authors found no distance metric could fairly cover the diversity of correct-answer formats; questions with binary, multiple-choice or small-integer (under 10) answers were deliberately excluded during curation to keep guessing from working.
| Model | Provider | Score | Card as of |
|---|---|---|---|
| Muse Spark | Meta | 33.0 | 2026-04 |