RealWorldQA

A multiple-choice test of everyday spatial understanding built from over 700 real-world photos, each paired with one verifiable question.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorymultimodal
Subcategoryreal-world spatial understanding
Page statusactive
Metricaccuracy
Directionhigher_is_better
Unit%
Dataset size765
Dataset licenceCC BY-ND 4.0
PublisherxAI

What it measures

RealWorldQA shows a model a single photo taken from a vehicle or another everyday real-world setting and asks a multiple-choice question about what is in the scene, for example distances, positions, counts, or the relation between objects. The questions are meant to be easy for a person but still expose weaknesses in a model's basic spatial and physical-world understanding, rather than testing specialist knowledge or long chains of reasoning.

Task format

multiple-choice visual question answering over a single photo, 2-4 answer options per question, most with 3

Models reporting this benchmark

These figures come from the model cards, which carry one collection date per card and no per-score attribution. They are shown as reported, not as verified evidence.
ModelProviderScoreCard as of
GPT-4oOpenAI75.42026-04
GPT-4o (2024-05-13)OpenAI75.42026-04
GPT-4o (2024-08-06)OpenAI75.42026-04
GPT-4o (2024-11-20)OpenAI75.42026-04
NVLM D 72BNVIDIA69.52026-04
MiniCPM V 4OpenBMB68.52026-04
MiniCPM V 4 5OpenBMB68.52026-04
MiniCPM V 4 5 ggufOpenBMB68.52026-04

Data

This page as JSON · Edit on GitHub