V*Bench

A 191-question visual-question-answering test built from high-resolution, visually crowded images, designed so a model cannot answer correctly without precisely locating a small target detail.

Also known as: VStar_Bench, vstar-bench, V-Star Bench

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorymultimodal
Subcategoryhigh-resolution visual search and grounding
Page statusactive
Metricaccuracy
Directionhigher_is_better
Unit%
Dataset size191
PublisherNew York University; UC San Diego

What it measures

V*Bench shows a model a high-resolution photograph (drawn from the SA-1B dataset, averaging 2246x1582 pixels) that is visually crowded with detail, and asks a multiple-choice question that can only be answered correctly by finding and closely inspecting one small, specific region -- an object's colour or material, or the relative position between two objects -- rather than reading the image at a glance. It targets a gap the authors identified in multimodal LLMs of the time: models that performed well on benchmarks built from smaller, sparser images failed once the informative detail was small relative to the whole frame, because they lacked any mechanism for deliberately searching the image the way a person visually scans a crowded scene for something specific.

Task format

Multiple-choice visual question answering: one high-resolution image plus a question in, one selected option out (four options for attribute questions, two for spatial-relationship questions).

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub