A 191-question visual-question-answering test built from high-resolution, visually crowded images, designed so a model cannot answer correctly without precisely locating a small target detail.
unassessed
| Category | multimodal |
|---|---|
| Subcategory | high-resolution visual search and grounding |
| Page status | active |
| Metric | accuracy |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 191 |
| Publisher | New York University; UC San Diego |
V*Bench shows a model a high-resolution photograph (drawn from the SA-1B dataset, averaging 2246x1582 pixels) that is visually crowded with detail, and asks a multiple-choice question that can only be answered correctly by finding and closely inspecting one small, specific region -- an object's colour or material, or the relative position between two objects -- rather than reading the image at a glance. It targets a gap the authors identified in multimodal LLMs of the time: models that performed well on benchmarks built from smaller, sparser images failed once the informative detail was small relative to the whole frame, because they lacked any mechanism for deliberately searching the image the way a person visually scans a crowded scene for something specific.
Multiple-choice visual question answering: one high-resolution image plus a question in, one selected option out (four options for attribute questions, two for spatial-relationship questions).
No model card in ModelSpec reports this benchmark yet.