V-FAT

V-FAT tests whether multimodal models answer from the image when corpus priors or misleading prompts conflict with what is shown.

Also known as: Visual Fidelity Against Text-bias

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorymultimodal
Subcategoryvisual grounding under text bias
Page statusactive
MetricVisual Robustness Score (VRS)
Directionhigher_is_better
Unitscore (0-1)
Dataset size4026
PublisherThe Chinese University of Hong Kong, Shenzhen, with Meituan and the University of Oxford

What it measures

V-FAT is a diagnostic visual question-answering test for multimodal large language models. Each item pairs an image with a question that can be answered from pixels, then adds increasing textual pressure: atypical scenes that fight common language priors, misleading instructions that assert a false visual fact, or both at once. The skill is visual fidelity rather than ordinary VQA accuracy, because a linguistically plausible guess can still be wrong relative to the image.

Task format

Image plus a multiple-choice or open-ended question at one of three bias levels; the model must report the visual fact, not the text prior or the prompt's false premise.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub