V-FAT tests whether multimodal models answer from the image when corpus priors or misleading prompts conflict with what is shown.
unassessed
| Category | multimodal |
|---|---|
| Subcategory | visual grounding under text bias |
| Page status | active |
| Metric | Visual Robustness Score (VRS) |
| Direction | higher_is_better |
| Unit | score (0-1) |
| Dataset size | 4026 |
| Publisher | The Chinese University of Hong Kong, Shenzhen, with Meituan and the University of Oxford |
V-FAT is a diagnostic visual question-answering test for multimodal large language models. Each item pairs an image with a question that can be answered from pixels, then adds increasing textual pressure: atypical scenes that fight common language priors, misleading instructions that assert a false visual fact, or both at once. The skill is visual fidelity rather than ordinary VQA accuracy, because a linguistically plausible guess can still be wrong relative to the image.
Image plus a multiple-choice or open-ended question at one of three bias levels; the model must report the visual fact, not the text prior or the prompt's false premise.
No model card in ModelSpec reports this benchmark yet.