MRAG-Bench evaluates whether vision-language models can retrieve and use visual knowledge for multimodal question answering.
unassessed
| Category | multimodal |
|---|---|
| Subcategory | retrieval-augmented generation |
| Metric | multiple-choice accuracy |
| Direction | higher_is_better |
| Unit | percent |
| Dataset size | 1353 |
| Publisher | MRAG-Bench authors |
MRAG-Bench tests retrieval-augmented multimodal models on scenarios where images provide more useful evidence than text. It contains 16,130 images and 1,353 human-annotated multiple-choice questions across nine scenarios.
Image retrieval plus multiple-choice visual question answering.
No model card in ModelSpec reports this benchmark yet.