MRAG-Bench

MRAG-Bench evaluates whether vision-language models can retrieve and use visual knowledge for multimodal question answering.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorymultimodal
Subcategoryretrieval-augmented generation
Metricmultiple-choice accuracy
Directionhigher_is_better
Unitpercent
Dataset size1353
PublisherMRAG-Bench authors

What it measures

MRAG-Bench tests retrieval-augmented multimodal models on scenarios where images provide more useful evidence than text. It contains 16,130 images and 1,353 human-annotated multiple-choice questions across nine scenarios.

Task format

Image retrieval plus multiple-choice visual question answering.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub