MULTI

MULTI evaluates Chinese multimodal understanding with more than 18,000 authentic examination questions and hard and in-context variants.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorymultimodal
Metricaccuracy
Directionhigher_is_better
Unitpercent
Dataset size18000
PublisherMULTI authors

What it measures

MULTI tests image-text comprehension, complex reasoning, and knowledge recall against real examination standards. MULTI-Elite is a 500-question hard subset, while MULTI-Extend adds more than 4,500 external knowledge context pieces.

Task format

Chinese image-text multiple-choice questions with optional retrieved context.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub