MULTI evaluates Chinese multimodal understanding with more than 18,000 authentic examination questions and hard and in-context variants.
unassessed
| Category | multimodal |
|---|---|
| Metric | accuracy |
| Direction | higher_is_better |
| Unit | percent |
| Dataset size | 18000 |
| Publisher | MULTI authors |
MULTI tests image-text comprehension, complex reasoning, and knowledge recall against real examination standards. MULTI-Elite is a 500-question hard subset, while MULTI-Extend adds more than 4,500 external knowledge context pieces.
Chinese image-text multiple-choice questions with optional retrieved context.
No model card in ModelSpec reports this benchmark yet.