{
 "body": "## What it measures\n\nMULTI evaluates Chinese multimodal understanding using authentic examination questions. Items combine images and text and test comprehension, complex reasoning, and knowledge recall. The project also provides MULTI-Elite, a selected 500-question hard subset, and MULTI-Extend, which adds external knowledge context for in-context learning.\n\n## How it is scored\n\nThe main metric is multiple-choice accuracy. The paper reports Qwen2-VL-72B at 76.9% on MULTI and 53.1% on MULTI-Elite, compared with human expert baselines of 86.1% and 73.1%. MULTI-Extend evaluates use of supplied context and should be reported separately.\n\n## Dataset and licence\n\nThe dataset contains more than 18,000 selected and refined questions, with 500 in MULTI-Elite and more than 4,500 external knowledge context pieces in MULTI-Extend. The consulted sources do not establish a unified licence or answer visibility, so those fields remain unknown.\n\n## Who publishes it\n\nZichen Zhu and 13 coauthors introduced MULTI in a 2024 arXiv paper, later published in Science China Information Sciences. The authors link the official project page. The project is a leaderboard and dataset rather than a judge-based evaluation.\n\n## Lineage\n\nMULTI is the family page. MULTI-Elite and MULTI-Extend are named variants. MULTI-Bench is a separate spoken-dialogue benchmark despite the shared word, and is listed as a related variant only for catalogue navigation.\n\n## Saturation and contamination\n\nThe gap between the leading reported model and human experts shows substantial headroom. The benchmark is open. Because examination material and release resources are public, contamination risk is medium; model-specific exposure is not established.\n\n## How to run it\n\nUse the official questions and multiple-choice extraction protocol. Report MULTI, MULTI-Elite, or MULTI-Extend, language, image handling, context retrieval, and model prompt. Keep context-augmented results separate from the base test.\n\n## Reading the numbers\n\nA high accuracy indicates strong performance on the selected Chinese visual and textual questions. It does not establish broad multimodal reasoning outside examination formats. Compare hard-subset and context-augmented results with human baselines. Inspect image-text and knowledge-recall slices when choosing a model.\n",
 "build": {
  "built_at": "2026-09-09T16:56:50+00:00",
  "commit": "0a599558854c0e238c03a0f0d725239cb28f9d11",
  "eligibility_as_of": "2026-09-09"
 },
 "disposition": {
  "canonical_id": "multi",
  "reasons": [],
  "status": "unassessed",
  "verified_results": []
 },
 "models_covered": [],
 "page": {
  "category": "multimodal",
  "contamination": {
   "note": "The questions derive from authentic examinations and the release is public; model-specific exposure is not established.",
   "risk": "medium"
  },
  "dataset": {
   "languages": [
    "Chinese"
   ],
   "modalities": [
    "text",
    "image"
   ],
   "public_test_set": null,
   "size": 18000,
   "size_note": "More than 18,000 questions; MULTI-Elite has 500 questions and MULTI-Extend has more than 4,500 context pieces.",
   "url": "https://opendfm.github.io/"
  },
  "freshness": {
   "researched": "2026-09-08",
   "researched_by": "GPT-5.6 Luna, luna-stream-a-008 (Codex coordinated)",
   "reviewed": "",
   "reviewed_by": ""
  },
  "harness": {
   "other": "Official MULTI evaluation resources."
  },
  "id": "multi",
  "lineage": {
   "variants": [
    "multi_bench"
   ]
  },
  "measures": "MULTI tests image-text comprehension, complex reasoning, and knowledge recall against real examination standards. MULTI-Elite is a 500-question hard subset, while MULTI-Extend adds more than 4,500 external knowledge context pieces.",
  "metric": {
   "baseline_note": "The paper reports human expert baselines of 86.1% on MULTI and 73.1% on MULTI-Elite.",
   "direction": "higher_is_better",
   "human_baseline": 86.1,
   "max_score": 100,
   "name": "accuracy",
   "unit": "percent"
  },
  "name": "MULTI",
  "page_kind": "family",
  "paper": {
   "arxiv": "2402.03173",
   "title": "MULTI: Multimodal Understanding Leaderboard with Text and Images",
   "url": "https://arxiv.org/abs/2402.03173",
   "year": 2024
  },
  "publisher": {
   "authors": [
    "Zichen Zhu",
    "Yang Xu",
    "Lu Chen",
    "Jingkai Yang",
    "Yichuan Ma",
    "Yiming Sun",
    "Hailin Wen",
    "Jiaqi Liu",
    "Jinyu Cai",
    "Yingzi Ma",
    "Situo Zhang",
    "Zihan Zhao",
    "Liangtai Sun",
    "Kai Yu"
   ],
   "org": "MULTI authors",
   "url": "https://arxiv.org/abs/2402.03173"
  },
  "released": "2024-02",
  "saturation": {
   "note": "The paper reports a large gap between leading model and human expert accuracy.",
   "status": "open"
  },
  "sources": [
   {
    "accessed": "2026-09-08",
    "title": "MULTI paper",
    "url": "https://arxiv.org/abs/2402.03173"
   },
   {
    "accessed": "2026-09-08",
    "title": "Official MULTI project page",
    "url": "https://opendfm.github.io/"
   }
  ],
  "summary": "MULTI evaluates Chinese multimodal understanding with more than 18,000 authentic examination questions and hard and in-context variants.",
  "tags": [
   "Chinese",
   "multimodal",
   "examinations",
   "visual-reasoning"
  ],
  "task_format": "Chinese image-text multiple-choice questions with optional retrieved context."
 }
}