11.5K college-level, image-paired exam questions across six disciplines, testing expert knowledge that requires reading a figure, chart or diagram.
unassessed
| Category | multimodal |
|---|---|
| Subcategory | expert multi-discipline knowledge |
| Page status | active |
| Metric | accuracy |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 11500 |
| Dataset licence | Apache-2.0 |
MMMU pairs text questions with images such as charts, diagrams, maps, tables, music sheets and chemical structures, drawn from real college exams, quizzes and textbooks. It spans six core disciplines (Art and Design, Business, Science, Health and Medicine, Humanities and Social Science, and Tech and Engineering) across 30 subjects and 183 subfields, and 30 distinct image types. The goal is to test whether a model can combine expert-level subject knowledge with genuine reading of the accompanying image, rather than treating the image as decoration for a question answerable from text alone.
Mostly four-option multiple-choice questions, each paired with one or more images; a smaller portion are open-ended.
| Model | Provider | Score | Card as of |
|---|---|---|---|
| Llama 4 Maverick 17B 128E Instruct | Meta | 73.4 | 2026-04 |
| Llama-4-Maverick-17B-128E-Instruct-FP8 | Meta | 73.4 | 2026-04 |
| Qwen3 VL 32B Instruct | Alibaba / Qwen Team | 72.8 | 2026-04 |
| GPT-4 | OpenAI | 72.5 | 2026-04 |
| GPT-4.1 | OpenAI | 72.5 | 2026-04 |
| GPT-4.1 mini | OpenAI | 72.5 | 2026-04 |
| GPT-4.1 nano | OpenAI | 72.5 | 2026-04 |
| Gemini 2.5 Pro | Google DeepMind | 72.1 | 2026-04 |
| Gemini 2.5 Pro Preview 05-06 | Google DeepMind | 72.1 | 2026-04 |
| Gemini 2.5 Pro Preview 06-05 | Google DeepMind | 72.1 | 2026-04 |
| Gemini 2.5 Pro Preview TTS | Google DeepMind | 72.1 | 2026-04 |
| Claude Opus 4 | Anthropic | 70.8 | 2026-04 |
| Claude Opus 4.6 | Anthropic | 70.8 | 2026-04 |
| Llama 4 Scout 17B 16E | Meta | 69.4 | 2026-04 |
| Llama 4 Scout 17B 16E Instruct | Meta | 69.4 | 2026-04 |
| Llama-4-Scout-17B-16E-Instruct-FP8 | Meta | 69.4 | 2026-04 |
| GPT-4o | OpenAI | 69.1 | 2026-04 |
| GPT-4o (2024-05-13) | OpenAI | 69.1 | 2026-04 |
| GPT-4o (2024-08-06) | OpenAI | 69.1 | 2026-04 |
| GPT-4o (2024-11-20) | OpenAI | 69.1 | 2026-04 |
| GPT-4o mini | OpenAI | 69.1 | 2026-04 |
| Claude Sonnet 4 | Anthropic | 68.2 | 2026-04 |
| Claude Sonnet 4.5 | Anthropic | 68.2 | 2026-04 |
| Claude Sonnet 4.5 (latest) | Anthropic | 68.2 | 2026-04 |
| Claude Sonnet 3.5 | Anthropic | 65.9 | 2026-04 |
| Claude Sonnet 3.5 v2 | Anthropic | 65.9 | 2026-04 |
| Qwen2.5-VL 72B Instruct | Alibaba / Qwen Team | 64.5 | 2026-04 |
| Gemini 1.5 Pro | Google DeepMind | 62.8 | 2026-04 |
| Gemini 2.0 Flash | Google DeepMind | 62.5 | 2026-04 |
| Gemini 2.0 Flash Lite | Google DeepMind | 62.5 | 2026-04 |
| Qwen3 VL 8B Instruct | Alibaba / Qwen Team | 62.5 | 2026-04 |
| Llama 3.2 90B Vision | Meta | 60.3 | 2026-04 |
| Llama 3.2 90B Vision Instruct | Meta | 60.3 | 2026-04 |
| Gemma 4 31B | Google DeepMind | 60.2 | 2026-04 |
| gemma 4 31B it | Google DeepMind | 60.2 | 2026-04 |
| gemma 4 31B it GGUF | Unsloth | 60.2 | 2026-04 |
| Gemma 4 31B IT NVFP4 | NVIDIA | 60.2 | 2026-04 |
| Gemma 3 12B | Google DeepMind | 59.6 | 2026-04 |
| Claude Opus 3 | Anthropic | 59.4 | 2026-04 |
| GPT-4 Turbo | OpenAI | 59.4 | 2026-04 |
| Pixtral Large (latest) | Mistral AI | 58.8 | 2026-04 |
| NVLM D 72B | NVIDIA | 58.7 | 2026-04 |
| Mistral Large (latest) | Mistral AI | 56.8 | 2026-04 |
| Mistral Large 3 | Mistral AI | 56.8 | 2026-04 |
| Gemini 1.5 Flash | Google DeepMind | 56.1 | 2026-04 |
| Gemini 1.5 Flash-8B | Google DeepMind | 56.1 | 2026-04 |
| Grok 2 | xAI | 55.5 | 2026-04 |
| Grok 2 Latest | xAI | 55.5 | 2026-04 |
| Grok 2 Vision | xAI | 55.5 | 2026-04 |
| Grok 2 Vision (1212) | xAI | 55.5 | 2026-04 |
| Grok 2 Vision Latest | xAI | 55.5 | 2026-04 |
| Gemma 3 27B | Google DeepMind | 55.2 | 2026-04 |
| Qwen2 VL 7B Instruct | Alibaba / Qwen Team | 54.1 | 2025-03 |
| Qwen2 VL 7B Instruct AWQ | Alibaba / Qwen Team | 54.1 | 2025-03 |
| Claude Sonnet 3 | Anthropic | 53.1 | 2026-04 |
| Command A Vision | Cohere | 52.1 | 2026-04 |
| MiniCPM V 4 | OpenBMB | 51.2 | 2026-04 |
| MiniCPM V 4 5 | OpenBMB | 51.2 | 2026-04 |
| MiniCPM V 4 5 gguf | OpenBMB | 51.2 | 2026-04 |
| Llama 3.2 11B Vision | Meta | 50.7 | 2026-04 |
| Llama 3.2 11B Vision Instruct | Meta | 50.7 | 2026-04 |
| Claude Haiku 3 | Anthropic | 50.2 | 2026-04 |
| Pixtral 12B | Mistral AI | 50.2 | 2026-04 |
| Gemma 3 4B | Google DeepMind | 48.8 | 2026-04 |
| gemma 3 4B pt | Google DeepMind | 48.8 | 2026-04 |
| Qwen2.5-VL 7B Instruct | Alibaba / Qwen Team | 48.2 | 2026-04 |
| Phi 3.5 vision instruct | Microsoft | 43.0 | 2026-04 |
| Qwen2.5 VL 3B Instruct | Alibaba / Qwen Team | 42.5 | 2026-04 |