Multiple-choice question answering over labeled grade-school science diagrams, testing whether a model can connect diagram text, structure and layout to a question.
unassessed
| Category | multimodal |
|---|---|
| Subcategory | diagram question answering |
| Page status | active |
| Metric | accuracy |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 3088 |
| Dataset licence | CC BY-SA |
| Publisher | Allen Institute for AI (AI2) |
Given an annotated science diagram (for example a food web, the water cycle, or a plant cell) and a multiple-choice question about it, the model must read embedded text labels and diagrammatic relationships between parts and choose the correct answer from four options.
Multiple-choice visual question answering: one diagram image, one question, four answer options, single best choice.
| Model | Provider | Score | Card as of |
|---|---|---|---|
| Gemini 2.5 Pro | Google DeepMind | 95.8 | 2026-04 |
| Gemini 2.5 Pro Preview 05-06 | Google DeepMind | 95.8 | 2026-04 |
| Gemini 2.5 Pro Preview 06-05 | Google DeepMind | 95.8 | 2026-04 |
| Gemini 2.5 Pro Preview TTS | Google DeepMind | 95.8 | 2026-04 |
| GPT-4 | OpenAI | 95.2 | 2026-04 |
| GPT-4.1 | OpenAI | 95.2 | 2026-04 |
| GPT-4.1 nano | OpenAI | 95.2 | 2026-04 |
| GPT-4o | OpenAI | 94.2 | 2026-04 |
| GPT-4o (2024-05-13) | OpenAI | 94.2 | 2026-04 |
| GPT-4o (2024-08-06) | OpenAI | 94.2 | 2026-04 |
| GPT-4o (2024-11-20) | OpenAI | 94.2 | 2026-04 |
| NVLM D 72B | NVIDIA | 94.2 | 2026-04 |
| Pixtral Large (latest) | Mistral AI | 93.8 | 2026-04 |
| Llama 3.2 90B Vision | Meta | 92.3 | 2026-04 |
| Llama 3.2 90B Vision Instruct | Meta | 92.3 | 2026-04 |
| GPT-4.1 mini | OpenAI | 91.5 | 2026-04 |
| GPT-4 Turbo | OpenAI | 89.4 | 2026-04 |
| Claude Haiku 3 | Anthropic | 88.7 | 2026-04 |
| Qwen2.5-VL 72B Instruct | Alibaba / Qwen Team | 88.7 | 2026-04 |
| Claude Opus 3 | Anthropic | 88.1 | 2026-04 |
| Claude Sonnet 3 | Anthropic | 88.1 | 2026-04 |
| Qwen3 VL 8B Instruct | Alibaba / Qwen Team | 85.7 | 2026-04 |
| Gemma 3 27B | Google DeepMind | 84.5 | 2026-04 |
| Gemma 3 12B | Google DeepMind | 84.2 | 2026-04 |
| Gemini 2.0 Flash | Google DeepMind | 83.3 | 2026-04 |
| Qwen2 VL 7B Instruct | Alibaba / Qwen Team | 83.0 | 2025-03 |
| Qwen2 VL 7B Instruct AWQ | Alibaba / Qwen Team | 83.0 | 2025-03 |
| MiniCPM V 4 | OpenBMB | 82.9 | 2026-04 |
| MiniCPM V 4 5 | OpenBMB | 82.9 | 2026-04 |
| MiniCPM V 4 5 gguf | OpenBMB | 82.9 | 2026-04 |
| Gemini 1.5 Pro | Google DeepMind | 82.5 | 2026-04 |
| Claude Sonnet 3.5 | Anthropic | 80.2 | 2026-04 |
| Claude Sonnet 3.5 v2 | Anthropic | 80.2 | 2026-04 |
| Pixtral 12B | Mistral AI | 79.5 | 2026-04 |
| Qwen2.5-VL 7B Instruct | Alibaba / Qwen Team | 79.5 | 2026-04 |
| Gemini 2.0 Flash Lite | Google DeepMind | 78.5 | 2026-04 |
| Phi 3.5 vision instruct | Microsoft | 78.1 | 2026-04 |
| Gemma 3 4B | Google DeepMind | 74.8 | 2026-04 |
| gemma 3 4B pt | Google DeepMind | 74.8 | 2026-04 |