Visual question answering over real-world chart images that require reading values off the chart and doing arithmetic or logical reasoning to answer.
unassessed
| Category | multimodal |
|---|---|
| Subcategory | chart and figure question answering |
| Page status | active |
| Metric | relaxed accuracy |
| Direction | higher_is_better |
| Unit | % |
| Dataset licence | GPL-3.0 |
ChartQA shows a model a chart image, drawn from real published sources, together with a natural-language question about it. The model must produce a short answer, which may require reading a value directly off the chart, comparing two or more values, or doing simple arithmetic or logical reasoning across several data points. Roughly three in ten question-answer pairs come from human annotators who wrote compositional, visually grounded questions; the rest were generated from human-written chart summaries and then checked. The task exercises visual perception of chart elements, such as axes, legends and bar heights, together with the reasoning needed to compute an answer, not just lookup.
Visual question answering over chart images; open-vocabulary short answers (numbers, phrases, yes/no)
| Model | Provider | Score | Card as of |
|---|---|---|---|
| Claude Sonnet 3.5 | Anthropic | 90.8 | 2026-04 |
| Claude Sonnet 3.5 v2 | Anthropic | 90.8 | 2026-04 |
| Llama 4 Maverick 17B 128E Instruct | Meta | 90.0 | 2026-04 |
| Llama-4-Maverick-17B-128E-Instruct-FP8 | Meta | 90.0 | 2026-04 |
| Llama 4 Scout 17B 16E | Meta | 88.8 | 2026-04 |
| Llama 4 Scout 17B 16E Instruct | Meta | 88.8 | 2026-04 |
| Llama-4-Scout-17B-16E-Instruct-FP8 | Meta | 88.8 | 2026-04 |
| GPT-4 | OpenAI | 88.5 | 2026-04 |
| GPT-4.1 | OpenAI | 88.5 | 2026-04 |
| GPT-4.1 mini | OpenAI | 88.5 | 2026-04 |
| GPT-4.1 nano | OpenAI | 88.5 | 2026-04 |
| Gemini 2.5 Pro | Google DeepMind | 87.5 | 2026-04 |
| Gemini 2.5 Pro Preview 05-06 | Google DeepMind | 87.5 | 2026-04 |
| Gemini 2.5 Pro Preview 06-05 | Google DeepMind | 87.5 | 2026-04 |
| Gemini 2.5 Pro Preview TTS | Google DeepMind | 87.5 | 2026-04 |
| GPT-4 Turbo | OpenAI | 87.2 | 2026-04 |
| Claude Opus 4 | Anthropic | 86.8 | 2026-04 |
| Claude Opus 4.6 | Anthropic | 86.8 | 2026-04 |
| NVLM D 72B | NVIDIA | 86.0 | 2026-04 |
| Llama 3.2 90B Vision | Meta | 85.5 | 2026-04 |
| Llama 3.2 90B Vision Instruct | Meta | 85.5 | 2026-04 |
| GPT-4o | OpenAI | 85.2 | 2026-04 |
| GPT-4o (2024-05-13) | OpenAI | 85.2 | 2026-04 |
| GPT-4o (2024-08-06) | OpenAI | 85.2 | 2026-04 |
| GPT-4o (2024-11-20) | OpenAI | 85.2 | 2026-04 |
| GPT-4o mini | OpenAI | 85.2 | 2026-04 |
| MiniCPM V 4 | OpenBMB | 84.4 | 2026-04 |
| MiniCPM V 4 5 | OpenBMB | 84.4 | 2026-04 |
| MiniCPM V 4 5 gguf | OpenBMB | 84.4 | 2026-04 |
| Claude Sonnet 4 | Anthropic | 84.2 | 2026-04 |
| Claude Sonnet 4.5 | Anthropic | 84.2 | 2026-04 |
| Claude Sonnet 4.5 (latest) | Anthropic | 84.2 | 2026-04 |
| Qwen2 VL 7B Instruct | Alibaba / Qwen Team | 83.0 | 2025-03 |
| Qwen2 VL 7B Instruct AWQ | Alibaba / Qwen Team | 83.0 | 2025-03 |
| Qwen2.5-VL 72B Instruct | Alibaba / Qwen Team | 82.5 | 2026-04 |
| Phi 3.5 vision instruct | Microsoft | 81.8 | 2026-04 |
| Claude Haiku 3 | Anthropic | 81.1 | 2026-04 |
| Claude Opus 3 | Anthropic | 80.8 | 2026-04 |
| Claude Sonnet 3 | Anthropic | 80.8 | 2026-04 |
| Gemini 1.5 Pro | Google DeepMind | 79.2 | 2026-04 |
| Gemini 2.0 Flash | Google DeepMind | 78.5 | 2026-04 |
| Gemini 2.0 Flash Lite | Google DeepMind | 78.5 | 2026-04 |
| Gemma 4 31B | Google DeepMind | 78.5 | 2026-04 |
| gemma 4 31B it | Google DeepMind | 78.5 | 2026-04 |
| gemma 4 31B it GGUF | Unsloth | 78.5 | 2026-04 |
| Gemma 4 31B IT NVFP4 | NVIDIA | 78.5 | 2026-04 |
| Pixtral Large (latest) | Mistral AI | 76.2 | 2026-04 |
| Gemma 3 12B | Google DeepMind | 75.7 | 2026-04 |
| Gemini 1.5 Flash | Google DeepMind | 74.8 | 2026-04 |
| Gemini 1.5 Flash-8B | Google DeepMind | 74.8 | 2026-04 |
| Mistral Large (latest) | Mistral AI | 74.5 | 2026-04 |
| Mistral Large 3 | Mistral AI | 74.5 | 2026-04 |
| Grok 2 | xAI | 73.2 | 2026-04 |
| Grok 2 Latest | xAI | 73.2 | 2026-04 |
| Grok 2 Vision | xAI | 73.2 | 2026-04 |
| Grok 2 Vision (1212) | xAI | 73.2 | 2026-04 |
| Grok 2 Vision Latest | xAI | 73.2 | 2026-04 |
| Gemma 3 27B | Google DeepMind | 72.5 | 2026-04 |
| Qwen2.5 VL 3B Instruct | Alibaba / Qwen Team | 72.5 | 2026-04 |
| Command A Vision | Cohere | 70.5 | 2026-04 |
| Gemma 3 4B | Google DeepMind | 68.8 | 2026-04 |
| gemma 3 4B pt | Google DeepMind | 68.8 | 2026-04 |
| Pixtral 12B | Mistral AI | 66.5 | 2026-04 |
| Qwen2.5-VL 7B Instruct | Alibaba / Qwen Team | 65.8 | 2026-04 |