Question answering over scanned and typed document images, scored by fuzzy text match against reference answers.
unassessed
| Category | multimodal |
|---|---|
| Subcategory | document understanding |
| Page status | active |
| Metric | ANLS (Average Normalized Levenshtein Similarity) |
| Direction | higher_is_better |
| Unit | score (0-1) |
| Dataset size | 50000 |
| Publisher | Computer Vision Center (Universitat Autònoma de Barcelona) and IIIT Hyderabad, with Amazon |
DocVQA gives a model an image of a real document — a letter, memo, form, table or report — plus a natural-language question about its content, and asks for a short free-text answer. Questions require reading text in context (totals in a table, a form field's value, a handwritten note, a figure label) rather than isolated OCR, so the task combines text recognition, layout understanding and reading comprehension in a single English-language, single-image, single-turn call.
Document image plus a natural-language question in; a short free-text answer out.
| Model | Provider | Score | Card as of |
|---|---|---|---|
| Qwen3 VL 32B Instruct | Alibaba / Qwen Team | 96.5 | 2026-04 |
| Qwen3 VL 8B Instruct | Alibaba / Qwen Team | 96.1 | 2026-04 |
| Claude Sonnet 3.5 | Anthropic | 95.2 | 2026-04 |
| Claude Sonnet 3.5 v2 | Anthropic | 95.2 | 2026-04 |
| Qwen2 VL 7B Instruct | Alibaba / Qwen Team | 94.5 | 2025-03 |
| Qwen2 VL 7B Instruct AWQ | Alibaba / Qwen Team | 94.5 | 2025-03 |
| Llama 4 Maverick 17B 128E Instruct | Meta | 94.4 | 2026-04 |
| Llama 4 Scout 17B 16E | Meta | 94.4 | 2026-04 |
| Llama 4 Scout 17B 16E Instruct | Meta | 94.4 | 2026-04 |
| Llama-4-Maverick-17B-128E-Instruct-FP8 | Meta | 94.4 | 2026-04 |
| Llama-4-Scout-17B-16E-Instruct-FP8 | Meta | 94.4 | 2026-04 |
| GPT-4 | OpenAI | 94.1 | 2026-04 |
| GPT-4.1 | OpenAI | 94.1 | 2026-04 |
| GPT-4.1 mini | OpenAI | 94.1 | 2026-04 |
| GPT-4.1 nano | OpenAI | 94.1 | 2026-04 |
| Gemini 2.5 Pro | Google DeepMind | 93.8 | 2026-04 |
| Gemini 2.5 Pro Preview 05-06 | Google DeepMind | 93.8 | 2026-04 |
| Gemini 2.5 Pro Preview 06-05 | Google DeepMind | 93.8 | 2026-04 |
| Gemini 2.5 Pro Preview TTS | Google DeepMind | 93.8 | 2026-04 |
| Claude Opus 4 | Anthropic | 93.2 | 2026-04 |
| Claude Opus 4.6 | Anthropic | 93.2 | 2026-04 |
| MiniCPM V 4 | OpenBMB | 92.9 | 2026-04 |
| MiniCPM V 4 5 | OpenBMB | 92.9 | 2026-04 |
| MiniCPM V 4 5 gguf | OpenBMB | 92.9 | 2026-04 |
| GPT-4o | OpenAI | 92.8 | 2026-04 |
| GPT-4o (2024-05-13) | OpenAI | 92.8 | 2026-04 |
| GPT-4o (2024-08-06) | OpenAI | 92.8 | 2026-04 |
| GPT-4o (2024-11-20) | OpenAI | 92.8 | 2026-04 |
| GPT-4o mini | OpenAI | 92.8 | 2026-04 |
| NVLM D 72B | NVIDIA | 92.6 | 2026-04 |
| Claude Sonnet 4 | Anthropic | 91.5 | 2026-04 |
| Claude Sonnet 4.5 | Anthropic | 91.5 | 2026-04 |
| Claude Sonnet 4.5 (latest) | Anthropic | 91.5 | 2026-04 |
| Llama 3.2 90B Vision | Meta | 90.1 | 2026-04 |
| Llama 3.2 90B Vision Instruct | Meta | 90.1 | 2026-04 |
| Qwen2.5-VL 72B Instruct | Alibaba / Qwen Team | 90.1 | 2026-04 |
| Claude Opus 3 | Anthropic | 89.3 | 2026-04 |
| Claude Sonnet 3 | Anthropic | 89.3 | 2026-04 |
| Claude Haiku 3 | Anthropic | 88.8 | 2026-04 |
| Gemini 1.5 Pro | Google DeepMind | 88.5 | 2026-04 |
| Llama 3.2 11B Vision | Meta | 88.4 | 2026-04 |
| Llama 3.2 11B Vision Instruct | Meta | 88.4 | 2026-04 |
| Gemini 2.0 Flash | Google DeepMind | 88.1 | 2026-04 |
| Gemini 2.0 Flash Lite | Google DeepMind | 88.1 | 2026-04 |
| GPT-4 Turbo | OpenAI | 87.2 | 2026-04 |
| Gemma 3 12B | Google DeepMind | 87.1 | 2026-04 |
| Gemma 4 31B | Google DeepMind | 86.8 | 2026-04 |
| gemma 4 31B it | Google DeepMind | 86.8 | 2026-04 |
| gemma 4 31B it GGUF | Unsloth | 86.8 | 2026-04 |
| Gemma 4 31B IT NVFP4 | NVIDIA | 86.8 | 2026-04 |
| Pixtral Large (latest) | Mistral AI | 85.5 | 2026-04 |
| Gemini 1.5 Flash | Google DeepMind | 85.2 | 2026-04 |
| Gemini 1.5 Flash-8B | Google DeepMind | 85.2 | 2026-04 |
| Qwen2.5 VL 3B Instruct | Alibaba / Qwen Team | 85.2 | 2026-04 |
| Mistral Large (latest) | Mistral AI | 83.5 | 2026-04 |
| Mistral Large 3 | Mistral AI | 83.5 | 2026-04 |
| Grok 2 | xAI | 82.5 | 2026-04 |
| Grok 2 Latest | xAI | 82.5 | 2026-04 |
| Grok 2 Vision | xAI | 82.5 | 2026-04 |
| Grok 2 Vision (1212) | xAI | 82.5 | 2026-04 |
| Grok 2 Vision Latest | xAI | 82.5 | 2026-04 |
| Gemma 3 27B | Google DeepMind | 82.1 | 2026-04 |
| Command A Vision | Cohere | 81.2 | 2026-04 |
| Pixtral 12B | Mistral AI | 78.8 | 2026-04 |
| Qwen2.5-VL 7B Instruct | Alibaba / Qwen Team | 78.5 | 2026-04 |
| Gemma 3 4B | Google DeepMind | 75.8 | 2026-04 |
| gemma 3 4B pt | Google DeepMind | 75.8 | 2026-04 |