1,000 hand-verified questions across five OCR task types, testing whether a multimodal model can read and reason about text embedded in images.
unassessed
| Category | multimodal |
|---|---|
| Subcategory | OCR and document understanding |
| Page status | active |
| Metric | aggregate score across five task types |
| Direction | higher_is_better |
| Unit | points |
| Dataset size | 1000 |
| Dataset licence | MIT |
OCRBench evaluates optical character recognition ability in large multimodal models across five task types: plain text recognition, scene-text-centric visual question answering, document-oriented visual question answering, key information extraction, and handwritten mathematical expression recognition. Items are drawn from 29 existing OCR-related datasets and re-verified by hand into 1,000 question-answer pairs, so the benchmark is broader than any single source dataset while staying small enough to evaluate cheaply.
Open-ended, short-answer visual question answering over an image containing text; the model must read the text to answer.
| Model | Provider | Score | Card as of |
|---|---|---|---|
| Qwen3 VL 8B Instruct | Alibaba / Qwen Team | 89.6 | 2026-04 |
| MiniCPM V 4 | OpenBMB | 89.4 | 2026-04 |
| MiniCPM V 4 5 | OpenBMB | 89.4 | 2026-04 |
| MiniCPM V 4 5 gguf | OpenBMB | 89.4 | 2026-04 |
| Qwen2.5-VL 72B Instruct | Alibaba / Qwen Team | 88.5 | 2026-04 |
| Qwen3 VL 32B Instruct | Alibaba / Qwen Team | 87.5 | 2026-04 |
| Qwen2.5-VL 7B Instruct | Alibaba / Qwen Team | 86.4 | 2026-04 |
| Gemini 2.5 Pro | Google DeepMind | 85.2 | 2026-04 |
| Gemini 2.5 Pro Preview 05-06 | Google DeepMind | 85.2 | 2026-04 |
| Gemini 2.5 Pro Preview 06-05 | Google DeepMind | 85.2 | 2026-04 |
| Gemini 2.5 Pro Preview TTS | Google DeepMind | 85.2 | 2026-04 |
| NVLM D 72B | NVIDIA | 85.2 | 2026-04 |
| Claude Sonnet 3.5 | Anthropic | 78.8 | 2026-04 |
| Claude Sonnet 3.5 v2 | Anthropic | 78.8 | 2026-04 |
| Pixtral Large (latest) | Mistral AI | 78.5 | 2026-04 |
| GPT-4o | OpenAI | 73.6 | 2026-04 |
| GPT-4o (2024-05-13) | OpenAI | 73.6 | 2026-04 |
| GPT-4o (2024-08-06) | OpenAI | 73.6 | 2026-04 |
| GPT-4o (2024-11-20) | OpenAI | 73.6 | 2026-04 |
| Pixtral 12B | Mistral AI | 68.9 | 2026-04 |
| Gemini 2.0 Flash | Google DeepMind | 68.6 | 2026-04 |