OCRBench

1,000 hand-verified questions across five OCR task types, testing whether a multimodal model can read and reason about text embedded in images.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorymultimodal
SubcategoryOCR and document understanding
Page statusactive
Metricaggregate score across five task types
Directionhigher_is_better
Unitpoints
Dataset size1000
Dataset licenceMIT

What it measures

OCRBench evaluates optical character recognition ability in large multimodal models across five task types: plain text recognition, scene-text-centric visual question answering, document-oriented visual question answering, key information extraction, and handwritten mathematical expression recognition. Items are drawn from 29 existing OCR-related datasets and re-verified by hand into 1,000 question-answer pairs, so the benchmark is broader than any single source dataset while staying small enough to evaluate cheaply.

Task format

Open-ended, short-answer visual question answering over an image containing text; the model must read the text to answer.

Models reporting this benchmark

These figures come from the model cards, which carry one collection date per card and no per-score attribution. They are shown as reported, not as verified evidence.
ModelProviderScoreCard as of
Qwen3 VL 8B InstructAlibaba / Qwen Team89.62026-04
MiniCPM V 4OpenBMB89.42026-04
MiniCPM V 4 5OpenBMB89.42026-04
MiniCPM V 4 5 ggufOpenBMB89.42026-04
Qwen2.5-VL 72B InstructAlibaba / Qwen Team88.52026-04
Qwen3 VL 32B InstructAlibaba / Qwen Team87.52026-04
Qwen2.5-VL 7B InstructAlibaba / Qwen Team86.42026-04
Gemini 2.5 ProGoogle DeepMind85.22026-04
Gemini 2.5 Pro Preview 05-06Google DeepMind85.22026-04
Gemini 2.5 Pro Preview 06-05Google DeepMind85.22026-04
Gemini 2.5 Pro Preview TTSGoogle DeepMind85.22026-04
NVLM D 72BNVIDIA85.22026-04
Claude Sonnet 3.5Anthropic78.82026-04
Claude Sonnet 3.5 v2Anthropic78.82026-04
Pixtral Large (latest)Mistral AI78.52026-04
GPT-4oOpenAI73.62026-04
GPT-4o (2024-05-13)OpenAI73.62026-04
GPT-4o (2024-08-06)OpenAI73.62026-04
GPT-4o (2024-11-20)OpenAI73.62026-04
Pixtral 12BMistral AI68.92026-04
Gemini 2.0 FlashGoogle DeepMind68.62026-04

Data

This page as JSON · Edit on GitHub