A six-language VLM benchmark of ~122 OECD PISA exam items with images, scored as accuracy in lm-eval groups pisa and pisa_llm_judged.
unassessed
| Category | multimodal |
|---|---|
| Subcategory | multilingual vision-language multiple choice from OECD PISA items |
| Page status | unknown |
| Metric | accuracy |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 122 |
| Publisher | Humboldt-Universität zu Berlin and DFKI |
PISA-Bench (Haller, Barth, Golde, Rehm, Akbik; arXiv:2510.24792) is a vision-language benchmark built from publicly available OECD Programme for International Student Assessment materials from 2012 and earlier. Each item pairs an English diagram or table with a question that needs the image. Human annotators kept complete, clear, multimodal questions; GPT-4o filled missing multiple-choice options and assigned one of four categories: spatial and geometric reasoning, quantitative reasoning, graph and pattern analysis, and text and diagram understanding. GPT-4 translations yield a parallel corpus in English, German, Spanish, French, Italian, and Chinese. This page is that VLM eval, not the OECD student league table and not a PISA IRT score.
Image plus text. lm-eval default tasks are generate_until multiple-choice with options A–D, prompt "Given the provided image <image>, answer following questions:". Group pisa uses substring matching; group pisa_llm_judged uses an OpenAI chat judge (default model gpt-4.1-mini via MODEL_VERSION). Paper §4 describes free-form generation judged by GPT-4; Table 3's caption and the Hub card name chatgpt-4o-mini / GPT-4o-mini.
No model card in ModelSpec reports this benchmark yet.