CharXiv

2,323 hand-picked charts from arXiv papers, paired with descriptive and reasoning questions built to resist the score inflation seen on template-based chart benchmarks.

Also known as: CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorymultimodal
Subcategoryscientific chart understanding (descriptive and reasoning questions)
Page statusactive
Metricaccuracy (GPT-4o judged)
Directionhigher_is_better
Unit%
Dataset size2323
Dataset licenceCC BY-SA 4.0 (QA annotations); charts retain the copyright of their original papers; code Apache-2.0
PublisherPrinceton Language and Intelligence (PLI), Princeton University

What it measures

CharXiv gives a model an image of a chart pulled from a real arXiv paper and one of two kinds of question about it. Descriptive questions ask about basic chart elements: reading off a value, counting marks, extracting a label. Reasoning questions require synthesising information across multiple visual elements of the chart -- comparing series, working out a trend, cross-referencing the legend against the axes -- rather than reading a single labelled point. Every chart and question was handpicked and verified by a human, specifically to avoid the simplified, homogeneous, template-generated charts the authors argue inflate scores on earlier chart-QA benchmarks.

Task format

Open-ended, free-text short-answer questions about a chart image (not multiple-choice). One reasoning question and four descriptive questions are written per chart. Models answer in natural language; there is no fixed answer template.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub