2,323 hand-picked charts from arXiv papers, paired with descriptive and reasoning questions built to resist the score inflation seen on template-based chart benchmarks.
unassessed
| Category | multimodal |
|---|---|
| Subcategory | scientific chart understanding (descriptive and reasoning questions) |
| Page status | active |
| Metric | accuracy (GPT-4o judged) |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 2323 |
| Dataset licence | CC BY-SA 4.0 (QA annotations); charts retain the copyright of their original papers; code Apache-2.0 |
| Publisher | Princeton Language and Intelligence (PLI), Princeton University |
CharXiv gives a model an image of a chart pulled from a real arXiv paper and one of two kinds of question about it. Descriptive questions ask about basic chart elements: reading off a value, counting marks, extracting a label. Reasoning questions require synthesising information across multiple visual elements of the chart -- comparing series, working out a trend, cross-referencing the legend against the axes -- rather than reading a single labelled point. Every chart and question was handpicked and verified by a human, specifically to avoid the simplified, homogeneous, template-generated charts the authors argue inflate scores on earlier chart-QA benchmarks.
Open-ended, free-text short-answer questions about a chart image (not multiple-choice). One reasoning question and four descriptive questions are written per chart. Models answer in natural language; there is no fixed answer template.
No model card in ModelSpec reports this benchmark yet.