The reasoning half of CharXiv: one open-ended question per chart that requires combining several visual elements, not just reading a labelled value.
unassessed
| Category | multimodal |
|---|---|
| Subcategory | chart reasoning over scientific figures |
| Page status | active |
| Metric | accuracy (GPT-4o judged) |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 2323 |
| Dataset licence | CC BY-SA 4.0 (QA annotations); charts retain the copyright of their original papers; code Apache-2.0 |
| Publisher | Princeton Language and Intelligence (PLI), Princeton University |
CharXiv Reasoning is the reasoning-question half of the CharXiv family: for each of 2,323 real charts pulled from arXiv papers, the model must answer one open-ended question that requires synthesising information across multiple parts of the chart -- comparing series, computing a derived value, cross-referencing the legend against the plotted data -- rather than reading a single labelled point. The benchmark's own taxonomy further splits reasoning questions into four answer-type categories: text-in-chart, text-in-general, number-in-chart and number-in-general, depending on whether the answer is text or a number and whether it is read directly off the chart or requires outside synthesis.
Open-ended, free-text short-answer question about one chart image, one per chart, not multiple-choice. Models typically reason before committing to a final answer; there is no fixed answer template beyond what individual harnesses impose.
| Model | Provider | Score | Card as of |
|---|---|---|---|
| Muse Spark | Meta | 86.4 | 2026-04 |
| Claude Mythos Preview | Anthropic | 86.1 | 2026-04 |
| Claude Opus 4.6 | Anthropic | 61.5 | 2026-04 |