The same 226 LAB-Bench FigQA questions scored when a model has an image-cropping tool and extended reasoning available, instead of answering from a single look at the figure.
unassessed
| Category | domain |
|---|---|
| Subcategory | biology research - figure interpretation, tool-augmented |
| Page status | active |
| Metric | precision (correct / attempted) |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 226 |
| Dataset licence | CC BY-SA 4.0 |
| Publisher | FutureHouse |
This id captures LAB-Bench FigQA scores produced with tool access, rather than from a single pass over the figure. The one documented case found in this research is specific and narrower than "tools" in general: Anthropic's Claude Opus 4.5 System Card (November 2025, Section 2.21) evaluates models "with a simple image cropping tool and a reasoning token budget of 32,768 tokens," contrasted against a baseline with neither tool nor extended thinking. That is a combination of one visual tool and extended reasoning, not web search or code execution. Treat any other source's "with tools" FigQA number as unverified until its own configuration is confirmed; the original LAB-Bench paper itself predicted FigQA specifically was unlikely to benefit much from tool use, since it is a perception task rather than a retrieval one.
Same figure-only multiple-choice questions as lab_bench_figqa, but the model may use an image-cropping tool (per Anthropic's documented setting) and, in that same setting, an extended reasoning budget before answering.
| Model | Provider | Score | Card as of |
|---|---|---|---|
| Claude Mythos Preview | Anthropic | 89.0 | 2026-04 |
| Claude Opus 4.6 | Anthropic | 75.1 | 2026-04 |