226 multiple-choice questions that give a model only a biology-paper figure image, no caption or text, and ask it to reason about the figure's content.
unassessed
| Category | domain |
|---|---|
| Subcategory | biology research - figure interpretation |
| Page status | active |
| Metric | precision (correct / attempted) |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 226 |
| Dataset licence | CC BY-SA 4.0 |
| Publisher | FutureHouse |
FigQA is one of seven task categories in FutureHouse's LAB-Bench, a suite built to test practical biology-research skills rather than textbook recall. FigQA shows the model an image of a figure taken from a biology research paper, with no caption, surrounding text or paper title, and asks a multiple-choice question that usually requires integrating information from several elements of the figure at once. The task's authors describe it as a visual analogue of multi-hop text benchmarks such as HotpotQA. It requires a model to be multi-modal but was designed, by its authors, to need no external tool use: the image alone should contain everything needed to answer.
Multiple-choice question over a single figure image (no caption or paper text provided), five or more options including an explicit option to decline for lack of information.
| Model | Provider | Score | Card as of |
|---|---|---|---|
| Claude Mythos Preview | Anthropic | 79.7 | 2026-04 |
| Claude Opus 4.6 | Anthropic | 58.5 | 2026-04 |