LAB-Bench: FigQA

226 multiple-choice questions that give a model only a biology-paper figure image, no caption or text, and ask it to reason about the figure's content.

Also known as: FigQA

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorydomain
Subcategorybiology research - figure interpretation
Page statusactive
Metricprecision (correct / attempted)
Directionhigher_is_better
Unit%
Dataset size226
Dataset licenceCC BY-SA 4.0
PublisherFutureHouse

What it measures

FigQA is one of seven task categories in FutureHouse's LAB-Bench, a suite built to test practical biology-research skills rather than textbook recall. FigQA shows the model an image of a figure taken from a biology research paper, with no caption, surrounding text or paper title, and asks a multiple-choice question that usually requires integrating information from several elements of the figure at once. The task's authors describe it as a visual analogue of multi-hop text benchmarks such as HotpotQA. It requires a model to be multi-modal but was designed, by its authors, to need no external tool use: the image alone should contain everything needed to answer.

Task format

Multiple-choice question over a single figure image (no caption or paper text provided), five or more options including an explicit option to decline for lack of information.

Models reporting this benchmark

These figures come from the model cards, which carry one collection date per card and no per-score attribution. They are shown as reported, not as verified evidence.
ModelProviderScoreCard as of
Claude Mythos PreviewAnthropic79.72026-04
Claude Opus 4.6Anthropic58.52026-04

Data

This page as JSON · Edit on GitHub