LAB-Bench

FutureHouse's eight-category, 2,457-question suite testing practical biology-research skills -- literature QA, figure/table reading, database and sequence work, and molecular cloning -- rather than textbook recall.

Also known as: Language Agent Biology Benchmark, lab-bench

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorydomain
Subcategorybiology research capabilities: literature, figures, tables, databases, protocols, sequences
Page statusactive
Metricprecision (correct / attempted), with accuracy and coverage reported alongside
Directionhigher_is_better
Unit%
Dataset size2457
Dataset licenceCC BY-SA 4.0
PublisherFutureHouse

What it measures

LAB-Bench tests whether a model can perform the practical, day-to-day tasks of biology research rather than answer textbook-style knowledge questions. It comprises eight categories: LitQA2 (answer questions by finding and reading full-text scientific papers) and SuppQA (the same, but the answer lives in a paper's supplementary material) test literature retrieval; FigQA and TableQA test interpreting a figure or table image stripped of its caption and surrounding text; DbQA tests navigating real bioinformatics databases across 10 narrower subtasks; ProtocolQA tests spotting an inserted error in a published wet-lab protocol; SeqQA tests reasoning about DNA, RNA and protein sequences across 15 subtasks; and CloningScenarios is a set of 41 deliberately "human-hard" multi-step molecular-cloning questions the authors expect may take a trained biologist over ten minutes each. LitQA2, SuppQA and DbQA are explicitly meant to be run with retrieval tools; SeqQA needs sequence-manipulation tools; FigQA, TableQA and ProtocolQA are designed to be tool-free tests of reasoning over what is already given.

Task format

Multiple-choice questions, each with an explicit "insufficient information" option to decline; FigQA and TableQA additionally supply a figure or table image. Category-appropriate agent or tool access is part of several categories' intended protocol rather than optional, so "LAB-Bench" results are not one uniform task format across the suite.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub