BrowseComp

1,266 deliberately hard-to-find, easy-to-verify questions that measure whether an agent can persistently search the web to pin down a single fact.

Also known as: BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryagentic
Subcategoryweb browsing and deep-research agents
Page statusactive
Metricaccuracy
Directionhigher_is_better
Unit%
Dataset size1266
PublisherOpenAI

What it measures

BrowseComp gives an agent a short, deliberately obscure question whose answer requires combining several hard-to-find facts from different web pages, such as identifying a specific person, place, date or event from constraints that never appear together on any one page. The agent must use search and browsing tools across many steps and return a single short answer. The task targets persistence and query reformulation rather than single-hop lookup: questions were constructed and filtered so they are not solvable by GPT-4o or o1 without live browsing, and so the top search results for the question do not already contain the answer.

Task format

Open-ended short-answer question answering with live web search/browsing tools; a single final answer is graded against a reference

Models reporting this benchmark

These figures come from the model cards, which carry one collection date per card and no per-score attribution. They are shown as reported, not as verified evidence.
ModelProviderScoreCard as of
Claude Mythos PreviewAnthropic86.92026-04
Claude Opus 4.6Anthropic83.72026-04
GLM 5.1Z.ai (Zhipu AI)68.02026-04

Data

This page as JSON · Edit on GitHub