1,266 deliberately hard-to-find, easy-to-verify questions that measure whether an agent can persistently search the web to pin down a single fact.
unassessed
| Category | agentic |
|---|---|
| Subcategory | web browsing and deep-research agents |
| Page status | active |
| Metric | accuracy |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 1266 |
| Publisher | OpenAI |
BrowseComp gives an agent a short, deliberately obscure question whose answer requires combining several hard-to-find facts from different web pages, such as identifying a specific person, place, date or event from constraints that never appear together on any one page. The agent must use search and browsing tools across many steps and return a single short answer. The task targets persistence and query reformulation rather than single-hop lookup: questions were constructed and filtered so they are not solvable by GPT-4o or o1 without live browsing, and so the top search results for the question do not already contain the answer.
Open-ended short-answer question answering with live web search/browsing tools; a single final answer is graded against a reference
| Model | Provider | Score | Card as of |
|---|---|---|---|
| Claude Mythos Preview | Anthropic | 86.9 | 2026-04 |
| Claude Opus 4.6 | Anthropic | 83.7 | 2026-04 |
| GLM 5.1 | Z.ai (Zhipu AI) | 68.0 | 2026-04 |