DeepSearchQA

900 multi-step web research tasks that grade an agent's full, deduplicated answer set rather than one fact.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryagentic
Subcategorydeep research / web search agents
Page statusactive
MetricF1 (over Fully Correct / Fully Incorrect / Correct with Excessive Answers)
Directionhigher_is_better
Unitscore
Dataset size900
Dataset licenceApache-2.0
PublisherGoogle DeepMind

What it measures

DeepSearchQA measures whether an LLM or web-connected agent can carry out a multi-step research task and return every correct piece of information, not just one plausible-sounding fact. Each of its 900 prompts spans one of 17 subject areas and requires a "causal chain" of lookups, where finding the answer to one step depends on having correctly resolved the step before it. About two-thirds of the prompts expect a set of several correct answers rather than a single value, so the task specifically targets comprehensiveness and stopping judgment, not just single-hop retrieval.

Task format

Open-ended research prompt in; an LLM or agent with web access must return a complete, deduplicated set (or single value) of correct answers, graded by a separate autorater model against a gold answer.

Models reporting this benchmark

These figures come from the model cards, which carry one collection date per card and no per-score attribution. They are shown as reported, not as verified evidence.
ModelProviderScoreCard as of
Muse SparkMeta74.82026-04

Data

This page as JSON · Edit on GitHub