900 multi-step web research tasks that grade an agent's full, deduplicated answer set rather than one fact.
unassessed
| Category | agentic |
|---|---|
| Subcategory | deep research / web search agents |
| Page status | active |
| Metric | F1 (over Fully Correct / Fully Incorrect / Correct with Excessive Answers) |
| Direction | higher_is_better |
| Unit | score |
| Dataset size | 900 |
| Dataset licence | Apache-2.0 |
| Publisher | Google DeepMind |
DeepSearchQA measures whether an LLM or web-connected agent can carry out a multi-step research task and return every correct piece of information, not just one plausible-sounding fact. Each of its 900 prompts spans one of 17 subject areas and requires a "causal chain" of lookups, where finding the answer to one step depends on having correctly resolved the step before it. About two-thirds of the prompts expect a set of several correct answers rather than a single value, so the task specifically targets comprehensiveness and stopping judgment, not just single-hop retrieval.
Open-ended research prompt in; an LLM or agent with web access must return a complete, deduplicated set (or single value) of correct answers, graded by a separate autorater model against a gold answer.
| Model | Provider | Score | Card as of |
|---|---|---|---|
| Muse Spark | Meta | 74.8 | 2026-04 |