SAGE

1,200 scientific-literature queries over a closed paper corpus that tests whether deep-research agents retrieve the right papers, not just browse the open web.

Also known as: Sage, Scientific AGentic retrieval Evaluation

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryagentic
Subcategoryscientific literature retrieval for deep-research agents
Page statusactive
Metricexact match (short-form); weighted recall (open-ended)
Directionhigher_is_better
Unit%
Dataset size1200
PublisherNYU Shanghai, Yale University, and New York University Center for Data Science

What it measures

SAGE (Scientific AGentic retrieval Evaluation) scores a deep-research agent on finding papers, not on writing a report. Each item is an English research query over a domain-specific corpus of open-access PDFs. Short-form items have one target paper and mix venue metadata, figure or table details, and citation overlap. Open-ended items mimic a literature-review request and have a ranked set of relevant papers. The skill under test is multi-step retrieval with sub-queries, not single-shot RAG. The authors contrast native web search with a controlled corpus-search tool.

Task format

English query in; the agent may think, issue search sub-queries, and answer with text plus citations. Short-form items ask for one paper. Open-ended items ask for a set of related papers.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub