1,200 scientific-literature queries over a closed paper corpus that tests whether deep-research agents retrieve the right papers, not just browse the open web.
unassessed
| Category | agentic |
|---|---|
| Subcategory | scientific literature retrieval for deep-research agents |
| Page status | active |
| Metric | exact match (short-form); weighted recall (open-ended) |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 1200 |
| Publisher | NYU Shanghai, Yale University, and New York University Center for Data Science |
SAGE (Scientific AGentic retrieval Evaluation) scores a deep-research agent on finding papers, not on writing a report. Each item is an English research query over a domain-specific corpus of open-access PDFs. Short-form items have one target paper and mix venue metadata, figure or table details, and citation overlap. Open-ended items mimic a literature-review request and have a ranked set of relevant papers. The skill under test is multi-step retrieval with sub-queries, not single-shot RAG. The authors contrast native web search with a controlled corpus-search tool.
English query in; the agent may think, issue search sub-queries, and answer with text plus citations. Short-form items ask for one paper. Open-ended items ask for a set of related papers.
No model card in ModelSpec reports this benchmark yet.