Dr. Bench evaluates deep research agents that decompose tasks, retrieve sources, reason across them and produce structured long reports.
unassessed
| Category | agentic |
|---|---|
| Subcategory | deep research report evaluation |
| Page status | active |
| Metric | task success rate |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 214 |
| Publisher | Dr. Bench authors |
Semantic quality, topical focus and retrieval trustworthiness of deep-research reports.
214 expert-curated tasks across 10 domains with manually constructed reference bundles.
No model card in ModelSpec reports this benchmark yet.