Dr. Bench

Dr. Bench evaluates deep research agents that decompose tasks, retrieve sources, reason across them and produce structured long reports.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryagentic
Subcategorydeep research report evaluation
Page statusactive
Metrictask success rate
Directionhigher_is_better
Unit%
Dataset size214
PublisherDr. Bench authors

What it measures

Semantic quality, topical focus and retrieval trustworthiness of deep-research reports.

Task format

214 expert-curated tasks across 10 domains with manually constructed reference bundles.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub