XL-DocBench

XL-DocBench tests evidence-grounded QA on extra-long professional documents, with page-level evidence, typed rules, and unanswerable cases.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorylong-context
Subcategoryevidence-grounded extra-long document question answering
Page statusactive
Metricrule-based accuracy
Directionhigher_is_better
Unit%
Dataset size1345
Dataset licenceother
PublisherMicrosoft Research Asia and Wuhan University

What it measures

XL-DocBench gives a system one or more long professional PDFs and a question that usually cannot be answered from a single local snippet. The model must find supporting pages, read text together with tables, charts, or figures, apply a typed verification rule, and abstain when the documents do not contain the required support. Twelve reasoning labels separate comparison, reference chains, ranking, coverage, set difference, compliance, counterfactuals, and related failures from a single accuracy number.

Task format

Extra-long document QA: page-image or OCR context (or a PDF agent), a free-form or typed answer, scored by a deterministic rule; gold evidence pages are withheld at inference.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub