XL-DocBench tests evidence-grounded QA on extra-long professional documents, with page-level evidence, typed rules, and unanswerable cases.
unassessed
| Category | long-context |
|---|---|
| Subcategory | evidence-grounded extra-long document question answering |
| Page status | active |
| Metric | rule-based accuracy |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 1345 |
| Dataset licence | other |
| Publisher | Microsoft Research Asia and Wuhan University |
XL-DocBench gives a system one or more long professional PDFs and a question that usually cannot be answered from a single local snippet. The model must find supporting pages, read text together with tables, charts, or figures, apply a typed verification rule, and abstain when the documents do not contain the required support. Twelve reasoning labels separate comparison, reference chains, ranking, coverage, set difference, compliance, counterfactuals, and related failures from a single accuracy number.
Extra-long document QA: page-image or OCR context (or a PDF agent), a free-form or typed answer, scored by a deterministic rule; gold evidence pages are withheld at inference.
No model card in ModelSpec reports this benchmark yet.