∞Bench's summarisation split: produce a concise summary of a full English novel averaging around 104K tokens, scored by ROUGE-L-Sum against a web-sourced reference summary.
unassessed
| Category | long-context |
|---|---|
| Subcategory | long-document summarisation of English novels |
| Page status | active |
| Metric | ROUGE-L-Sum |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 103 |
| Dataset licence | MIT, per the OpenBMB/InfiniteBench GitHub repository; see the infinitebench family page. |
| Publisher | Department of Computer Science and Technology, Tsinghua University |
Given a full English novel (averaging about 103,500 tokens of context), the model must produce a concise summary, capped at 1,200 output tokens in the reference setup. Reference summaries are sourced from the web and manually filtered to remove non-summary content such as reader comments, then, like the other English book tasks, subjected to ∞Bench's key-entity replacement to reduce the chance a model recognises the specific book from pretraining alone. Unlike En.MC and En.QA, which test locating and combining specific facts, En.Sum tests whether a model can compress an entire long document's content, not just retrieve isolated details from it.
Free-text summary generation over a full-length novel, evaluated automatically against a single reference summary rather than by an LLM judge or human raters.
No model card in ModelSpec reports this benchmark yet.