∞Bench: En.Sum (English Summarisation)

∞Bench's summarisation split: produce a concise summary of a full English novel averaging around 104K tokens, scored by ROUGE-L-Sum against a web-sourced reference summary.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorylong-context
Subcategorylong-document summarisation of English novels
Page statusactive
MetricROUGE-L-Sum
Directionhigher_is_better
Unit%
Dataset size103
Dataset licenceMIT, per the OpenBMB/InfiniteBench GitHub repository; see the infinitebench family page.
PublisherDepartment of Computer Science and Technology, Tsinghua University

What it measures

Given a full English novel (averaging about 103,500 tokens of context), the model must produce a concise summary, capped at 1,200 output tokens in the reference setup. Reference summaries are sourced from the web and manually filtered to remove non-summary content such as reader comments, then, like the other English book tasks, subjected to ∞Bench's key-entity replacement to reduce the chance a model recognises the specific book from pretraining alone. Unlike En.MC and En.QA, which test locating and combining specific facts, En.Sum tests whether a model can compress an entire long document's content, not just retrieve isolated details from it.

Task format

Free-text summary generation over a full-length novel, evaluated automatically against a single reference summary rather than by an LLM judge or human raters.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub