SummScreen

Pairs of TV series transcripts and human-written episode recaps, testing long-document abstractive summarization where plot detail is scattered across dialogue.

Also known as: SummScreen: A Dataset for Abstractive Screenplay Summarization

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorygeneration
Subcategorylong-document abstractive summarization of TV episode transcripts into human-written recaps
Page statusactive
MetricROUGE (standard) plus two entity-centric metrics proposed by the paper for character/plot fidelity
Directionhigher_is_better
Unitscore
Dataset size26851
Dataset licenceNot established: the GitHub repository page did not state an explicit licence in the sources opened for this page.
PublisherToyota Technological Institute at Chicago (TTIC)

What it measures

SummScreen pairs a full television episode transcript (dialogue plus scene descriptions) with a human-written recap of that episode's plot. A model must produce an abstractive summary that identifies and integrates plot-relevant information scattered non-contiguously across long dialogue, while omitting comedic asides and character-development detail that do not advance the plot. This makes it a long-input, entity-centric summarization test in English: input transcripts run into the thousands of tokens (averaging roughly 6,400-7,600 tokens depending on the source), while target recaps are much shorter (roughly 110-380 tokens on average).

Task format

Free-text generation: given a full episode transcript, produce an abstractive plot recap; scored automatically and with two entity-centric metrics proposed by the paper.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub