Pairs of TV series transcripts and human-written episode recaps, testing long-document abstractive summarization where plot detail is scattered across dialogue.
unassessed
| Category | generation |
|---|---|
| Subcategory | long-document abstractive summarization of TV episode transcripts into human-written recaps |
| Page status | active |
| Metric | ROUGE (standard) plus two entity-centric metrics proposed by the paper for character/plot fidelity |
| Direction | higher_is_better |
| Unit | score |
| Dataset size | 26851 |
| Dataset licence | Not established: the GitHub repository page did not state an explicit licence in the sources opened for this page. |
| Publisher | Toyota Technological Institute at Chicago (TTIC) |
SummScreen pairs a full television episode transcript (dialogue plus scene descriptions) with a human-written recap of that episode's plot. A model must produce an abstractive summary that identifies and integrates plot-relevant information scattered non-contiguously across long dialogue, while omitting comedic asides and character-development detail that do not advance the plot. This makes it a long-input, entity-centric summarization test in English: input transcripts run into the thousands of tokens (averaging roughly 6,400-7,600 tokens depending on the source), while target recaps are much shorter (roughly 110-380 tokens on average).
Free-text generation: given a full episode transcript, produce an abstractive plot recap; scored automatically and with two entity-centric metrics proposed by the paper.
No model card in ModelSpec reports this benchmark yet.