LCSTS (Large-scale Chinese Short Text Summarization)

Chinese short-text summarization from Sina Weibo author summaries; OpenCompass scores the 725-pair test split with jieba-tokenized ROUGE.

Also known as: Large-scale Chinese Short Text Summarization

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorygeneration
SubcategoryChinese Weibo short-text summarization (OpenCompass generation split)
Page statusactive
Metricjieba-tokenized ROUGE-1/2/L F-measure (OpenCompass)
Directionhigher_is_better
Unit%
Dataset size725
Dataset licenceApache-2.0 on the ModelScope opencompass/LCSTS card; original 2015 release terms not restated there
PublisherHarbin Institute of Technology (dataset); OpenCompass (harness config)

What it measures

LCSTS asks a model to write a short Chinese summary of a Sina Weibo post. Gold summaries were written by the post author, not by a third-party abstractor. OpenCompass feeds the post text and scores the generated line against that author summary. Character-level Chinese, generation, not multiple choice.

Task format

OpenCompass LCSTSDataset, abbreviation lcsts, path opencompass/LCSTS. Zero-shot GenInferencer. Default config lcsts_gen.py includes lcsts_gen_8ee1fe (chat-style HUMAN prompt). Alternate lcsts_gen_9b0b89 uses a single string template. Evaluator JiebaRougeEvaluator.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub