GEM (BIG-bench)

A 14,802-example BIG-bench collection of 13 GEM generation subtasks, scored with ROUGE-Lsum, not the full GEM workshop suite.

Also known as: BIG-bench gem, Various Generation Skills Benchmark

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorygeneration
SubcategoryBIG-bench few-shot port of GEM NLG datasets (summarization, simplification, data-to-text)
Page statusunknown
MetricrougeLsum
Directionhigher_is_better
Dataset size14802
Dataset licenceApache-2.0
PublisherGEM organizers (BIG-bench submission); Google (BIG-bench collaboration)

What it measures

The BIG-bench task named gem packs thirteen modified GEM datasets into one JSON parent: sentence simplification (ASSET, TURK), concept-to-sentence (CommonGen), Czech restaurant acts, E2E restaurant descriptions, German and Spanish news summarization (MLSUM), schema-guided dialogue, WebNLG English and Russian, WikiLingua English and multilingual, and XSum. The model writes free text. The skill is generation (summarize, simplify, verbalize a table), not classification. Prompts are mixed English and other languages. This is the BIG-bench few-shot/zero-shot slice, not the living GEM workshop benchmark with 30-plus metrics.

Task format

Free-response JSON subtasks. Preferred metric rougeLsum; BLEU and ROUGE also listed. Parent task.json has no examples; each subtask folder has its own task.json. Canary GUID embedded. Dummy-model header: 0 multiple-choice and 14,802 free-text queries.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub