A 14,802-example BIG-bench collection of 13 GEM generation subtasks, scored with ROUGE-Lsum, not the full GEM workshop suite.
unassessed
| Category | generation |
|---|---|
| Subcategory | BIG-bench few-shot port of GEM NLG datasets (summarization, simplification, data-to-text) |
| Page status | unknown |
| Metric | rougeLsum |
| Direction | higher_is_better |
| Dataset size | 14802 |
| Dataset licence | Apache-2.0 |
| Publisher | GEM organizers (BIG-bench submission); Google (BIG-bench collaboration) |
The BIG-bench task named gem packs thirteen modified GEM datasets into one JSON parent: sentence simplification (ASSET, TURK), concept-to-sentence (CommonGen), Czech restaurant acts, E2E restaurant descriptions, German and Spanish news summarization (MLSUM), schema-guided dialogue, WebNLG English and Russian, WikiLingua English and multilingual, and XSum. The model writes free text. The skill is generation (summarize, simplify, verbalize a table), not classification. Prompts are mixed English and other languages. This is the BIG-bench few-shot/zero-shot slice, not the living GEM workshop benchmark with 30-plus metrics.
Free-response JSON subtasks. Preferred metric rougeLsum; BLEU and ROUGE also listed. Parent task.json has no examples; each subtask folder has its own task.json. Canary GUID embedded. Dummy-model header: 0 multiple-choice and 14,802 free-text queries.
No model card in ModelSpec reports this benchmark yet.