A 76-item BIG-bench similarities test that asks how two objects are alike, scoring the abstract option over concrete distractors.
unassessed
| Category | reasoning |
|---|---|
| Subcategory | BIG-bench MoCA/WAIS-style object similarities, abstract vs concrete (76 items) |
| Page status | unknown |
| Metric | multiple_choice_grade |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 76 |
| Dataset licence | Apache-2.0 |
| Publisher | Google (BIG-bench collaboration) |
similarities_abstraction names two objects and asks how they are alike, after a fruit example in task_prefix, following MoCA/WAIS similarities practice. M. Yee wrote one abstract gold and several concrete distractors per item (truthful but too specific). Direct count: 76 examples (66 with four choices, 7 with five, 3 with six). Preferred metric multiple_choice_grade; BLEU and ROUGE are also listed for optional free response. Dummy-model header: 76 multiple-choice and 76 free-text queries. Canary GUID embedded. Not in BIG-Bench Hard.
Multiple choice with append_choices_to_input false, plus optional free-response targets for BLEU and ROUGE. JSON task. Keywords analogical reasoning, context-free question answering, free response, human-like behavior, json, multiple choice.
No model card in ModelSpec reports this benchmark yet.