RiddleSense (BIG-bench)

49 five-way riddle questions shipped in BIG-bench; not the full 5,715-item RiddleSense dataset.

Also known as: BIG-bench riddle_sense

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryreasoning
SubcategoryBIG-bench 49-item 5-way slice of the RiddleSense riddle QA dataset
Page statusunknown
Metricmultiple_choice_grade
Directionhigher_is_better
Unit%
Dataset size49
Dataset licenceApache-2.0
PublisherGoogle (BIG-bench collaboration); original dataset INK Lab, University of Southern California

What it measures

riddle_sense asks a short English riddle with five answer strings and scores the correct concept. Authors Bill Yuchen Lin, Ziyi Wu, Yichi Yang, Dong-Ho Lee, and Xiang Ren. The BIG-bench README says the task is based on the RiddleSense paper (ACL 2021 Findings; arXiv:2101.00376) and uses only the development set. task.json has 49 items, not the original 1,021-item validation split. Related work in that paper is [CommonsenseQA](commonsense_qa.md).

Task format

Five-way multiple choice, preferred_score multiple_choice_grade. append_choices_to_input is true. Dummy-model header: 49 multiple-choice queries. Canary GUID embedded.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub