A 201-item BIG-bench Yes/No task that asks models to reason inside fantasy premises that violate real-world common sense.
unassessed
| Category | reasoning |
|---|---|
| Subcategory | BIG-bench Yes/No reasoning under fantasy premises (201 items) |
| Page status | unknown |
| Metric | multiple_choice_grade |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 201 |
| Dataset licence | Apache-2.0 |
| Publisher | Google (BIG-bench collaboration); University of Amsterdam authors |
fantasy_reasoning gives an English fantasy or science-fiction premise from r/writingprompts, then a yes/no question that is only answerable if the model accepts the premise's local rules. The authors' claim is that ordinary commonsense benches can be solved by interpolating web text, whereas these contexts should be rare. Comments in task.json mark items implicit or explicit. It is not [bbh](bbh.md) and not a story-generation task.
Two-option multiple choice (Yes. / No.), preferred metric multiple_choice_grade. Input is context concatenated with the question. append_choices_to_input is false. Canary GUID embedded.
No model card in ModelSpec reports this benchmark yet.