Fantasy Reasoning

A 201-item BIG-bench Yes/No task that asks models to reason inside fantasy premises that violate real-world common sense.

Also known as: BIG-bench fantasy_reasoning

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryreasoning
SubcategoryBIG-bench Yes/No reasoning under fantasy premises (201 items)
Page statusunknown
Metricmultiple_choice_grade
Directionhigher_is_better
Unit%
Dataset size201
Dataset licenceApache-2.0
PublisherGoogle (BIG-bench collaboration); University of Amsterdam authors

What it measures

fantasy_reasoning gives an English fantasy or science-fiction premise from r/writingprompts, then a yes/no question that is only answerable if the model accepts the premise's local rules. The authors' claim is that ordinary commonsense benches can be solved by interpolating web text, whereas these contexts should be rare. Comments in task.json mark items implicit or explicit. It is not [bbh](bbh.md) and not a story-generation task.

Task format

Two-option multiple choice (Yes. / No.), preferred metric multiple_choice_grade. Input is context concatenated with the question. append_choices_to_input is false. Canary GUID embedded.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub