BIG-bench task built on a clinical theory-of-mind battery that asks a model to infer characters' beliefs, intentions and non-literal meaning from short narratives.
unassessed
| Category | reasoning |
|---|---|
| Subcategory | theory of mind / social reasoning |
| Page status | active |
| Metric | multiple choice grade |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 174 |
| Dataset licence | Apache-2.0 |
| Publisher | Google (BIG-bench collaboration) |
Strange Stories gives the model a short naturalistic narrative in which a character says something that is not literally true -- a lie, a joke, a white lie, sarcasm, a misunderstanding -- and asks a forced-choice question about a character's mental state or intent, such as why they said what they said. It adapts a clinical psychology battery originally used to test theory-of-mind (ToM) impairment in autism and other conditions, where ToM is the ability to infer others' unobservable beliefs, desires and intentions. The task targets social/emotional reasoning that typically develops in children from about age 4, rather than factual recall or symbolic manipulation.
Zero-shot forced choice per item: multiple choice among several answer options (multiple_choice subtask) or a boolean true/false judgment (boolean subtask).
No model card in ModelSpec reports this benchmark yet.