Strange Stories

BIG-bench task built on a clinical theory-of-mind battery that asks a model to infer characters' beliefs, intentions and non-literal meaning from short narratives.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryreasoning
Subcategorytheory of mind / social reasoning
Page statusactive
Metricmultiple choice grade
Directionhigher_is_better
Unit%
Dataset size174
Dataset licenceApache-2.0
PublisherGoogle (BIG-bench collaboration)

What it measures

Strange Stories gives the model a short naturalistic narrative in which a character says something that is not literally true -- a lie, a joke, a white lie, sarcasm, a misunderstanding -- and asks a forced-choice question about a character's mental state or intent, such as why they said what they said. It adapts a clinical psychology battery originally used to test theory-of-mind (ToM) impairment in autism and other conditions, where ToM is the ability to infer others' unobservable beliefs, desires and intentions. The task targets social/emotional reasoning that typically develops in children from about age 4, rather than factual recall or symbolic manipulation.

Task format

Zero-shot forced choice per item: multiple choice among several answer options (multiple_choice subtask) or a boolean true/false judgment (boolean subtask).

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub