Causal Judgment (BIG-bench)

190 yes/no questions, from psychology papers, asking whether a typical person would say X caused Y in a short moral or counterfactual story.

Also known as: causal_judgement, BBH causal_judgement

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryreasoning
Subcategoryyes/no causal attribution in short stories
Page statusactive
Metricmultiple_choice_grade
Directionhigher_is_better
Unit%
Dataset size190
Dataset licenceApache-2.0
PublisherGoogle (BIG-bench collaboration)

What it measures

causal_judgment gives a short story with several candidate causes and a yes/no question such as “Did X cause Y?” or “Did the CEO intentionally harm the environment?” Labels are majority human answers from the original experiments, not a formal causal graph. The point is folk causal judgment (norms, intent, counterfactuals), not COPA-style commonsense pairing.

Task format

Two-way multiple choice (Yes / No), preferred_score multiple_choice_grade. BIG-bench JSON has 190 examples (99 Yes, 91 No). BIG-bench Hard ships 187 examples as causal_judgement.json (British spelling) and usually adds chain-of-thought.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub