190 yes/no questions, from psychology papers, asking whether a typical person would say X caused Y in a short moral or counterfactual story.
unassessed
| Category | reasoning |
|---|---|
| Subcategory | yes/no causal attribution in short stories |
| Page status | active |
| Metric | multiple_choice_grade |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 190 |
| Dataset licence | Apache-2.0 |
| Publisher | Google (BIG-bench collaboration) |
causal_judgment gives a short story with several candidate causes and a yes/no question such as “Did X cause Y?” or “Did the CEO intentionally harm the environment?” Labels are majority human answers from the original experiments, not a formal causal graph. The point is folk causal judgment (norms, intent, counterfactuals), not COPA-style commonsense pairing.
Two-way multiple choice (Yes / No), preferred_score multiple_choice_grade. BIG-bench JSON has 190 examples (99 Yes, 91 No). BIG-bench Hard ships 187 examples as causal_judgement.json (British spelling) and usually adds chain-of-thought.
No model card in ModelSpec reports this benchmark yet.