A 44-item BIG-bench multiple-choice slice of CRASS: what would have happened if a simple event had gone the other way.
unassessed
| Category | reasoning |
|---|---|
| Subcategory | BIG-bench 44-item four-way slice of CRASS counterfactual conditionals |
| Page status | unknown |
| Metric | multiple_choice_grade |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 44 |
| Dataset licence | Apache-2.0 |
| Publisher | Google (BIG-bench collaboration); CRASS paper from Jörg Frohberg and Frank Binder |
crass_ai is the BIG-bench JSON slice of CRASS (counterfactual reasoning assessment). Each item states a simple event, then asks a questionized counterfactual ("What would have happened if …?"). The model picks one of four English answers. The LREC 2022 paper's full set is 274 premise-counterfactual tuples; task.json holds 44 of them. The skill is everyday counterfactual consequence, not COPA cause/effect and not BBH causal_judgement.
Four-option multiple choice; preferred metric multiple_choice_grade. append_choices_to_input is true. Canary GUID embedded.
No model card in ModelSpec reports this benchmark yet.