CRASS (BIG-bench crass_ai)

A 44-item BIG-bench multiple-choice slice of CRASS: what would have happened if a simple event had gone the other way.

Also known as: CRASS, crass_ai, Counterfactual Conditionals

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryreasoning
SubcategoryBIG-bench 44-item four-way slice of CRASS counterfactual conditionals
Page statusunknown
Metricmultiple_choice_grade
Directionhigher_is_better
Unit%
Dataset size44
Dataset licenceApache-2.0
PublisherGoogle (BIG-bench collaboration); CRASS paper from Jörg Frohberg and Frank Binder

What it measures

crass_ai is the BIG-bench JSON slice of CRASS (counterfactual reasoning assessment). Each item states a simple event, then asks a questionized counterfactual ("What would have happened if …?"). The model picks one of four English answers. The LREC 2022 paper's full set is 274 premise-counterfactual tuples; task.json holds 44 of them. The skill is everyday counterfactual consequence, not COPA cause/effect and not BBH causal_judgement.

Task format

Four-option multiple choice; preferred metric multiple_choice_grade. append_choices_to_input is true. Canary GUID embedded.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub