A BIG-bench programmatic slice of Com2Sense: complementary true/false sentences scored with pairwise accuracy.
unassessed
| Category | reasoning |
|---|---|
| Subcategory | BIG-bench complementary true/false commonsense (programmatic subsample) |
| Page status | unknown |
| Metric | pair-wise-accuracy |
| Direction | higher_is_better |
| Dataset size | 1874 |
| Dataset licence | Apache-2.0 |
| Publisher | PlusLabNLP; Google (BIG-bench collaboration) |
com2sense asks whether a short English sentence is commonsensical (true) or not (false), then requires the complementary partner sentence to be judged correctly as well. Items cover physical, social, and temporal knowledge under causal or comparative (and optional numeracy) frames. The BIG-bench task is a programmatic subsample of the ACL Findings 2021 Com2Sense test set, not the full 3,985-pair release.
Programmatic true/false (Yes/No when use_question_prefix is true, the default). Preferred metric pair-wise-accuracy; standard-accuracy is also reported. Default evaluate_model uses max_examples=30 (15 pairs) and seed 42. Canary GUID embedded.
No model card in ModelSpec reports this benchmark yet.