Com2Sense

A BIG-bench programmatic slice of Com2Sense: complementary true/false sentences scored with pairwise accuracy.

Also known as: COM2SENSE, BIG-bench com2sense

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryreasoning
SubcategoryBIG-bench complementary true/false commonsense (programmatic subsample)
Page statusunknown
Metricpair-wise-accuracy
Directionhigher_is_better
Dataset size1874
Dataset licenceApache-2.0
PublisherPlusLabNLP; Google (BIG-bench collaboration)

What it measures

com2sense asks whether a short English sentence is commonsensical (true) or not (false), then requires the complementary partner sentence to be judged correctly as well. Items cover physical, social, and temporal knowledge under causal or comparative (and optional numeracy) frames. The BIG-bench task is a programmatic subsample of the ACL Findings 2021 Com2Sense test set, not the full 3,985-pair release.

Task format

Programmatic true/false (Yes/No when use_question_prefix is true, the default). Preferred metric pair-wise-accuracy; standard-accuracy is also reported. Default evaluate_model uses max_examples=30 (15 pairs) and seed 42. Canary GUID embedded.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub