A 60-item BIG-bench true/false task that labels author-written English sentences laced with hedges, fragments, or partial truths.
unassessed
| Category | reasoning |
|---|---|
| Subcategory | BIG-bench true/false English sentences with hedges, fragments, and partial truths (60 items) |
| Page status | unknown |
| Metric | multiple_choice_grade |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 60 |
| Dataset licence | Apache-2.0 |
| Publisher | Google (BIG-bench collaboration) |
sentence_ambiguity gives one English claim and asks True or False. Sahib Singh wrote the items so that hedges ("might", "likely"), arithmetic fragments, or half-true clauses make a hasty fact lookup fail. Gold is a single boolean. Dummy-model header: 60 multiple-choice queries and 0 free-text queries. Direct count of task.json: 60 examples (28 True, 32 False). Preferred metric multiple_choice_grade. Canary GUID embedded. Not in BIG-Bench Hard.
Two-way multiple choice (True/False). example_input_prefix "Claim: "; example_output_prefix "True or False? ". append_choices_to_input is false. JSON task, keywords common sense, json, multiple choice, reading comprehension, word sense disambiguation.
No model card in ModelSpec reports this benchmark yet.