A 32-item BIG-bench task that asks which statement most strengthens or weakens a short argument, in five-option GRE style.
unassessed
| Category | reasoning |
|---|---|
| Subcategory | BIG-bench GRE-style argument strengthen/weaken multiple-choice (32 items) |
| Page status | unknown |
| Metric | multiple_choice_grade |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 32 |
| Dataset licence | Apache-2.0 |
| Publisher | Google (BIG-bench collaboration) |
logical_args gives a model a short argumentative passage and a question about that argument's structure -- typically which extra fact would most weaken or strengthen the conclusion, or which assumption the argument makes -- and five answer options. The author wrote the items for this task in the style of GRE logical-reasoning questions. The stated skills are inferring implicit warrants, telling main claims from support, and judging strength of evidence, rather than retrieving a memorised fact.
Five-option multiple-choice, scored with BIG-bench `multiple_choice_grade`. The task.json `task_prefix` asks the model to "answer the following questions about the structure of logical arguments." A canary GUID is embedded.
No model card in ModelSpec reports this benchmark yet.