Logical Arguments

A 32-item BIG-bench task that asks which statement most strengthens or weakens a short argument, in five-option GRE style.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryreasoning
SubcategoryBIG-bench GRE-style argument strengthen/weaken multiple-choice (32 items)
Page statusunknown
Metricmultiple_choice_grade
Directionhigher_is_better
Unit%
Dataset size32
Dataset licenceApache-2.0
PublisherGoogle (BIG-bench collaboration)

What it measures

logical_args gives a model a short argumentative passage and a question about that argument's structure -- typically which extra fact would most weaken or strengthen the conclusion, or which assumption the argument makes -- and five answer options. The author wrote the items for this task in the style of GRE logical-reasoning questions. The stated skills are inferring implicit warrants, telling main claims from support, and judging strength of evidence, rather than retrieving a memorised fact.

Task format

Five-option multiple-choice, scored with BIG-bench `multiple_choice_grade`. The task.json `task_prefix` asks the model to "answer the following questions about the structure of logical arguments." A canary GUID is embedded.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub