Empirical Judgments

A 99-item BIG-bench task that labels an English sentence as asserting a causal, correlative, or merely conceptual relation.

Also known as: BIG-bench empirical_judgments

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryreasoning
SubcategoryBIG-bench three-way causal vs correlative vs neutral sentence labels (99 items)
Page statusunknown
Metricmultiple_choice_grade
Directionhigher_is_better
Unit%
Dataset size99
Dataset licenceApache-2.0
PublisherGoogle (BIG-bench collaboration)

What it measures

empirical_judgments presents one English sentence and asks whether it asserts a causal relation between events, a correlative (regularly conjoined) relation, or neither (neutral). Jennifer Marsh and Samuel Schoenholz wrote the 99 items after Kant's contrast between objectively valid "judgments of experience" and subjectively valid "judgments of perception." The first 66 items pair similar event descriptions in causal versus correlative wording. The last 33 copy if/whenever syntax but relate concepts, not events. The skill is recognising those linguistic claims, not identifying causes from data and not [cause_and_effect](cause_and_effect.md).

Task format

Three-option multiple choice (causal / correlative / neutral), preferred metric multiple_choice_grade. task_prefix tells the model to respond with those three words. example_input_prefix is "Sentence:"; example_output_prefix is "Relation:". append_choices_to_input is false. Canary GUID embedded.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub