PIQA

Binary-choice physical commonsense reasoning built from instructables.com how-to text; a 2019 benchmark now close to its human baseline for most current models.

Also known as: Physical Interaction QA, Physical IQA

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryreasoning
Subcategoryphysical commonsense reasoning
Page statusactive
Metricaccuracy
Directionhigher_is_better
Unit%
Dataset size21000

What it measures

PIQA tests physical commonsense reasoning: given a goal stated in a short sentence, such as how to make a hole in a piece of wood, and two candidate solutions, a model must pick the more physically sensible one. The dataset was built from instructables.com, a site of how-to instructions for building, cooking and everyday physical tasks, so the knowledge required concerns how everyday materials and actions behave in the physical world rather than facts a model could simply recite. Solutions were engineered to require choosing between a typical, correct-seeming approach and an atypical or physically implausible one, rather than between an obviously right and an obviously wrong answer.

Task format

Given a goal sentence and two candidate solutions, the model selects the more physically appropriate solution; exactly one of the two is correct.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub