Binary-choice physical commonsense reasoning built from instructables.com how-to text; a 2019 benchmark now close to its human baseline for most current models.
unassessed
| Category | reasoning |
|---|---|
| Subcategory | physical commonsense reasoning |
| Page status | active |
| Metric | accuracy |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 21000 |
PIQA tests physical commonsense reasoning: given a goal stated in a short sentence, such as how to make a hole in a piece of wood, and two candidate solutions, a model must pick the more physically sensible one. The dataset was built from instructables.com, a site of how-to instructions for building, cooking and everyday physical tasks, so the knowledge required concerns how everyday materials and actions behave in the physical world rather than facts a model could simply recite. Solutions were engineered to require choosing between a typical, correct-seeming approach and an atypical or physically implausible one, rather than between an obviously right and an obviously wrong answer.
Given a goal sentence and two candidate solutions, the model selects the more physically appropriate solution; exactly one of the two is correct.
No model card in ModelSpec reports this benchmark yet.