18,736 four-choice English questions from 14 templates that probe physical reasoning about objects in space and time, meant to be used zero-shot.
unassessed
| Category | reasoning |
|---|---|
| Subcategory | zero-shot four-choice English physical reasoning over 10 object concepts |
| Page status | unknown |
| Metric | accuracy (acc); acc_norm also reported |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 18736 |
| Dataset licence | Apache-2.0 |
| Publisher | University of Colorado Boulder (NaLa lab) |
PROST (Physical Reasoning about Objects Through Space and Time) asks which of four everyday objects best fits a short English scene. Concepts are direction, mass, height, circumference, stackable, rollable, graspable, breakable, slideable, and bounceable. Aroca-Ouellette, Paik, Roncone, and Kann wrote 14 templates and expanded them to 18,736 items. The paper and the lm-eval README require zero-shot use. It is not [piqa](piqa.md), which is two-way how-to commonsense from Instructables.
Four-option multiple choice. lm-eval task prost loads hf://datasets/corypaik/prost/data/default.jsonl, test split. Prompt is context, then "Question: {ex_question}", then "Answer:". Choices are fields A–D. Target is the integer label. should_decontaminate true. Metrics acc and acc_norm. YAML filename corypaik_prost.yaml; runnable name is prost.
No model card in ModelSpec reports this benchmark yet.