IBM's boolean- and multiple-choice test of seven atomic reasoning skills needed for planning -- action applicability, reachability, justification, landmarks and more -- across 13 formal domains.
unassessed
| Category | reasoning |
|---|---|
| Subcategory | planning and action reasoning |
| Page status | active |
| Metric | accuracy |
| Direction | higher_is_better |
| Unit | % |
| Dataset licence | CDLA-Permissive-2.0 |
| Publisher | IBM Research |
ACPBench gives a model a natural-language description of a planning problem (translated from a formal PDDL domain and problem definition, covering domains such as Blocksworld, Logistics, Grippers and Alfworld) and asks a targeted question about one of seven atomic reasoning skills: whether an action is applicable in the current state, what a state looks like after an action (progression), whether a goal atom is reachable, whether an action sequence validly achieves a goal, whether an action could ever become applicable on some future path (action reachability), whether an action in a plan is unjustified and can be dropped, or which facts every valid plan must pass through (landmarks). It isolates the specific reasoning sub-skills end-to-end plan generation depends on, rather than asking a model to produce a full plan itself.
Boolean (yes/no) or four-option multiple-choice questions over a natural-language-translated planning problem; the base ACPBench tasks require no free-text generation (a separate, harder generative version, ACPBench-Hard, exists -- see Lineage).
No model card in ModelSpec reports this benchmark yet.