NPHardEval tests large language models on 900 algorithmic questions spanning complexity classes through NP-hard problems.
unassessed
| Category | reasoning |
|---|---|
| Subcategory | algorithmic reasoning |
| Page status | active |
| Metric | weighted accuracy |
| Direction | higher_is_better |
| Unit | percent |
| Dataset size | 900 |
| Dataset licence | Apache-2.0 |
| Publisher | CASM Lab |
NPHardEval evaluates reasoning about graph, routing, matching, and path problems. Its official description says the questions span complexity classes below and through NP-hard problems.
Zero-shot natural-language problem with structured final answers and brief reasoning.
No model card in ModelSpec reports this benchmark yet.