NPHardEval

NPHardEval tests large language models on 900 algorithmic questions spanning complexity classes through NP-hard problems.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryreasoning
Subcategoryalgorithmic reasoning
Page statusactive
Metricweighted accuracy
Directionhigher_is_better
Unitpercent
Dataset size900
Dataset licenceApache-2.0
PublisherCASM Lab

What it measures

NPHardEval evaluates reasoning about graph, routing, matching, and path problems. Its official description says the questions span complexity classes below and through NP-hard problems.

Task format

Zero-shot natural-language problem with structured final answers and brief reasoning.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub