EvoEval

EvoEval evaluates code generation on evolved HumanEval-style problems across difficult, creative, subtle, combined, and tool-use domains.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorycoding
Metricpass rate
Directionhigher_is_better
Unitpercent
PublisherEvoEval authors

What it measures

EvoEval measures program synthesis beyond familiar benchmark items by transforming existing coding problems into targeted variants. Its suites test difficult reasoning, creative solutions, subtle edge cases, composition, and use of provided tools.

Task format

Code generation from natural-language problem statements, scored by execution tests.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub