EvoEval evaluates code generation on evolved HumanEval-style problems across difficult, creative, subtle, combined, and tool-use domains.
unassessed
| Category | coding |
|---|---|
| Metric | pass rate |
| Direction | higher_is_better |
| Unit | percent |
| Publisher | EvoEval authors |
EvoEval measures program synthesis beyond familiar benchmark items by transforming existing coding problems into targeted variants. Its suites test difficult reasoning, creative solutions, subtle edge cases, composition, and use of provided tools.
Code generation from natural-language problem statements, scored by execution tests.
No model card in ModelSpec reports this benchmark yet.