Evo-Bench evaluates whether language models can autonomously evolve agent harnesses across Search, Office and General domains.
unassessed
| Category | agentic |
|---|---|
| Subcategory | harness evolution |
| Page status | active |
| Metric | absolute performance gain from harness evolution |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 0 |
| Publisher | Evo-Bench authors |
Intrinsic agent-harness evolution and cross-suite gains after autonomous harness optimization.
Long-horizon agent tasks where the model can modify or improve its operating harness.
No model card in ModelSpec reports this benchmark yet.