Evo-Bench

Evo-Bench evaluates whether language models can autonomously evolve agent harnesses across Search, Office and General domains.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryagentic
Subcategoryharness evolution
Page statusactive
Metricabsolute performance gain from harness evolution
Directionhigher_is_better
Unit%
Dataset size0
PublisherEvo-Bench authors

What it measures

Intrinsic agent-harness evolution and cross-suite gains after autonomous harness optimization.

Task format

Long-horizon agent tasks where the model can modify or improve its operating harness.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub