{
 "body": "\n## What it measures\n\nEvo-Bench targets an agent\u2019s ability to improve its own operating harness, rather than only solve static tasks. It spans Search, Office and General agent domains and uses sensitivity-aware construction and stratified splitting to isolate harness effects from base model strength.\n\n## How it is scored\n\nThe paper reports absolute gains from autonomous harness evolution and compares them with human-engineered baselines. Its abstract gives a top gain of 16.6 points but does not define the full metric formula; report the exact task score and gain protocol.\n\n## Dataset and licence\n\nThe primary source establishes the benchmark release, but the opened materials do not establish a single dataset licence. Confirm the current release terms and split accounting before redistribution.\n\n## Who publishes it\n\nEvo-Bench is described in the 2026 paper and released through the linked benchmark materials; the opened sources do not provide a complete author list.\n\n## Lineage\n\nNo predecessor or successor was established in the opened primary source.\n\n## Saturation and contamination\n\nThe benchmark materials are public, which creates contamination opportunities. The opened source does not establish a contamination audit or current saturation ceiling.\n\n## How to run it\n\nFollow the official repository or paper protocol, recording the exact model, prompts, evaluator, task version, tool access and timeout. Preserve per-task outcomes when comparing runs.\n\n## Reading the numbers\n\nHigher scores indicate more successful tasks under the selected protocol. Results can depend on evaluator models, prompts, environment setup and aggregation, so compare only matched configurations.\n\n## Protocol cautions\n\nHarness evolution can change tool calls, prompts, memory and control flow, so the starting harness and allowed modifications are part of the experimental condition. Report both the pre-evolution score and the post-evolution score, plus the policy model and search budget. A gain without those controls cannot isolate harness improvement.\n\nThe abstract also reports that synthesized harnesses can transfer across policy models and that autonomous evolution behaves differently by domain: it performs strongly in Search, while Office tasks with specialized workflows remain difficult. These findings motivate reporting domain-level scores instead of only one aggregate.\nDomain breakdowns are therefore essential for interpretation.\n",
 "build": {
  "built_at": "2026-09-09T16:56:50+00:00",
  "commit": "0a599558854c0e238c03a0f0d725239cb28f9d11",
  "eligibility_as_of": "2026-09-09"
 },
 "disposition": {
  "canonical_id": "evo_bench",
  "reasons": [],
  "status": "unassessed",
  "verified_results": []
 },
 "models_covered": [],
 "page": {
  "aliases": [],
  "category": "agentic",
  "contamination": {
   "note": "The benchmark materials are public; the opened source does not establish a contamination audit or private rotating holdout.",
   "risk": "medium"
  },
  "dataset": {
   "languages": [
    "en"
   ],
   "license": "",
   "modalities": [
    "text",
    "code"
   ],
   "public_test_set": true,
   "size": 0,
   "size_note": "The paper abstract does not state a single item count.",
   "splits": "Unknown unless specified by the official release.",
   "url": "https://huggingface.co/datasets/RUC-AIBOX/Evo-Bench"
  },
  "freshness": {
   "researched": "2026-09-09",
   "researched_by": "GPT-5.6 Luna, luna-stream-c-005 (Codex coordinated)",
   "reviewed": "",
   "reviewed_by": ""
  },
  "harness": {
   "bigbench": "",
   "helm": "",
   "inspect_evals": "",
   "lm_eval": "",
   "opencompass": "",
   "other": "Use the official release protocol and record its evaluator and prompt settings."
  },
  "id": "evo_bench",
  "last_updated": "",
  "leaderboard_url": "",
  "lineage": {
   "family": "agent evaluation",
   "predecessor": "",
   "successors": [],
   "variants": []
  },
  "measures": "Intrinsic agent-harness evolution and cross-suite gains after autonomous harness optimization.",
  "metric": {
   "baseline_note": "No universal random or human baseline was established in the opened primary source.",
   "direction": "higher_is_better",
   "human_baseline": null,
   "max_score": 100,
   "name": "absolute performance gain from harness evolution",
   "random_baseline": null,
   "unit": "%"
  },
  "name": "Evo-Bench",
  "page_kind": "benchmark",
  "paper": {
   "arxiv": "2608.09096",
   "title": "Evo-Bench: Can Language Models Improve Agent Harness?",
   "url": "https://arxiv.org/abs/2608.09096v2",
   "year": 2026
  },
  "publisher": {
   "authors": [],
   "org": "Evo-Bench authors",
   "url": "https://huggingface.co/datasets/RUC-AIBOX/Evo-Bench"
  },
  "released": "2026-08",
  "repo_url": "https://huggingface.co/datasets/RUC-AIBOX/Evo-Bench",
  "saturation": {
   "as_of": "",
   "note": "No current saturation ceiling was established in the opened source.",
   "status": "unknown",
   "top_score": null
  },
  "sources": [
   {
    "accessed": "2026-09-09",
    "title": "Primary paper",
    "url": "https://arxiv.org/abs/2608.09096v2"
   },
   {
    "accessed": "2026-09-09",
    "title": "Official repository",
    "url": "https://huggingface.co/datasets/RUC-AIBOX/Evo-Bench"
   }
  ],
  "status": "active",
  "subcategory": "harness evolution",
  "summary": "Evo-Bench evaluates whether language models can autonomously evolve agent harnesses across Search, Office and General domains.",
  "tags": [
   "agents",
   "harness",
   "search",
   "office"
  ],
  "task_format": "Long-horizon agent tasks where the model can modify or improve its operating harness."
 }
}