{
 "body": "\n## What it measures\n\nTerminal-Bench 2.0 evaluates agents completing software, science and system tasks in Docker environments. An agent receives an instruction and works inside a terminal environment. The task may involve programming, debugging, system administration, data processing or scientific computation, depending on the release.\n\nSuccess requires both useful interaction and a completed artifact. The benchmark measures agent planning, tool use, environment handling and end-to-end task completion. It is not a static coding-question exam.\n\n## How it is scored\n\nTerminal-Bench-style releases verify each task with task-specific tests or an oracle solution. A run normally reports the fraction of tasks whose tests pass. The source documentation does not establish a random or human baseline.\n\nResults depend on the agent adapter, model, container image, time limit, concurrency, network and exact dataset version. A score from Harbor on one release should not be compared with another without recording those settings. LILT reports pass rate across multilingual tasks, while versioned Harbor releases use executable task verification.\n\n## Dataset and licence\n\nThe opened lead does not identify a release manifest or task count. Tasks include instructions, environments and tests; some releases also mirror solutions or answer keys. The dataset card states Apache-2.0. Task-level source licences may still differ.\n\nVersion and mirror provenance matter. Hugging Face cards for Terminal-Bench 2.0 and 3.0 identify themselves as mirrors and point readers to source repositories or Harbor Hub.\n\n## Who publishes it\n\nTerminal-Bench Team is the credited publisher or author information in the opened source. The Harbor and Terminal-Bench projects maintain the execution framework and versioned task releases. No universal leaderboard was established for every assigned version.\n\n## Lineage\n\nTerminal-Bench is a versioned family of terminal-agent evaluations. Terminal-Bench 2.0, its verified derivative, Terminal-Bench 3.0 and Terminal-Bench-Science are distinct releases or task collections, not interchangeable scores. Terminal-Bench-LILT is a multilingual coding extension with a separate paper and task suite.\n\nThe v2.1 and v4.0 leads supplied for this assignment resolve only to an Artificial Analysis logo asset. They do not establish benchmark identities, releases or aliases.\n\n## Saturation and contamination\n\nNo current saturation ceiling was established. Public task instructions, tests, mirrors and, for some releases, solutions create contamination risk. A result should state whether answer keys were available to the evaluated agent and whether network access was enabled.\n\nThe science mirror states that its preview release is private upstream because tests and solutions are included. That distinction matters when interpreting scores.\n\n## How to run it\n\nUse Harbor or the Terminal-Bench CLI with the exact dataset name and version. Terminal-Bench 2.0 uses Harbor dataset terminal-bench@2.0; Terminal-Bench 3.0 documentation identifies version 3.0.0; Terminal-Bench-Science identifies v0.1.0. Run the matching release and record agent, model, container provider, time limit, concurrency and network policy.\n\n## Reading the numbers\n\nA high pass rate means the agent completed and passed selected terminal tasks under a particular environment. It does not isolate model reasoning from tools, container images, tests, timeouts or agent scaffolding. Versioned releases can change task difficulty and infrastructure. Compare only runs with the same release and execution protocol.\n\n",
 "build": {
  "built_at": "2026-09-09T16:56:50+00:00",
  "commit": "0a599558854c0e238c03a0f0d725239cb28f9d11",
  "eligibility_as_of": "2026-09-09"
 },
 "disposition": {
  "canonical_id": "terminal_bench_2_0",
  "reasons": [],
  "status": "unassessed",
  "verified_results": []
 },
 "models_covered": [],
 "page": {
  "aliases": [],
  "category": "agentic",
  "contamination": {
   "note": "Task instructions and, in several releases, tests or solutions are public or mirrored; no rotating private test policy was established.",
   "risk": "medium"
  },
  "dataset": {
   "languages": [
    "en"
   ],
   "license": "Apache-2.0",
   "modalities": [
    "text"
   ],
   "public_test_set": true,
   "size": null,
   "size_note": "A stable task count was not established from the opened primary source.",
   "splits": "versioned task collection; no train/test split established",
   "url": "https://huggingface.co/datasets/harborframework/terminal-bench-2.0"
  },
  "freshness": {
   "researched": "2026-09-09",
   "researched_by": "GPT-5.6 Luna, luna-stream-c-002 (Codex coordinated)",
   "reviewed": "",
   "reviewed_by": ""
  },
  "harness": {
   "bigbench": "",
   "helm": "",
   "inspect_evals": "",
   "lm_eval": "",
   "opencompass": "",
   "other": "Harbor/Terminal-Bench execution harness"
  },
  "id": "terminal_bench_2_0",
  "last_updated": "",
  "leaderboard_url": "",
  "lineage": {
   "family": "terminal_bench",
   "predecessor": "",
   "successors": [],
   "variants": []
  },
  "measures": "Terminal-Bench 2.0 evaluates agents completing software, science and system tasks in Docker environments. The benchmark gives an agent an English task, a terminal environment and a verification procedure; success depends on completing the task and passing its tests.\n",
  "metric": {
   "baseline_note": "The opened primary documentation does not publish a random or human baseline.",
   "direction": "higher_is_better",
   "human_baseline": null,
   "max_score": 100,
   "name": "task success verified by tests",
   "random_baseline": null,
   "unit": "%"
  },
  "name": "Terminal-Bench 2.0",
  "page_kind": "benchmark",
  "paper": {
   "arxiv": "",
   "title": "Terminal-Bench",
   "url": "https://huggingface.co/datasets/harborframework/terminal-bench-2.0",
   "year": 2025
  },
  "publisher": {
   "authors": [
    "Terminal-Bench Team"
   ],
   "org": "Terminal-Bench Team",
   "url": "https://github.com/harbor-framework/terminal-bench-2"
  },
  "released": "2025",
  "repo_url": "https://github.com/harbor-framework/terminal-bench-2",
  "saturation": {
   "as_of": "",
   "note": "No current saturation result was established from the opened primary source.",
   "status": "unknown",
   "top_score": null
  },
  "sources": [
   {
    "accessed": "2026-09-09",
    "title": "Terminal-Bench 2.0 primary dataset or paper source",
    "url": "https://huggingface.co/datasets/harborframework/terminal-bench-2.0"
   },
   {
    "accessed": "2026-09-09",
    "title": "Terminal-Bench 2.0 source repository",
    "url": "https://github.com/harbor-framework/terminal-bench-2"
   }
  ],
  "status": "unknown",
  "subcategory": "Agent execution on real terminal tasks in containers",
  "summary": "Terminal-Bench 2.0 evaluates agents completing software, science and system tasks in Docker environments.",
  "tags": [
   "terminal",
   "agents",
   "tool-use",
   "container"
  ],
  "task_format": "Agent interacts with a containerized terminal; task-specific tests determine success."
 }
}