{
 "body": "\n## What it measures\n\nTerminal-Bench measures whether an AI agent can complete real tasks inside a live terminal: given a\nnatural-language instruction and a sandboxed shell (a Docker container, sometimes several linked ones), it\nmust issue commands to get the job done, the way a person working at a command line would. The 80 tasks in\nthe launch dataset, Terminal-Bench-Core-v0, covered scientific workflows, network configuration, games,\ndata analysis, calling APIs and fixing security vulnerabilities \u2014 deliberately broader than\nsoftware-engineering-only benchmarks such as SWE-bench, and closer to general computer use.\n\n## How it is scored\n\nAn agent gets one attempt per task (accuracy is averaged across repeated runs for some agents, with error\nbars). After the agent finishes or times out, an automated test script checks the resulting container\nstate and the task counts as passed or failed, with no partial credit described. Each task also ships a\nhuman-verified reference (\"oracle\") solution used to confirm the task is actually solvable. The harness\nsupports three ways of wiring in an agent \u2014 installing it directly in the task container, integrating it\nthrough a Python interface, or connecting it over an MCP server that exposes a tmux session \u2014 and which\nmethod a given agent uses can itself affect what it is able to do in a task with an unusual or broken\nenvironment.\n\n## Dataset and licence\n\nTerminal-Bench-Core-v0 launched with 80 hand-crafted, human-verified tasks, each with its own Dockerfile\nor compose file, task specification, reference solution and test script, released under an Apache-2.0\nlicence. The announcement stated an intent to add \"several hundred\" more tasks in the following months,\nand the dataset kept growing and being re-versioned in the releases that followed (see Lineage).\n\n## Who publishes it\n\nTerminal-Bench was written by Mike Merrill, Alex Shaw, Chris Rytting, Ludwig Schmidt and Andy Konwinski,\nand released under the Laude Institute's GitHub organisation; the PyPI package's release history shows the\nfirst version published 2025-05-19, matching the project's own later reference to a \"launch in May.\" The\nproject has since moved to be hosted under the Harbor Framework organisation, alongside Harbor, a\ncompanion package for cloud-scale agent evaluation and training.\n\n## Lineage\n\nThis page covers Terminal-Bench-Core-v0, the original 80-task release (\"Terminal-Bench 1.0\"). It was\ndirectly superseded by Terminal-Bench 2.0 (`terminal_bench_2`), described by its authors as \"a harder,\nbetter verified version\" built to fix quality problems the community found in 1.0's tasks. The project\ncontinued past 2.0 with further major releases (v3.0.0 in July 2026 and v4.0.0 in August 2026, tagged in\nthe main terminal-bench repository), which appear to merge in a separate, harder professional-task set and\nare not catalogued in this repository as of this writing.\n\n## Saturation and contamination\n\nThe launch announcement showed a resolution-rate chart across agent-model combinations on\nTerminal-Bench-Core-v0 but did not print the underlying numbers in a form this research could recover, and\nthe project's live leaderboard now defaults to a much later major version rather than v0.1.1, so neither an\noriginal nor a current top score for this specific version could be confirmed here. Contamination risk is\nlower than SWE-bench's, since tasks are hand-crafted rather than mined from a pre-existing public corpus,\nbut every task's instructions and human-verified solution become public on release, and the benchmark\nexplicitly invites open-source contribution, so solved instances could plausibly reach later training\ndata; no formal publisher statement on contamination was found for this version.\n\n## How to run it\n\nThe reference implementation is the `terminal-bench` pip package (CLI command `tb`), pointed at a dataset\nname and version, for example `tb run --agent terminus --dataset-name terminal-bench-core\n--dataset-version 0.1.1`. Terminus, the project's own reference agent, is deliberately restricted to a\nsingle tool \u2014 an interactive tmux session \u2014 so that comparisons across models are not biased by a richer,\nmodel-specific toolset; third-party agents such as Claude Code, Codex CLI and Goose can also be run\nthrough the same harness, and which one is used affects both raw scores and how fairly they compare.\n\n## Reading the numbers\n\nA Terminal-Bench-Core-v0 score describes an agent's ability to plan and execute multi-step, tool-using\nwork in a real shell across a wide variety of domains, not just coding \u2014 a broader claim than a SWE-bench\nscore. Because the underlying dataset, harness and hosting have all changed substantially in the versions\nsince (2.0, 2.1, 3.0, 4.0), a \"Terminal-Bench\" number without a version attached should not be assumed to\nbe this original release; check which dataset version and which agent scaffold (installed,\nPython-integrated, MCP, or the neutral Terminus baseline) produced it before comparing across sources.\n",
 "build": {
  "built_at": "2026-09-09T16:56:50+00:00",
  "commit": "0a599558854c0e238c03a0f0d725239cb28f9d11",
  "eligibility_as_of": "2026-09-09"
 },
 "disposition": {
  "canonical_id": "terminal_bench",
  "reasons": [],
  "status": "unassessed",
  "verified_results": []
 },
 "models_covered": [
  {
   "as_of": "2026-04",
   "attribution": "unverified-legacy",
   "display_name": "o3",
   "model_id": "openai/o3",
   "provider": "openai",
   "provider_display": "OpenAI",
   "score": 58.2,
   "source": "lmarena.ai, provider-reports, domain-evals"
  },
  {
   "as_of": "2026-04",
   "attribution": "unverified-legacy",
   "display_name": "o3-deep-research",
   "model_id": "openai/o3-deep-research",
   "provider": "openai",
   "provider_display": "OpenAI",
   "score": 58.2,
   "source": "lmarena.ai, provider-reports"
  },
  {
   "as_of": "2026-04",
   "attribution": "unverified-legacy",
   "display_name": "o3-mini",
   "model_id": "openai/o3-mini",
   "provider": "openai",
   "provider_display": "OpenAI",
   "score": 58.2,
   "source": "lmarena.ai, provider-reports, llm-stats, domain-evals"
  },
  {
   "as_of": "2026-04",
   "attribution": "unverified-legacy",
   "display_name": "o3-pro",
   "model_id": "openai/o3-pro",
   "provider": "openai",
   "provider_display": "OpenAI",
   "score": 58.2,
   "source": "lmarena.ai, provider-reports"
  },
  {
   "as_of": "2026-04",
   "attribution": "unverified-legacy",
   "display_name": "Claude Opus 4",
   "model_id": "anthropic/claude-opus-4-20250514",
   "provider": "anthropic",
   "provider_display": "Anthropic",
   "score": 55.8,
   "source": "lmarena.ai, provider-reports, multimodal-evals, safety-evals, preference-evals, domain-evals"
  },
  {
   "as_of": "2026-04",
   "attribution": "unverified-legacy",
   "display_name": "Claude Opus 4.6",
   "model_id": "anthropic/claude-opus-4-6",
   "provider": "anthropic",
   "provider_display": "Anthropic",
   "score": 55.8,
   "source": "lmarena.ai, provider-reports, multimodal-evals, safety-evals, preference-evals, domain-evals, anthropic-system-card-mythos"
  },
  {
   "as_of": "2026-04",
   "attribution": "unverified-legacy",
   "display_name": "Claude Opus 4.1",
   "model_id": "anthropic/claude-opus-4-1-20250805",
   "provider": "anthropic",
   "provider_display": "Anthropic",
   "score": 52.1,
   "source": "lmarena.ai, provider-reports"
  },
  {
   "as_of": "2026-04",
   "attribution": "unverified-legacy",
   "display_name": "Claude Opus 4.1 (latest)",
   "model_id": "anthropic/claude-opus-4-1",
   "provider": "anthropic",
   "provider_display": "Anthropic",
   "score": 52.1,
   "source": "lmarena.ai, provider-reports"
  },
  {
   "as_of": "2026-04",
   "attribution": "unverified-legacy",
   "display_name": "GPT-5",
   "model_id": "openai/gpt-5",
   "provider": "openai",
   "provider_display": "OpenAI",
   "score": 48.8,
   "source": "lmarena.ai, provider-reports, domain-evals"
  },
  {
   "as_of": "2026-04",
   "attribution": "unverified-legacy",
   "display_name": "GPT-5 Chat (latest)",
   "model_id": "openai/gpt-5-chat-latest",
   "provider": "openai",
   "provider_display": "OpenAI",
   "score": 48.8,
   "source": "lmarena.ai, provider-reports"
  },
  {
   "as_of": "2026-04",
   "attribution": "unverified-legacy",
   "display_name": "GPT-5 Mini",
   "model_id": "openai/gpt-5-mini",
   "provider": "openai",
   "provider_display": "OpenAI",
   "score": 48.8,
   "source": "lmarena.ai, provider-reports"
  },
  {
   "as_of": "2026-04",
   "attribution": "unverified-legacy",
   "display_name": "GPT-5 Nano",
   "model_id": "openai/gpt-5-nano",
   "provider": "openai",
   "provider_display": "OpenAI",
   "score": 48.8,
   "source": "lmarena.ai, provider-reports"
  },
  {
   "as_of": "2026-04",
   "attribution": "unverified-legacy",
   "display_name": "GPT-5 Pro",
   "model_id": "openai/gpt-5-pro",
   "provider": "openai",
   "provider_display": "OpenAI",
   "score": 48.8,
   "source": "lmarena.ai, provider-reports"
  },
  {
   "as_of": "2026-04",
   "attribution": "unverified-legacy",
   "display_name": "GPT-5-Codex",
   "model_id": "openai/gpt-5-codex",
   "provider": "openai",
   "provider_display": "OpenAI",
   "score": 48.8,
   "source": "lmarena.ai, provider-reports"
  },
  {
   "as_of": "2026-04",
   "attribution": "unverified-legacy",
   "display_name": "GPT-5.1",
   "model_id": "openai/gpt-5-1",
   "provider": "openai",
   "provider_display": "OpenAI",
   "score": 48.8,
   "source": "lmarena.ai, provider-reports, domain-evals"
  },
  {
   "as_of": "2026-04",
   "attribution": "unverified-legacy",
   "display_name": "GPT-5.1 Chat",
   "model_id": "openai/gpt-5-1-chat-latest",
   "provider": "openai",
   "provider_display": "OpenAI",
   "score": 48.8,
   "source": "lmarena.ai, provider-reports"
  },
  {
   "as_of": "2026-04",
   "attribution": "unverified-legacy",
   "display_name": "GPT-5.1 Codex",
   "model_id": "openai/gpt-5-1-codex",
   "provider": "openai",
   "provider_display": "OpenAI",
   "score": 48.8,
   "source": "lmarena.ai, provider-reports"
  },
  {
   "as_of": "2026-04",
   "attribution": "unverified-legacy",
   "display_name": "GPT-5.1 Codex Max",
   "model_id": "openai/gpt-5-1-codex-max",
   "provider": "openai",
   "provider_display": "OpenAI",
   "score": 48.8,
   "source": "lmarena.ai, provider-reports"
  },
  {
   "as_of": "2026-04",
   "attribution": "unverified-legacy",
   "display_name": "GPT-5.1 Codex mini",
   "model_id": "openai/gpt-5-1-codex-mini",
   "provider": "openai",
   "provider_display": "OpenAI",
   "score": 48.8,
   "source": "lmarena.ai, provider-reports"
  },
  {
   "as_of": "2026-04",
   "attribution": "unverified-legacy",
   "display_name": "GPT-5.2",
   "model_id": "openai/gpt-5-2",
   "provider": "openai",
   "provider_display": "OpenAI",
   "score": 48.8,
   "source": "lmarena.ai, provider-reports, domain-evals"
  },
  {
   "as_of": "2026-04",
   "attribution": "unverified-legacy",
   "display_name": "GPT-5.2 Chat",
   "model_id": "openai/gpt-5-2-chat-latest",
   "provider": "openai",
   "provider_display": "OpenAI",
   "score": 48.8,
   "source": "lmarena.ai, provider-reports"
  },
  {
   "as_of": "2026-04",
   "attribution": "unverified-legacy",
   "display_name": "GPT-5.2 Codex",
   "model_id": "openai/gpt-5-2-codex",
   "provider": "openai",
   "provider_display": "OpenAI",
   "score": 48.8,
   "source": "lmarena.ai, provider-reports"
  },
  {
   "as_of": "2026-04",
   "attribution": "unverified-legacy",
   "display_name": "GPT-5.2 Pro",
   "model_id": "openai/gpt-5-2-pro",
   "provider": "openai",
   "provider_display": "OpenAI",
   "score": 48.8,
   "source": "lmarena.ai, provider-reports"
  },
  {
   "as_of": "2026-04",
   "attribution": "unverified-legacy",
   "display_name": "GPT-5.3 Chat (latest)",
   "model_id": "openai/gpt-5-3-chat-latest",
   "provider": "openai",
   "provider_display": "OpenAI",
   "score": 48.8,
   "source": "lmarena.ai, provider-reports"
  },
  {
   "as_of": "2026-04",
   "attribution": "unverified-legacy",
   "display_name": "GPT-5.3 Codex",
   "model_id": "openai/gpt-5-3-codex",
   "provider": "openai",
   "provider_display": "OpenAI",
   "score": 48.8,
   "source": "lmarena.ai, provider-reports"
  },
  {
   "as_of": "2026-04",
   "attribution": "unverified-legacy",
   "display_name": "GPT-5.3 Codex Spark",
   "model_id": "openai/gpt-5-3-codex-spark",
   "provider": "openai",
   "provider_display": "OpenAI",
   "score": 48.8,
   "source": "lmarena.ai, provider-reports"
  },
  {
   "as_of": "2026-04",
   "attribution": "unverified-legacy",
   "display_name": "GPT-5.4",
   "model_id": "openai/gpt-5-4",
   "provider": "openai",
   "provider_display": "OpenAI",
   "score": 48.8,
   "source": "lmarena.ai, provider-reports, anthropic-system-card-mythos, domain-evals"
  },
  {
   "as_of": "2026-04",
   "attribution": "unverified-legacy",
   "display_name": "GPT-5.4 mini",
   "model_id": "openai/gpt-5-4-mini",
   "provider": "openai",
   "provider_display": "OpenAI",
   "score": 48.8,
   "source": "lmarena.ai, provider-reports"
  },
  {
   "as_of": "2026-04",
   "attribution": "unverified-legacy",
   "display_name": "GPT-5.4 nano",
   "model_id": "openai/gpt-5-4-nano",
   "provider": "openai",
   "provider_display": "OpenAI",
   "score": 48.8,
   "source": "lmarena.ai, provider-reports"
  },
  {
   "as_of": "2026-04",
   "attribution": "unverified-legacy",
   "display_name": "GPT-5.4 Pro",
   "model_id": "openai/gpt-5-4-pro",
   "provider": "openai",
   "provider_display": "OpenAI",
   "score": 48.8,
   "source": "lmarena.ai, provider-reports"
  },
  {
   "as_of": "2026-04",
   "attribution": "unverified-legacy",
   "display_name": "Claude Sonnet 4",
   "model_id": "anthropic/claude-sonnet-4-20250514",
   "provider": "anthropic",
   "provider_display": "Anthropic",
   "score": 48.2,
   "source": "lmarena.ai, provider-reports, multimodal-evals, safety-evals, preference-evals"
  },
  {
   "as_of": "2026-04",
   "attribution": "unverified-legacy",
   "display_name": "Claude Sonnet 4.5",
   "model_id": "anthropic/claude-sonnet-4-5-20250929",
   "provider": "anthropic",
   "provider_display": "Anthropic",
   "score": 48.2,
   "source": "lmarena.ai, provider-reports, multimodal-evals, safety-evals, preference-evals"
  },
  {
   "as_of": "2026-04",
   "attribution": "unverified-legacy",
   "display_name": "Claude Sonnet 4.5 (latest)",
   "model_id": "anthropic/claude-sonnet-4-5",
   "provider": "anthropic",
   "provider_display": "Anthropic",
   "score": 48.2,
   "source": "lmarena.ai, provider-reports, multimodal-evals, safety-evals, preference-evals"
  },
  {
   "as_of": "2026-04",
   "attribution": "unverified-legacy",
   "display_name": "Gemini 2.5 Pro",
   "model_id": "google/gemini-2-5-pro",
   "provider": "google",
   "provider_display": "Google DeepMind",
   "score": 45.5,
   "source": "lmarena.ai, provider-reports, multimodal-evals, safety-evals, preference-evals, domain-evals, llm-stats, intlpull"
  },
  {
   "as_of": "2026-04",
   "attribution": "unverified-legacy",
   "display_name": "GPT-4.1",
   "model_id": "openai/gpt-4-1",
   "provider": "openai",
   "provider_display": "OpenAI",
   "score": 42.1,
   "source": "lmarena.ai, provider-reports, multimodal-evals, safety-evals, domain-evals preference-evals, llm-stats, intlpull"
  },
  {
   "as_of": "2026-04",
   "attribution": "unverified-legacy",
   "display_name": "Claude Haiku 4.5",
   "model_id": "anthropic/claude-haiku-4-5-20251001",
   "provider": "anthropic",
   "provider_display": "Anthropic",
   "score": 41.0,
   "source": "lmarena.ai, provider-reports, domain-evals"
  },
  {
   "as_of": "2026-04",
   "attribution": "unverified-legacy",
   "display_name": "Claude Haiku 4.5 (latest)",
   "model_id": "anthropic/claude-haiku-4-5",
   "provider": "anthropic",
   "provider_display": "Anthropic",
   "score": 41.0,
   "source": "lmarena.ai, provider-reports, domain-evals"
  },
  {
   "as_of": "2026-04",
   "attribution": "unverified-legacy",
   "display_name": "Claude Opus 4 (latest)",
   "model_id": "anthropic/claude-opus-4-0",
   "provider": "anthropic",
   "provider_display": "Anthropic",
   "score": 39.2,
   "source": "lmarena.ai, provider-reports"
  },
  {
   "as_of": "2026-04",
   "attribution": "unverified-legacy",
   "display_name": "DeepSeek Chat",
   "model_id": "deepseek/deepseek-chat",
   "provider": "deepseek",
   "provider_display": "DeepSeek",
   "score": 37.7,
   "source": "lmarena.ai, provider-reports, safety-evals, preference-evals, open-llm-leaderboard-v2, llm-stats"
  },
  {
   "as_of": "2026-04",
   "attribution": "unverified-legacy",
   "display_name": "DeepSeek V3",
   "model_id": "deepseek/deepseek-v3",
   "provider": "deepseek",
   "provider_display": "DeepSeek",
   "score": 37.7,
   "source": "lmarena.ai, provider-reports, safety-evals, preference-evals, domain-evals open-llm-leaderboard-v2, llm-stats"
  },
  {
   "as_of": "2026-04",
   "attribution": "unverified-legacy",
   "display_name": "DeepSeek V3.2",
   "model_id": "deepseek/deepseek-v3-2",
   "provider": "deepseek",
   "provider_display": "DeepSeek",
   "score": 37.7,
   "source": "lmarena.ai, provider-reports, safety-evals, preference-evals, open-llm-leaderboard-v2, llm-stats"
  },
  {
   "as_of": "2026-04",
   "attribution": "unverified-legacy",
   "display_name": "DeepSeek V3.2 Exp",
   "model_id": "deepseek/deepseek-v3-2-exp",
   "provider": "deepseek",
   "provider_display": "DeepSeek",
   "score": 37.7,
   "source": "lmarena.ai, provider-reports, safety-evals, preference-evals, open-llm-leaderboard-v2, llm-stats"
  },
  {
   "as_of": "2026-04",
   "attribution": "unverified-legacy",
   "display_name": "Claude Sonnet 4 (latest)",
   "model_id": "anthropic/claude-sonnet-4-0",
   "provider": "anthropic",
   "provider_display": "Anthropic",
   "score": 35.5,
   "source": "lmarena.ai, provider-reports"
  },
  {
   "as_of": "2026-04",
   "attribution": "unverified-legacy",
   "display_name": "Claude Sonnet 3.7",
   "model_id": "anthropic/claude-3-7-sonnet-20250219",
   "provider": "anthropic",
   "provider_display": "Anthropic",
   "score": 35.2,
   "source": "lmarena.ai, provider-reports, domain-evals"
  },
  {
   "as_of": "2026-04",
   "attribution": "unverified-legacy",
   "display_name": "GPT-4o",
   "model_id": "openai/gpt-4o",
   "provider": "openai",
   "provider_display": "OpenAI",
   "score": 35.2,
   "source": "lmarena.ai, provider-reports, multimodal-evals, safety-evals, preference-evals, domain-evals, llm-stats, intlpull"
  },
  {
   "as_of": "2026-04",
   "attribution": "unverified-legacy",
   "display_name": "GPT-4o (2024-05-13)",
   "model_id": "openai/gpt-4o-2024-05-13",
   "provider": "openai",
   "provider_display": "OpenAI",
   "score": 35.2,
   "source": "lmarena.ai, provider-reports, multimodal-evals, safety-evals, preference-evals, domain-evals, llm-stats, intlpull"
  },
  {
   "as_of": "2026-04",
   "attribution": "unverified-legacy",
   "display_name": "GPT-4o (2024-08-06)",
   "model_id": "openai/gpt-4o-2024-08-06",
   "provider": "openai",
   "provider_display": "OpenAI",
   "score": 35.2,
   "source": "lmarena.ai, provider-reports, multimodal-evals, safety-evals, preference-evals, domain-evals, llm-stats, intlpull"
  },
  {
   "as_of": "2026-04",
   "attribution": "unverified-legacy",
   "display_name": "GPT-4o (2024-11-20)",
   "model_id": "openai/gpt-4o-2024-11-20",
   "provider": "openai",
   "provider_display": "OpenAI",
   "score": 35.2,
   "source": "lmarena.ai, provider-reports, multimodal-evals, safety-evals, preference-evals, domain-evals, llm-stats, intlpull"
  },
  {
   "as_of": "2026-04",
   "attribution": "unverified-legacy",
   "display_name": "GPT-4o mini",
   "model_id": "openai/gpt-4o-mini",
   "provider": "openai",
   "provider_display": "OpenAI",
   "score": 35.2,
   "source": "lmarena.ai, provider-reports, multimodal-evals, safety-evals, preference-evals, domain-evals, llm-stats, intlpull"
  },
  {
   "as_of": "2026-04",
   "attribution": "unverified-legacy",
   "display_name": "DeepSeek V3.1",
   "model_id": "deepseek/deepseek-v3-1",
   "provider": "deepseek",
   "provider_display": "DeepSeek",
   "score": 31.3,
   "source": "lmarena.ai, provider-reports, safety-evals, preference-evals, open-llm-leaderboard-v2, llm-stats"
  },
  {
   "as_of": "2026-04",
   "attribution": "unverified-legacy",
   "display_name": "DeepSeek R1",
   "model_id": "deepseek/deepseek-r1",
   "provider": "deepseek",
   "provider_display": "DeepSeek",
   "score": 5.7,
   "source": "lmarena.ai, provider-reports, preference-evals, open-llm-leaderboard-v2, domain-evals"
  },
  {
   "as_of": "2026-04",
   "attribution": "unverified-legacy",
   "display_name": "DeepSeek R1 0528",
   "model_id": "deepseek/deepseek-r1-0528",
   "provider": "deepseek",
   "provider_display": "DeepSeek",
   "score": 5.7,
   "source": "lmarena.ai, provider-reports, preference-evals, open-llm-leaderboard-v2, domain-evals"
  },
  {
   "as_of": "2026-04",
   "attribution": "unverified-legacy",
   "display_name": "DeepSeek R1 0528 NVFP4 v2",
   "model_id": "nvidia/deepseek-r1-0528-nvfp4-v2",
   "provider": "nvidia",
   "provider_display": "NVIDIA",
   "score": 5.7,
   "source": "lmarena.ai, provider-reports, preference-evals, open-llm-leaderboard-v2"
  }
 ],
 "page": {
  "aliases": [
   "Terminal-Bench 1.0",
   "Terminal-Bench-Core",
   "T-Bench"
  ],
  "category": "agentic",
  "contamination": {
   "note": "Tasks are hand-crafted for the benchmark rather than mined from a pre-existing public corpus (unlike SWE-bench's pull requests), which limits initial exposure. But each task's instructions, Docker environment and human-verified solution become public once released, and the benchmark actively invites open-source contribution, so solved instances could plausibly reach later training data. No formal publisher statement on contamination was found for this version.\n",
   "risk": "medium"
  },
  "dataset": {
   "languages": [
    "English"
   ],
   "license": "Apache-2.0",
   "modalities": [
    "text"
   ],
   "public_test_set": true,
   "size": 80,
   "size_note": "Terminal-Bench-Core-v0, the launch dataset, held 80 hand-crafted tasks, each with a dedicated Docker environment, a human-verified solution, and an automated test script. The announcement stated an intent to grow this to \"several hundred\" tasks over the following weeks and months; the dataset continued to expand under later versioned releases (see Lineage).\n",
   "splits": "Terminal-Bench-Core-v0 (80 tasks at launch); re-versioned afterward as the dataset grew",
   "url": "https://pypi.org/project/terminal-bench/"
  },
  "freshness": {
   "researched": "2026-09-08",
   "researched_by": "sonnet-5 agent, batch 1b, slice M",
   "reviewed": "",
   "reviewed_by": ""
  },
  "harness": {
   "bigbench": "",
   "helm": "",
   "inspect_evals": "",
   "lm_eval": "",
   "opencompass": "",
   "other": "Official Terminal-Bench harness (`tb` CLI, pip package `terminal-bench`), which also ships Terminus, a deliberately minimal reference agent restricted to a single tmux tool, built so that model comparisons are not biased by one agent's own tool design. Third-party agents (Claude Code, Codex CLI, Goose, etc.) can also be run through the harness, which affects comparability across leaderboard rows.\n"
  },
  "id": "terminal_bench",
  "last_updated": "",
  "leaderboard_url": "https://www.tbench.ai/",
  "lineage": {
   "family": "",
   "predecessor": "",
   "successors": [
    "terminal_bench_2"
   ],
   "variants": []
  },
  "measures": "Terminal-Bench gives an agent a natural-language instruction and a live, sandboxed terminal (a Docker container, sometimes several linked containers) and asks it to complete the task by issuing commands, the way a person would work at a shell. Tasks at release covered scientific workflows, network configuration, games, data analysis, calling APIs and fixing security vulnerabilities, deliberately going beyond software-engineering-only benchmarks like SWE-bench. Everything happens in text, matching how language models are trained, but the range of tools, state and multi-step planning required is broader than a single coding task.\n",
  "metric": {
   "baseline_note": "No formal human baseline published; each task instead ships a human-verified oracle solution used to confirm the task is solvable.",
   "direction": "higher_is_better",
   "human_baseline": null,
   "max_score": 100,
   "name": "task resolution rate (tasks passed / tasks attempted)",
   "random_baseline": 0,
   "unit": "%"
  },
  "name": "Terminal-Bench",
  "page_kind": "benchmark",
  "paper": {
   "arxiv": "",
   "title": "",
   "url": "",
   "year": null
  },
  "publisher": {
   "authors": [
    "Mike Merrill",
    "Alex Shaw",
    "Chris Rytting",
    "Ludwig Schmidt",
    "Andy Konwinski"
   ],
   "org": "Originally released under the Laude Institute's GitHub organisation; the project is now hosted under the Harbor Framework",
   "url": "https://www.tbench.ai/"
  },
  "released": "2025-05",
  "repo_url": "https://github.com/harbor-framework/terminal-bench",
  "saturation": {
   "as_of": "",
   "note": "The launch announcement showed a resolution-rate chart across agent-model pairs on Terminal-Bench-Core-v0 but did not print the underlying numbers in a form this research could read, and the site's live leaderboard now defaults to a much later major version rather than v0.1.1, so a current or original top score could not be confirmed here.\n",
   "status": "unknown",
   "top_score": null
  },
  "sources": [
   {
    "accessed": "2026-09-08",
    "title": "Terminal-Bench (launch announcement)",
    "url": "https://www.tbench.ai/news/announcement"
   },
   {
    "accessed": "2026-09-08",
    "title": "harbor-framework/terminal-bench repository",
    "url": "https://github.com/harbor-framework/terminal-bench"
   },
   {
    "accessed": "2026-09-08",
    "title": "terminal-bench on PyPI (release history)",
    "url": "https://pypi.org/project/terminal-bench/#history"
   },
   {
    "accessed": "2026-09-08",
    "title": "Terminus (reference agent announcement)",
    "url": "https://www.tbench.ai/news/terminus"
   }
  ],
  "status": "superseded",
  "subcategory": "general-purpose terminal / command-line agent task completion",
  "summary": "Terminal-Bench 1.0 measures whether an AI agent can complete real command-line tasks in a sandboxed Docker terminal, graded by automated tests; superseded by later major versions.",
  "tags": [
   "agentic",
   "terminal",
   "command-line",
   "docker",
   "tool-use"
  ],
  "task_format": "An agent receives a task instruction and a terminal session inside a Docker environment. The Terminal-Bench harness orchestrates the agent (installed directly in the container, integrated via a Python interface, or connected through an MCP server exposing a tmux session), logs its actions, and checks final container state against a task-specific automated test script after the agent finishes or times out. Each task also ships a human-verified reference solution.\n"
 }
}