A harder, more heavily verified 89-task remake of Terminal-Bench, released with the Harbor evaluation package; Terminal-Bench 2.1 later patched 28 of its tasks.
unassessed
| Category | agentic |
|---|---|
| Subcategory | general-purpose terminal / command-line agent task completion, verified task set |
| Page status | active |
| Metric | task resolution rate (tasks passed / tasks attempted) |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 89 |
| Dataset licence | Apache-2.0 |
| Publisher | The Terminal-Bench Team, hosted under the Harbor Framework project |
Terminal-Bench 2.0 measures the same thing as the original Terminal-Bench: whether an AI agent can carry out a real task by issuing commands in a live, sandboxed terminal. The task domains and text-only interface are unchanged; what changed is quality control. The authors write that they "weren't satisfied with the level of verification" in the original dataset — for example, a task that scraped YouTube broke whenever YouTube's anti-bot defences changed — so 2.0 puts each task through substantial manual and LM-assisted review before inclusion.
Unchanged from the original Terminal-Bench: an agent receives a task instruction and a sandboxed Docker terminal, and must complete the task through shell commands. Terminal-Bench 2.0 shipped alongside Harbor, a rebuilt evaluation package (cloud-deployed containers, rollout interfaces for RL/SFT training, a simpler any-agent interface) that replaced the original harness as the way tasks are run and graded.
| Model | Provider | Score | Card as of |
|---|---|---|---|
| Claude Mythos Preview | Anthropic | 82.0 | 2026-04 |
| GPT-5 | OpenAI | 77.3 | 2026-04 |
| GPT-5 Chat (latest) | OpenAI | 77.3 | 2026-04 |
| GPT-5.3 Chat (latest) | OpenAI | 77.3 | 2026-04 |
| GPT-5.3 Codex | OpenAI | 77.3 | 2026-04 |
| GPT-5.3 Codex Spark | OpenAI | 77.3 | 2026-04 |
| GPT-5.4 | OpenAI | 75.1 | 2026-04 |
| Gemini 3.1 Pro Preview | Google DeepMind | 68.5 | 2026-04 |
| Claude Opus 4 | Anthropic | 65.4 | 2026-04 |
| Claude Opus 4.6 | Anthropic | 65.4 | 2026-04 |
| GPT-5.2 | OpenAI | 64.0 | 2026-04 |
| GPT-5.2 Chat | OpenAI | 64.0 | 2026-04 |
| GPT-5.2 Codex | OpenAI | 64.0 | 2026-04 |
| GLM 5.1 | Z.ai (Zhipu AI) | 63.5 | 2026-04 |
| Claude Opus 4.5 | Anthropic | 59.3 | 2026-04 |
| Claude Opus 4.5 (latest) | Anthropic | 59.3 | 2026-04 |
| Claude Sonnet 4 | Anthropic | 59.1 | 2026-04 |
| Claude Sonnet 4.6 | Anthropic | 59.1 | 2026-04 |
| Muse Spark | Meta | 59.0 | 2026-04 |
| GPT-5.1 | OpenAI | 52.8 | 2026-04 |
| GPT-5.1 Chat | OpenAI | 52.8 | 2026-04 |
| GPT-5.1 Codex | OpenAI | 52.8 | 2026-04 |
| GPT-5.1 Codex Max | OpenAI | 52.8 | 2026-04 |
| GPT-5.1 Codex mini | OpenAI | 52.8 | 2026-04 |
| DeepSeek Chat | DeepSeek | 46.4 | 2026-04 |
| DeepSeek V3 | DeepSeek | 46.4 | 2026-04 |
| DeepSeek V3.2 | DeepSeek | 46.4 | 2026-04 |
| DeepSeek V3.2 Exp | DeepSeek | 46.4 | 2026-04 |
| Qwen3-Coder 480B-A35B Instruct | Alibaba / Qwen Team | 37.5 | 2026-04 |