Terminal-Bench 2.0

A harder, more heavily verified 89-task remake of Terminal-Bench, released with the Harbor evaluation package; Terminal-Bench 2.1 later patched 28 of its tasks.

Also known as: Terminal-Bench 2, TB2, Terminal-Bench 2.1

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryagentic
Subcategorygeneral-purpose terminal / command-line agent task completion, verified task set
Page statusactive
Metrictask resolution rate (tasks passed / tasks attempted)
Directionhigher_is_better
Unit%
Dataset size89
Dataset licenceApache-2.0
PublisherThe Terminal-Bench Team, hosted under the Harbor Framework project

What it measures

Terminal-Bench 2.0 measures the same thing as the original Terminal-Bench: whether an AI agent can carry out a real task by issuing commands in a live, sandboxed terminal. The task domains and text-only interface are unchanged; what changed is quality control. The authors write that they "weren't satisfied with the level of verification" in the original dataset — for example, a task that scraped YouTube broke whenever YouTube's anti-bot defences changed — so 2.0 puts each task through substantial manual and LM-assisted review before inclusion.

Task format

Unchanged from the original Terminal-Bench: an agent receives a task instruction and a sandboxed Docker terminal, and must complete the task through shell commands. Terminal-Bench 2.0 shipped alongside Harbor, a rebuilt evaluation package (cloud-deployed containers, rollout interfaces for RL/SFT training, a simpler any-agent interface) that replaced the original harness as the way tasks are run and graded.

Models reporting this benchmark

These figures come from the model cards, which carry one collection date per card and no per-score attribution. They are shown as reported, not as verified evidence.
ModelProviderScoreCard as of
Claude Mythos PreviewAnthropic82.02026-04
GPT-5OpenAI77.32026-04
GPT-5 Chat (latest)OpenAI77.32026-04
GPT-5.3 Chat (latest)OpenAI77.32026-04
GPT-5.3 CodexOpenAI77.32026-04
GPT-5.3 Codex SparkOpenAI77.32026-04
GPT-5.4OpenAI75.12026-04
Gemini 3.1 Pro PreviewGoogle DeepMind68.52026-04
Claude Opus 4Anthropic65.42026-04
Claude Opus 4.6Anthropic65.42026-04
GPT-5.2OpenAI64.02026-04
GPT-5.2 ChatOpenAI64.02026-04
GPT-5.2 CodexOpenAI64.02026-04
GLM 5.1Z.ai (Zhipu AI)63.52026-04
Claude Opus 4.5Anthropic59.32026-04
Claude Opus 4.5 (latest)Anthropic59.32026-04
Claude Sonnet 4Anthropic59.12026-04
Claude Sonnet 4.6Anthropic59.12026-04
Muse SparkMeta59.02026-04
GPT-5.1OpenAI52.82026-04
GPT-5.1 ChatOpenAI52.82026-04
GPT-5.1 CodexOpenAI52.82026-04
GPT-5.1 Codex MaxOpenAI52.82026-04
GPT-5.1 Codex miniOpenAI52.82026-04
DeepSeek ChatDeepSeek46.42026-04
DeepSeek V3DeepSeek46.42026-04
DeepSeek V3.2DeepSeek46.42026-04
DeepSeek V3.2 ExpDeepSeek46.42026-04
Qwen3-Coder 480B-A35B InstructAlibaba / Qwen Team37.52026-04

Data

This page as JSON · Edit on GitHub