Terminal-Bench

Terminal-Bench 1.0 measures whether an AI agent can complete real command-line tasks in a sandboxed Docker terminal, graded by automated tests; superseded by later major versions.

Also known as: Terminal-Bench 1.0, Terminal-Bench-Core, T-Bench

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryagentic
Subcategorygeneral-purpose terminal / command-line agent task completion
Page statussuperseded
Metrictask resolution rate (tasks passed / tasks attempted)
Directionhigher_is_better
Unit%
Dataset size80
Dataset licenceApache-2.0
PublisherOriginally released under the Laude Institute's GitHub organisation; the project is now hosted under the Harbor Framework

What it measures

Terminal-Bench gives an agent a natural-language instruction and a live, sandboxed terminal (a Docker container, sometimes several linked containers) and asks it to complete the task by issuing commands, the way a person would work at a shell. Tasks at release covered scientific workflows, network configuration, games, data analysis, calling APIs and fixing security vulnerabilities, deliberately going beyond software-engineering-only benchmarks like SWE-bench. Everything happens in text, matching how language models are trained, but the range of tools, state and multi-step planning required is broader than a single coding task.

Task format

An agent receives a task instruction and a terminal session inside a Docker environment. The Terminal-Bench harness orchestrates the agent (installed directly in the container, integrated via a Python interface, or connected through an MCP server exposing a tmux session), logs its actions, and checks final container state against a task-specific automated test script after the agent finishes or times out. Each task also ships a human-verified reference solution.

Models reporting this benchmark

These figures come from the model cards, which carry one collection date per card and no per-score attribution. They are shown as reported, not as verified evidence.
ModelProviderScoreCard as of
o3OpenAI58.22026-04
o3-deep-researchOpenAI58.22026-04
o3-miniOpenAI58.22026-04
o3-proOpenAI58.22026-04
Claude Opus 4Anthropic55.82026-04
Claude Opus 4.6Anthropic55.82026-04
Claude Opus 4.1Anthropic52.12026-04
Claude Opus 4.1 (latest)Anthropic52.12026-04
GPT-5OpenAI48.82026-04
GPT-5 Chat (latest)OpenAI48.82026-04
GPT-5 MiniOpenAI48.82026-04
GPT-5 NanoOpenAI48.82026-04
GPT-5 ProOpenAI48.82026-04
GPT-5-CodexOpenAI48.82026-04
GPT-5.1OpenAI48.82026-04
GPT-5.1 ChatOpenAI48.82026-04
GPT-5.1 CodexOpenAI48.82026-04
GPT-5.1 Codex MaxOpenAI48.82026-04
GPT-5.1 Codex miniOpenAI48.82026-04
GPT-5.2OpenAI48.82026-04
GPT-5.2 ChatOpenAI48.82026-04
GPT-5.2 CodexOpenAI48.82026-04
GPT-5.2 ProOpenAI48.82026-04
GPT-5.3 Chat (latest)OpenAI48.82026-04
GPT-5.3 CodexOpenAI48.82026-04
GPT-5.3 Codex SparkOpenAI48.82026-04
GPT-5.4OpenAI48.82026-04
GPT-5.4 miniOpenAI48.82026-04
GPT-5.4 nanoOpenAI48.82026-04
GPT-5.4 ProOpenAI48.82026-04
Claude Sonnet 4Anthropic48.22026-04
Claude Sonnet 4.5Anthropic48.22026-04
Claude Sonnet 4.5 (latest)Anthropic48.22026-04
Gemini 2.5 ProGoogle DeepMind45.52026-04
GPT-4.1OpenAI42.12026-04
Claude Haiku 4.5Anthropic41.02026-04
Claude Haiku 4.5 (latest)Anthropic41.02026-04
Claude Opus 4 (latest)Anthropic39.22026-04
DeepSeek ChatDeepSeek37.72026-04
DeepSeek V3DeepSeek37.72026-04
DeepSeek V3.2DeepSeek37.72026-04
DeepSeek V3.2 ExpDeepSeek37.72026-04
Claude Sonnet 4 (latest)Anthropic35.52026-04
Claude Sonnet 3.7Anthropic35.22026-04
GPT-4oOpenAI35.22026-04
GPT-4o (2024-05-13)OpenAI35.22026-04
GPT-4o (2024-08-06)OpenAI35.22026-04
GPT-4o (2024-11-20)OpenAI35.22026-04
GPT-4o miniOpenAI35.22026-04
DeepSeek V3.1DeepSeek31.32026-04
DeepSeek R1DeepSeek5.72026-04
DeepSeek R1 0528DeepSeek5.72026-04
DeepSeek R1 0528 NVFP4 v2NVIDIA5.72026-04

Data

This page as JSON · Edit on GitHub