Terminal-Bench 2.0 evaluates agents completing software, science and system tasks in Docker environments.
unassessed
| Category | agentic |
|---|---|
| Subcategory | Agent execution on real terminal tasks in containers |
| Page status | unknown |
| Metric | task success verified by tests |
| Direction | higher_is_better |
| Unit | % |
| Dataset licence | Apache-2.0 |
| Publisher | Terminal-Bench Team |
Terminal-Bench 2.0 evaluates agents completing software, science and system tasks in Docker environments. The benchmark gives an agent an English task, a terminal environment and a verification procedure; success depends on completing the task and passing its tests.
Agent interacts with a containerized terminal; task-specific tests determine success.
No model card in ModelSpec reports this benchmark yet.