Terminal-Bench 3.0 is the v3.0.0 Harbor dataset release with task instructions, environments, tests and solutions.
unassessed
| Category | agentic |
|---|---|
| Subcategory | Versioned agent execution on real terminal tasks |
| Page status | unknown |
| Metric | task success verified by tests |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 74 |
| Dataset licence | Apache-2.0 |
| Publisher | Harbor Framework |
Terminal-Bench 3.0 is the v3.0.0 Harbor dataset release with task instructions, environments, tests and solutions. The benchmark gives an agent an English task, a terminal environment and a verification procedure; success depends on completing the task and passing its tests.
Agent interacts with a containerized terminal; task-specific tests determine success.
No model card in ModelSpec reports this benchmark yet.