Terminal-Bench-Science evaluates real computational research tasks across life, physical, earth, mathematical and engineering sciences.
unassessed
| Category | agentic |
|---|---|
| Subcategory | Agent execution on computational research workflows |
| Page status | unknown |
| Metric | task success verified by tests |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 70 |
| Dataset licence | Apache-2.0 |
| Publisher | Harbor Framework |
Terminal-Bench-Science evaluates real computational research tasks across life, physical, earth, mathematical and engineering sciences. The benchmark gives an agent an English task, a terminal environment and a verification procedure; success depends on completing the task and passing its tests.
Agent interacts with a containerized terminal; task-specific tests determine success.
No model card in ModelSpec reports this benchmark yet.