Terminal-Bench 2.0 Verified

This reviewed derivative fixes environment and instruction issues in Terminal-Bench 2.0 while preserving task logic where stated.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryagentic
SubcategoryReviewed Terminal-Bench 2.0 tasks for agent execution
Page statusunknown
Metrictask success verified by tests
Directionhigher_is_better
Unit%
Dataset size89
Dataset licenceApache-2.0
PublisherTerminal Bench 2 Verified

What it measures

This reviewed derivative fixes environment and instruction issues in Terminal-Bench 2.0 while preserving task logic where stated. The benchmark gives an agent an English task, a terminal environment and a verification procedure; success depends on completing the task and passing its tests.

Task format

Agent interacts with a containerized terminal; task-specific tests determine success.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub