Terminal-Bench v4.0

The supplied v4.0 lead resolves only to an Artificial Analysis logo asset, not an authoritative benchmark release or task registry.

unverified

This page is not in the default catalogue. Evidence required by the catalogue contract is missing or was not approved by a reviewer. That is a statement about the evidence we hold, not a claim that the benchmark is stale or illegitimate.

Recorded reasons:

Categoryagentic
SubcategoryTerminal-Bench version label
Page statusunknown
Metricnot established
Directionhigher_is_better
Unit%
PublisherArtificial Analysis lead

What it measures

The supplied v4.0 lead resolves only to an Artificial Analysis logo asset, not an authoritative benchmark release or task registry. The benchmark gives an agent an English task, a terminal environment and a verification procedure; success depends on completing the task and passing its tests.

Task format

Agent interacts with a containerized terminal; task-specific tests determine success.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub