The supplied v4.0 lead resolves only to an Artificial Analysis logo asset, not an authoritative benchmark release or task registry.
unverified
Recorded reasons:
| Category | agentic |
|---|---|
| Subcategory | Terminal-Bench version label |
| Page status | unknown |
| Metric | not established |
| Direction | higher_is_better |
| Unit | % |
| Publisher | Artificial Analysis lead |
The supplied v4.0 lead resolves only to an Artificial Analysis logo asset, not an authoritative benchmark release or task registry. The benchmark gives an agent an English task, a terminal environment and a verification procedure; success depends on completing the task and passing its tests.
Agent interacts with a containerized terminal; task-specific tests determine success.
No model card in ModelSpec reports this benchmark yet.