Terminal-Bench-LILT

300 authentic coding tasks in ten languages test agents on multilingual software-development problems.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryagentic
SubcategoryMultilingual agentic coding tasks grounded in language, region and culture
Page statusunknown
Metrictask pass rate
Directionhigher_is_better
Unit%
Dataset size300
PublisherKim Yunsu; Uhlig Kaden; Purohit Ashwin; Agarwal Milind; Simianer Patrick; Arslan Anil; Mokhtari Kiarash; Zenkel Thomas; Mosig Johannes; Bretschner Gabriel; Bose Shamik; Wuebker Joern; DeNero John

What it measures

300 authentic coding tasks in ten languages test agents on multilingual software-development problems. The benchmark gives an agent an English task, a terminal environment and a verification procedure; success depends on completing the task and passing its tests.

Task format

Agent interacts with a containerized terminal; task-specific tests determine success.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub