300 authentic coding tasks in ten languages test agents on multilingual software-development problems.
unassessed
| Category | agentic |
|---|---|
| Subcategory | Multilingual agentic coding tasks grounded in language, region and culture |
| Page status | unknown |
| Metric | task pass rate |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 300 |
| Publisher | Kim Yunsu; Uhlig Kaden; Purohit Ashwin; Agarwal Milind; Simianer Patrick; Arslan Anil; Mokhtari Kiarash; Zenkel Thomas; Mosig Johannes; Bretschner Gabriel; Bose Shamik; Wuebker Joern; DeNero John |
300 authentic coding tasks in ten languages test agents on multilingual software-development problems. The benchmark gives an agent an English task, a terminal environment and a verification procedure; success depends on completing the task and passing its tests.
Agent interacts with a containerized terminal; task-specific tests determine success.
No model card in ModelSpec reports this benchmark yet.