{
 "body": "\n## What it measures\n\n300 authentic coding tasks in ten languages test agents on multilingual software-development problems. An agent receives an instruction and works inside a terminal environment. The task may involve programming, debugging, system administration, data processing or scientific computation, depending on the release.\n\nSuccess requires both useful interaction and a completed artifact. The benchmark measures agent planning, tool use, environment handling and end-to-end task completion. It is not a static coding-question exam.\n\n## How it is scored\n\nTerminal-Bench-style releases verify each task with task-specific tests or an oracle solution. A run normally reports the fraction of tasks whose tests pass. The source documentation does not establish a random or human baseline.\n\nResults depend on the agent adapter, model, container image, time limit, concurrency, network and exact dataset version. A score from Harbor on one release should not be compared with another without recording those settings. LILT reports pass rate across multilingual tasks, while versioned Harbor releases use executable task verification.\n\n## Dataset and licence\n\nThe primary source states 300 tasks for this release. Tasks include instructions, environments and tests; some releases also mirror solutions or answer keys. A dataset licence was not established from the opened source. Task-level source licences may still differ.\n\nVersion and mirror provenance matter. Hugging Face cards for Terminal-Bench 2.0 and 3.0 identify themselves as mirrors and point readers to source repositories or Harbor Hub.\n\n## Who publishes it\n\nKim Yunsu; Uhlig Kaden; Purohit Ashwin; Agarwal Milind; Simianer Patrick; Arslan Anil; Mokhtari Kiarash; Zenkel Thomas; Mosig Johannes; Bretschner Gabriel; Bose Shamik; Wuebker Joern; DeNero John is the credited publisher or author information in the opened source. The Harbor and Terminal-Bench projects maintain the execution framework and versioned task releases. No universal leaderboard was established for every assigned version.\n\n## Lineage\n\nTerminal-Bench is a versioned family of terminal-agent evaluations. Terminal-Bench 2.0, its verified derivative, Terminal-Bench 3.0 and Terminal-Bench-Science are distinct releases or task collections, not interchangeable scores. Terminal-Bench-LILT is a multilingual coding extension with a separate paper and task suite.\n\nThe v2.1 and v4.0 leads supplied for this assignment resolve only to an Artificial Analysis logo asset. They do not establish benchmark identities, releases or aliases.\n\n## Saturation and contamination\n\nNo current saturation ceiling was established. Public task instructions, tests, mirrors and, for some releases, solutions create contamination risk. A result should state whether answer keys were available to the evaluated agent and whether network access was enabled.\n\nThe science mirror states that its preview release is private upstream because tests and solutions are included. That distinction matters when interpreting scores.\n\n## How to run it\n\nUse Harbor or the Terminal-Bench CLI with the exact dataset name and version. Terminal-Bench 2.0 uses Harbor dataset terminal-bench@2.0; Terminal-Bench 3.0 documentation identifies version 3.0.0; Terminal-Bench-Science identifies v0.1.0. Run the matching release and record agent, model, container provider, time limit, concurrency and network policy.\n\n## Reading the numbers\n\nA high pass rate means the agent completed and passed selected terminal tasks under a particular environment. It does not isolate model reasoning from tools, container images, tests, timeouts or agent scaffolding. Versioned releases can change task difficulty and infrastructure. Compare only runs with the same release and execution protocol.\n\n",
 "build": {
  "built_at": "2026-09-09T16:56:50+00:00",
  "commit": "0a599558854c0e238c03a0f0d725239cb28f9d11",
  "eligibility_as_of": "2026-09-09"
 },
 "disposition": {
  "canonical_id": "terminal_bench_lilt",
  "reasons": [],
  "status": "unassessed",
  "verified_results": []
 },
 "models_covered": [],
 "page": {
  "aliases": [],
  "category": "agentic",
  "contamination": {
   "note": "Task instructions and, in several releases, tests or solutions are public or mirrored; no rotating private test policy was established.",
   "risk": "medium"
  },
  "dataset": {
   "languages": [
    "en"
   ],
   "license": "",
   "modalities": [
    "text"
   ],
   "public_test_set": true,
   "size": 300,
   "size_note": "The primary source states 300 tasks.",
   "splits": "versioned task collection; no train/test split established",
   "url": "https://arxiv.org/abs/2608.28641"
  },
  "freshness": {
   "researched": "2026-09-09",
   "researched_by": "GPT-5.6 Luna, luna-stream-c-002 (Codex coordinated)",
   "reviewed": "",
   "reviewed_by": ""
  },
  "harness": {
   "bigbench": "",
   "helm": "",
   "inspect_evals": "",
   "lm_eval": "",
   "opencompass": "",
   "other": "Harbor/Terminal-Bench execution harness"
  },
  "id": "terminal_bench_lilt",
  "last_updated": "",
  "leaderboard_url": "",
  "lineage": {
   "family": "terminal_bench",
   "predecessor": "",
   "successors": [],
   "variants": []
  },
  "measures": "300 authentic coding tasks in ten languages test agents on multilingual software-development problems. The benchmark gives an agent an English task, a terminal environment and a verification procedure; success depends on completing the task and passing its tests.\n",
  "metric": {
   "baseline_note": "The opened primary documentation does not publish a random or human baseline.",
   "direction": "higher_is_better",
   "human_baseline": null,
   "max_score": 100,
   "name": "task pass rate",
   "random_baseline": null,
   "unit": "%"
  },
  "name": "Terminal-Bench-LILT",
  "page_kind": "benchmark",
  "paper": {
   "arxiv": "2608.28641",
   "title": "Terminal-Bench-LILT: Multilingual Agentic Coding Benchmark Grounded in Language, Region, and Culture",
   "url": "https://arxiv.org/abs/2608.28641",
   "year": 2026
  },
  "publisher": {
   "authors": [
    "Kim Yunsu"
   ],
   "org": "Kim Yunsu; Uhlig Kaden; Purohit Ashwin; Agarwal Milind; Simianer Patrick; Arslan Anil; Mokhtari Kiarash; Zenkel Thomas; Mosig Johannes; Bretschner Gabriel; Bose Shamik; Wuebker Joern; DeNero John",
   "url": "https://github.com/lilt/terminal-bench-lilt"
  },
  "released": "2026-08",
  "repo_url": "https://github.com/lilt/terminal-bench-lilt",
  "saturation": {
   "as_of": "",
   "note": "No current saturation result was established from the opened primary source.",
   "status": "unknown",
   "top_score": null
  },
  "sources": [
   {
    "accessed": "2026-09-09",
    "title": "Terminal-Bench-LILT primary dataset or paper source",
    "url": "https://arxiv.org/abs/2608.28641"
   },
   {
    "accessed": "2026-09-09",
    "title": "Terminal-Bench-LILT source repository",
    "url": "https://github.com/lilt/terminal-bench-lilt"
   }
  ],
  "status": "unknown",
  "subcategory": "Multilingual agentic coding tasks grounded in language, region and culture",
  "summary": "300 authentic coding tasks in ten languages test agents on multilingual software-development problems.",
  "tags": [
   "terminal",
   "agents",
   "tool-use",
   "container"
  ],
  "task_format": "Agent interacts with a containerized terminal; task-specific tests determine success."
 }
}