A 34-item BIG-bench Lite task that asks for a Python program's intermediate state or exception without running the code.
unassessed
| Category | coding |
|---|---|
| Subcategory | BIG-bench Lite Python program-state free-response (34 items) |
| Page status | unknown |
| Metric | exact_str_match |
| Direction | higher_is_better |
| Dataset size | 34 |
| Dataset licence | Apache-2.0 |
| Publisher | Google (BIG-bench collaboration) |
auto_debugging shows a short Python 3.7 snippet in a fenced block and asks a question about an intermediate variable, a final value, a printed output, or the exception the snippet raises. Authors Mo Tiwari, Chris Waites and Medina Baitemirova wrote the items so a correct answer requires tracing state without an interpreter. Several targets are lists of acceptable exception strings. The README frames a strong result as evidence a model could act as a static debugger.
Free-response JSON task scored with BIG-bench exact_str_match. A canary GUID is embedded. There is no multiple-choice YAML in lm-eval for this task.
No model card in ModelSpec reports this benchmark yet.