A 6,390-query BIG-bench task that tests few-shot learning of 426 P3 binary-string programs from machine-teaching witness sets.
unassessed
| Category | reasoning |
|---|---|
| Subcategory | BIG-bench few-shot P3 binary-string programs from machine-teaching witness sets (6,390 queries) |
| Page status | unknown |
| Metric | exact_str_match |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 6390 |
| Dataset licence | Apache-2.0 |
| Publisher | Google (BIG-bench collaboration); Universitat Politècnica de València |
simp_turing_concept shows Input/Output pairs of binary strings and asks for the next output, testing whether a model shares a simplicity prior with a machine teacher for P3 (a small Turing-complete language). 426 concepts; three nested teaching batches (witness set, AS I, AS II); five held-out test strings per concept (max length 5). Parent task.json has no examples; subtasks do. Dummy-model header: 0 multiple-choice and 6,390 free-text queries. Witness-set JSON counted 2,130 examples (426 × 5). Preferred metric exact_str_match. Must be run as BIG-bench zero-shot because shots are already in the prompt. Canary GUID embedded. Not in BIG-Bench Hard.
Free-text continuation. Prompt style Input: … Output: with predetermined few-shot pairs. output_regex "('.*?')". Keywords many-shot, logical reasoning, computer code, json, free response. Subtasks witness_set, additional_set_1, additional_set_2.
No model card in ModelSpec reports this benchmark yet.