A 6,000-item BIG-bench task that, given a UCI game prefix and a start square, asks for any legal destination square of that piece.
unassessed
| Category | reasoning |
|---|---|
| Subcategory | BIG-bench UCI legal-destination prediction (6 x 1,000 games) |
| Page status | unknown |
| Metric | exact_str_match |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 6000 |
| Dataset licence | Apache-2.0 |
| Publisher | Google (BIG-bench collaboration); authors at TTIC |
chess_state_tracking measures whether a language model has an implicit chess board. The prompt is a Universal Chess Interface (UCI) move list plus the start square of the current non-pawn, non-castling move. The target is any legal destination square for that piece. Authors Shubham Toshniwal, Sam Wiseman, Karen Livescu, and Kevin Gimpel split the work into real Lichess games and synthetic games from a random tree search, each at short, medium, and long length. The skill is legal-move geometry, not mate finding.
Free-text destination square. Preferred metric exact_str_match with output_regex [a-h][1-8]. Targets are lists of legal squares. task_prefix asks the model to complete the last move by filling in the destination. Canary GUID embedded.
No model card in ModelSpec reports this benchmark yet.