Codenames

An 85-item BIG-bench free-response task: given a Codenames-style clue and a word list, emit the associated words in alphabetical order.

Also known as: BIG-bench codenames

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryreasoning
SubcategoryBIG-bench Codenames-style free-response word association (85 items)
Page statusunknown
Metricbleu
Directionhigher_is_better
Dataset size85
Dataset licenceApache-2.0
PublisherGoogle (BIG-bench collaboration)

What it measures

codenames gives one English clue word and a board of candidate words, then asks the model to name the words that best match the clue. Isaac Noble, Lucy Noble, Emma Lam, and Lucas Lam wrote the clues and boards for BIG-bench; they are not dumps of the board game. The probe is analogical word association, not playing a full two-team Codenames match.

Task format

Free-text list. Preferred metric bleu (rouge is also listed). Inputs ask for 1, 2, 3, or 4 associated words and require alphabetical order. No multiple-choice options. Canary GUID embedded.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub