BIG-bench Lite task: pick which sentence correctly describes a 'structure' — a sequence of six emoji pieces — across five adversarial variants.
unassessed
| Category | reasoning |
|---|---|
| Subcategory | structured visual/logical reasoning via emoji-symbol interpretation (BIG-bench Lite) |
| Page status | unknown |
| Metric | multiple_choice_grade |
| Direction | higher_is_better |
| Dataset size | 990 |
| Dataset licence | Apache-2.0 |
| Publisher | Google (BIG-bench collaboration) |
The model is given a "structure": a sequence of six pieces represented by emojis, standing in for objects in a simple constructed world. It must choose, from a set of candidate sentences, the one that correctly and consistently describes two given structures. The task is split into five subtasks that vary how directly the emojis map to their described meaning: a "plain" version with direct emoji-to-name correspondence, an "adversarial" version with intentionally mismatched emoji-name associations, a "tricky" version with reversed object descriptions, and two "agnostic" versions that substitute generic placeholders for either the names or the emojis. Within each subtask, items escalate across difficulty tiers covering simple quantification, logical operators, and positional relationships between pieces.
Multiple-choice, zero-shot; each item asks which sentence is consistent with two given emoji structures.
No model card in ModelSpec reports this benchmark yet.