Manually localized Icelandic WinoGrande schemas scored as two-way fill-in-the-blank accuracy; lm-eval runs the public train split.
unassessed
| Category | reasoning |
|---|---|
| Subcategory | Icelandic commonsense pronoun resolution, localized from WinoGrande |
| Page status | unknown |
| Metric | accuracy |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 1088 |
| Dataset licence | CC-BY-4.0 |
| Publisher | Miðeind ehf. and University of Iceland |
Icelandic WinoGrande gives a short Icelandic sentence with a blank and two candidate noun phrases. The model must pick the filler that makes commonsense sense, not the one that merely agrees in grammar. Miðeind and the University of Iceland translated and adapted English WinoGrande test items by hand so gender, number, and case would not leak the answer. Items that could not be localized were skipped or rewritten. The skill is Icelandic coreference and everyday inference, not English WinoGrande and not IceBERT's other Icelandic tagging tasks.
Two-way fill-in-the-blank. Each row has sentence, option1, option2, and answer "1" or "2". lm-evaluation-harness scores it as multiple_choice with partial cloze likelihoods (Trinh and Le 2018), matching English winogrande in the same harness. The Hub eval.py compares causal-LM loss on the two full sentences, default 3-shot.
No model card in ModelSpec reports this benchmark yet.