1,133 UK-Linguistics-Olympiad-style puzzles across 90+ mostly low-resource languages, scored on direct accuracy and a no-context control that penalises memorisation.
unassessed
| Category | reasoning |
|---|---|
| Subcategory | olympiad-level linguistic reasoning puzzles in low-resource and extinct languages |
| Page status | active |
| Metric | exact match accuracy, plus a no-context delta |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 1133 |
| Dataset licence | The Hugging Face dataset card states CC BY-NC-ND 4.0 plus an acceptable-use policy that forbids redistributing questions or answers in a web-scrapable plain-text format and forbids training directly on the benchmark. The GitHub repository itself carries no machine-readable licence (GitHub's own API reports its licence as unassigned). |
| Publisher | University of Oxford (Oxford Internet Institute), with co-authors at Stanford University, the UK Linguistics Olympiad and Meedan |
LingOly gives a model a full linguistics-olympiad problem sheet -- background on an unfamiliar, usually very low-resource or extinct, language, a set of example words or sentences, and one or more sub-questions -- then asks it to answer specific sub-questions using only the patterns shown on the sheet. Six question formats appear (including translation into and out of the target language, pattern completion, and match-up tasks) across five levels of human difficulty. Because the target languages are deliberately obscure, a correct answer should come from in-context pattern generalisation rather than from facts the model already knew about the language, and the benchmark also tests whether a model can follow the sheet's often intricate formatting instructions.
Free-text answer generation from a full problem-sheet prompt (background, worked examples, and the specific sub-question), with the model told to return a JSON object keyed by sub-question number; answers are graded per sub-question, not per sheet.
No model card in ModelSpec reports this benchmark yet.