LingOly

1,133 UK-Linguistics-Olympiad-style puzzles across 90+ mostly low-resource languages, scored on direct accuracy and a no-context control that penalises memorisation.

Also known as: LINGOLY, Linguistic Olympiad Benchmark

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryreasoning
Subcategoryolympiad-level linguistic reasoning puzzles in low-resource and extinct languages
Page statusactive
Metricexact match accuracy, plus a no-context delta
Directionhigher_is_better
Unit%
Dataset size1133
Dataset licenceThe Hugging Face dataset card states CC BY-NC-ND 4.0 plus an acceptable-use policy that forbids redistributing questions or answers in a web-scrapable plain-text format and forbids training directly on the benchmark. The GitHub repository itself carries no machine-readable licence (GitHub's own API reports its licence as unassigned).
PublisherUniversity of Oxford (Oxford Internet Institute), with co-authors at Stanford University, the UK Linguistics Olympiad and Meedan

What it measures

LingOly gives a model a full linguistics-olympiad problem sheet -- background on an unfamiliar, usually very low-resource or extinct, language, a set of example words or sentences, and one or more sub-questions -- then asks it to answer specific sub-questions using only the patterns shown on the sheet. Six question formats appear (including translation into and out of the target language, pattern completion, and match-up tasks) across five levels of human difficulty. Because the target languages are deliberately obscure, a correct answer should come from in-context pattern generalisation rather than from facts the model already knew about the language, and the benchmark also tests whether a model can follow the sheet's often intricate formatting instructions.

Task format

Free-text answer generation from a full problem-sheet prompt (background, worked examples, and the specific sub-question), with the model told to return a JSON object keyed by sub-question number; answers are graded per sub-question, not per sheet.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub