ULQA (Uyghur language eval group)

lm-eval group of Uyghur textbook, exam, and cloze tasks covering basic language, literature, and last-word prediction.

Also known as: ulqa_, uleval, ULUT, CELEP1, CELEP2, lambada_uyghur

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorycomposite
SubcategoryUyghur language and literature exams, textbook exercises, and last-word cloze
Page statusactive
Metricper-task accuracy or exact match; ulut also reports a size-weighted acc mean
Directionhigher_is_better
Unit%
Dataset size1468
Dataset licenceMixed: Apache-2.0 on keramjan/ulqa, uleval, CELEP1, CELEP2, lambada_uyghur; MIT on keramjan/ulut. Underlying textbooks and exam papers are separately copyrighted.
Publisherkeramjan (Hugging Face); packaged for lm-eval as group ulqa

What it measures

ulqa is EleutherAI's group name for six Uyghur evaluations assembled by Hugging Face user keramjan, not a single published paper. The group mixes last-word cloze (lambada_uyghur), a five-task Uyghur language-understanding bundle (ulut), short generative language QA (ulqa_), four-way language MCQ (uleval), and 2011-2012 college-entrance literature items (celep1 multiple-choice, celep2 open-ended). All prompts are in Uyghur. The skill is Uyghur reading, vocabulary, grammar and literature exam answering, not a shared metric.

Task format

Mixed. lambada_uyghur is log-likelihood last-word prediction. ulut, uleval and celep1 are multiple-choice. ulqa_ and celep2 are greedy generate_until, stopping at "سۇئال:". ulqa_ is 2-shot; celep2 is 0-shot. Every YAML uses the Hugging Face train split as the test set. There is no held-out test split.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub