lm-eval group of Uyghur textbook, exam, and cloze tasks covering basic language, literature, and last-word prediction.
unassessed
| Category | composite |
|---|---|
| Subcategory | Uyghur language and literature exams, textbook exercises, and last-word cloze |
| Page status | active |
| Metric | per-task accuracy or exact match; ulut also reports a size-weighted acc mean |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 1468 |
| Dataset licence | Mixed: Apache-2.0 on keramjan/ulqa, uleval, CELEP1, CELEP2, lambada_uyghur; MIT on keramjan/ulut. Underlying textbooks and exam papers are separately copyrighted. |
| Publisher | keramjan (Hugging Face); packaged for lm-eval as group ulqa |
ulqa is EleutherAI's group name for six Uyghur evaluations assembled by Hugging Face user keramjan, not a single published paper. The group mixes last-word cloze (lambada_uyghur), a five-task Uyghur language-understanding bundle (ulut), short generative language QA (ulqa_), four-way language MCQ (uleval), and 2011-2012 college-entrance literature items (celep1 multiple-choice, celep2 open-ended). All prompts are in Uyghur. The skill is Uyghur reading, vocabulary, grammar and literature exam answering, not a shared metric.
Mixed. lambada_uyghur is log-likelihood last-word prediction. ulut, uleval and celep1 are multiple-choice. ulqa_ and celep2 are greedy generate_until, stopping at "سۇئال:". ulqa_ is 2-shot; celep2 is 0-shot. Every YAML uses the Hugging Face train split as the test set. There is no held-out test split.
No model card in ModelSpec reports this benchmark yet.