TurBLiMP's core group of 16 Turkish grammaticality phenomena, 1,000 minimal pairs each, run as a single lm-evaluation-harness task.
unassessed
| Category | knowledge |
|---|---|
| Subcategory | Turkish grammaticality, minimal-pair paradigms |
| Page status | active |
| Metric | accuracy (acc) and length-normalized accuracy (acc_norm) |
| Direction | higher_is_better |
| Unit | percent |
| Dataset size | 16000 |
| Dataset licence | CC BY 4.0 |
TurBLiMP tests whether a model's probabilities favour the grammatical member of a minimal pair of Turkish sentences, across 16 phenomena including subject and anaphor agreement, binding, island effects, scrambling, and suspended affixation, with particular attention to Turkish's flexible word order and morphological subordination.
Forced binary choice between two near-identical Turkish sentences that differ by one grammatical violation; the model is scored on which sentence it assigns higher log-probability to, not on a generated answer.
No model card in ModelSpec reports this benchmark yet.