TurBLiMP Core

TurBLiMP's core group of 16 Turkish grammaticality phenomena, 1,000 minimal pairs each, run as a single lm-evaluation-harness task.

Also known as: TurBLiMP

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryknowledge
SubcategoryTurkish grammaticality, minimal-pair paradigms
Page statusactive
Metricaccuracy (acc) and length-normalized accuracy (acc_norm)
Directionhigher_is_better
Unitpercent
Dataset size16000
Dataset licenceCC BY 4.0

What it measures

TurBLiMP tests whether a model's probabilities favour the grammatical member of a minimal pair of Turkish sentences, across 16 phenomena including subject and anaphor agreement, binding, island effects, scrambling, and suspended affixation, with particular attention to Turkish's flexible word order and morphological subordination.

Task format

Forced binary choice between two near-identical Turkish sentences that differ by one grammatical violation; the model is scored on which sentence it assigns higher log-probability to, not on a generated answer.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub