BLiMP-NL (Benchmark of Linguistic Minimal Pairs for Dutch)

8,400 Dutch minimal pairs across 84 paradigms and 22 phenomena, scored by whether a model prefers the grammatical sentence over a close ungrammatical match.

Also known as: BLiMP-NL, BLiMP-NL large

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryknowledge
SubcategoryDutch grammatical acceptability, minimal-pair paradigms
Page statusactive
Metricpairwise accuracy (grammatical sentence assigned the higher probability)
Directionhigher_is_better
Unit%
Dataset size8400
Dataset licenceCC-BY-SA-4.0
PublisherRadboud University (Centre for Language Studies), with University of Amsterdam

What it measures

BLiMP-NL tests whether a language model’s probabilities favour grammatical Dutch over a minimally different ungrammatical sentence. Each item is a pair, not a question. The contrasts cover 22 syntactic phenomena that matter in Dutch, including verb-second order, R-words, crossing dependencies, and parasitic gaps, rather than a translation of English BLiMP. The model is not asked to label sentences or explain a rule. A high score means the distribution ranks the good sentence above the bad one.

Task format

Zero-shot forced choice by likelihood. lm-evaluation-harness leaves the prompt empty and compares log-probability of sentence_good against sentence_bad (doc_to_target 0). The paper’s original scoring used masked models and syntactic log-odds ratios (SLOG); the harness does not implement SLOG.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub