72 auto-generated minimal-pair test sets probing whether a model's probabilities favour the grammatical sentence for subject-verb agreement, reflexive anaphora and negative polarity items.
unassessed
| Category | knowledge |
|---|---|
| Subcategory | targeted syntactic minimal pairs (agreement, reflexives, negative polarity items) |
| Page status | active |
| Metric | pairwise accuracy (grammatical sentence assigned the higher probability) |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 158084 |
| Dataset licence | MIT |
LM-SynEval measures what a language model implicitly knows about specific points of English syntax, not whether it can perform a task or explain a rule. Each item is a minimal pair of sentences differing in one grammatical property -- for example "The author laughs" versus "The author laugh" -- and the model is scored by whether it assigns a higher probability to the grammatical member of the pair. The pairs are organised into three phenomena: subject-verb agreement (tested across many constructions -- across a prepositional phrase, a sentential complement, subject and object relative clauses with and without an overt "that", and verb-phrase coordination), reflexive anaphora (does the reflexive pronoun's number match its antecedent, again tested within simple sentences and across relative clauses), and negative polarity item licensing (does "ever" or a similar NPI appear only where a licensing context such as "no" makes it grammatical). A model can score well on every construction here while being unable to state, in words, the agreement or licensing rule it is implicitly satisfying -- this is a probe of linguistic competence, closer to a psycholinguistic acceptability experiment than to a benchmark of task-solving ability like question answering.
Forced-choice by probability comparison: for each minimal pair, compare the model's log-probability on the grammatical sentence against its ungrammatical, minimally different counterpart; the harness presents exactly two choices per item and no explicit answer or generation is requested.
No model card in ModelSpec reports this benchmark yet.