BHS (Basque, Hindi, Swahili syntactic evaluation)

22 suites of 1,000 minimal pairs each that test whether models prefer grammatical Basque, Hindi, or Swahili continuations.

Also known as: BHS, Controlled Evaluation of Syntactic Knowledge

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryreasoning
Subcategorytargeted syntactic evaluation in Basque, Hindi, and Swahili
Page statusactive
Metricaccuracy
Directionhigher_is_better
Unit%
Dataset size22000
PublisherMassachusetts Institute of Technology (authors); Hugging Face mirror by jmichaelov

What it measures

BHS is a targeted syntactic evaluation (TSE) for three lower-resource languages. Each item is a minimal pair: two sentences that differ by one word, only one of which is grammatical given the rest of the sentence. The model should assign higher probability to the grammatical member. Basque suites probe auxiliary agreement with subject, direct object and indirect object across several word orders. Hindi suites probe perfective versus non-perfective verb form with and without the ergative clitic ne, with optional possessors in between. Swahili suites probe noun-class agreement on verbs and adjectives across intervening material. Items are generated from vocabularies, not scraped corpora.

Task format

Zero-shot minimal-pair scoring. lm-eval implements each suite as a two-way multiple-choice task (ending_good vs ending_bad) with metrics acc and acc_norm. The authors' script compares length-normalised log-probability of the last word, which lm-eval cannot match exactly.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub