PAWS-X

lm-eval multilingual paraphrase identification on PAWS-X: seven languages of high-overlap sentence pairs, scored by accuracy.

Also known as: paws-x, pawsx, PAWS-X: A Cross-lingual Adversarial Dataset for Paraphrase Identification

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryreasoning
Subcategorymultilingual binary paraphrase identification on translated PAWS-Wiki pairs
Page statusactive
Metricaccuracy (acc); group mean weighted by size
Directionhigher_is_better
Unit%
Dataset size23659
Dataset licenceother (same Google PAWS licence: free use with acknowledgement, AS IS)
PublisherGoogle Research

What it measures

This id is EleutherAI lm-eval group pawsx (directory paws-x), not English-only Inspect paws and not a translation quality test. Each item is a sentence pair in German, English, Spanish, French, Japanese, Korean or Chinese. The model must decide whether the pair is a paraphrase. Non-English evaluation pairs are human translations of PAWS-Wiki; training pairs are machine translated. lm-eval casts the decision as two cloze strings: "{s1}, right? No, {s2}" versus "{s1}, right? Yes, {s2}", with language-specific wording from Google Translate.

Task format

Multiple-choice likelihood (output_type multiple_choice) over No versus Yes continuations. Group pawsx aggregates paws_de, paws_en, paws_es, paws_fr, paws_ja, paws_ko, paws_zh. English prompt shown in the harness README; other languages use translated masks.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub