lm-eval multilingual paraphrase identification on PAWS-X: seven languages of high-overlap sentence pairs, scored by accuracy.
unassessed
| Category | reasoning |
|---|---|
| Subcategory | multilingual binary paraphrase identification on translated PAWS-Wiki pairs |
| Page status | active |
| Metric | accuracy (acc); group mean weighted by size |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 23659 |
| Dataset licence | other (same Google PAWS licence: free use with acknowledgement, AS IS) |
| Publisher | Google Research |
This id is EleutherAI lm-eval group pawsx (directory paws-x), not English-only Inspect paws and not a translation quality test. Each item is a sentence pair in German, English, Spanish, French, Japanese, Korean or Chinese. The model must decide whether the pair is a paraphrase. Non-English evaluation pairs are human translations of PAWS-Wiki; training pairs are machine translated. lm-eval casts the decision as two cloze strings: "{s1}, right? No, {s2}" versus "{s1}, right? Yes, {s2}", with language-specific wording from Google Translate.
Multiple-choice likelihood (output_type multiple_choice) over No versus Yes continuations. Group pawsx aggregates paws_de, paws_en, paws_es, paws_fr, paws_ja, paws_ko, paws_zh. English prompt shown in the harness README; other languages use translated masks.
No model card in ModelSpec reports this benchmark yet.