PAWS

Inspect Evals yes/no paraphrase detection on the 8,000-item PAWS-Wiki labeled-final test set of high-overlap sentence pairs.

Also known as: PAWS-Wiki, Paraphrase Adversaries from Word Scrambling

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryreasoning
Subcategorybinary paraphrase identification on high lexical-overlap English sentence pairs
Page statusactive
Metricaccuracy (Inspect includes scorer against Yes/No)
Directionhigher_is_better
Unit%
Dataset size8000
Dataset licenceother (Google: may be freely used for any purpose, acknowledgement appreciated; AS IS, no warranty)
PublisherGoogle Research

What it measures

This id is UKGovernmentBEIS inspect_evals task paws, not PAWS-X and not GLUE MRPC. Each item is two English sentences with high bag-of-words overlap. The model must answer Yes if they are paraphrases and No otherwise. Zhang, Baldridge and He (NAACL 2019) built the pairs from Wikipedia (and a separate QQP-derived set) by word swapping and back-translation so that overlap no longer implies equivalence. Inspect uses Hugging Face config labeled_final, test split only.

Task format

Yes/No generation. Inspect prompt: answer Yes or No whether two sentences are paraphrases, with Sentence1/Sentence2 lines, no other tokens. Default shuffle true. English text.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub