Inspect Evals yes/no paraphrase detection on the 8,000-item PAWS-Wiki labeled-final test set of high-overlap sentence pairs.
unassessed
| Category | reasoning |
|---|---|
| Subcategory | binary paraphrase identification on high lexical-overlap English sentence pairs |
| Page status | active |
| Metric | accuracy (Inspect includes scorer against Yes/No) |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 8000 |
| Dataset licence | other (Google: may be freely used for any purpose, acknowledgement appreciated; AS IS, no warranty) |
| Publisher | Google Research |
This id is UKGovernmentBEIS inspect_evals task paws, not PAWS-X and not GLUE MRPC. Each item is two English sentences with high bag-of-words overlap. The model must answer Yes if they are paraphrases and No otherwise. Zhang, Baldridge and He (NAACL 2019) built the pairs from Wikipedia (and a separate QQP-derived set) by word swapping and back-translation so that overlap no longer implies equivalence. Inspect uses Hugging Face config labeled_final, test split only.
Yes/No generation. Inspect prompt: answer Yes or No whether two sentences are paraphrases, with Sentence1/Sentence2 lines, no other tokens. Default shuffle true. English text.
No model card in ModelSpec reports this benchmark yet.