Crowdsourced test of whether a language model prefers stereotypical over anti-stereotypical associations across gender, race, religion and profession, while staying fluent.
unassessed
| Category | safety |
|---|---|
| Subcategory | stereotypical bias |
| Page status | active |
| Metric | Idealized CAT Score (ICAT) |
| Direction | higher_is_better |
| Unit | points |
| Dataset size | 16995 |
| Dataset licence | CC BY-SA 4.0 |
| Publisher | MIT (Nadeem, Bethke) and McGill University / Mila (Reddy) |
StereoSet gives a model a context sentence about a person or group (a gender, race, religion or profession target) and asks it to choose among three completions: one that reflects a common stereotype about the target, one that reflects an anti-stereotype, and one that is unrelated or meaningless. The intrasentence variant fills in a blank within a single sentence with a stereotype, anti-stereotype or unrelated word; the intersentence variant picks the most plausible of three follow-on sentences. The design deliberately separates two things a language model could get wrong: producing fluent, on-topic language at all, and doing so by leaning on a stereotype rather than a neutral or counter-stereotypical association.
Three-way forced choice per item (stereotype / anti-stereotype / unrelated), either filling a blank within a sentence (intrasentence) or selecting a following sentence (intersentence).
No model card in ModelSpec reports this benchmark yet.