StereoSet

Crowdsourced test of whether a language model prefers stereotypical over anti-stereotypical associations across gender, race, religion and profession, while staying fluent.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorysafety
Subcategorystereotypical bias
Page statusactive
MetricIdealized CAT Score (ICAT)
Directionhigher_is_better
Unitpoints
Dataset size16995
Dataset licenceCC BY-SA 4.0
PublisherMIT (Nadeem, Bethke) and McGill University / Mila (Reddy)

What it measures

StereoSet gives a model a context sentence about a person or group (a gender, race, religion or profession target) and asks it to choose among three completions: one that reflects a common stereotype about the target, one that reflects an anti-stereotype, and one that is unrelated or meaningless. The intrasentence variant fills in a blank within a single sentence with a stereotype, anti-stereotype or unrelated word; the intersentence variant picks the most plausible of three follow-on sentences. The design deliberately separates two things a language model could get wrong: producing fluent, on-topic language at all, and doing so by leaning on a stereotype rather than a neutral or counter-stereotypical association.

Task format

Three-way forced choice per item (stereotype / anti-stereotype / unrelated), either filling a blank within a sentence (intrasentence) or selecting a following sentence (intersentence).

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub