BOLD (Bias in Open-Ended Language Generation Dataset)

23,679 Wikipedia-derived prompts across five demographic domains, used to check whether a model's open-ended completions differ in sentiment, toxicity or regard depending on the group named in the prompt.

Also known as: Bias in Open-Ended Language Generation Dataset

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorysafety
Subcategorybias in open-ended text generation, read as gaps between demographic groups rather than one score
Page statusactive
Metricsentiment, toxicity, regard, gender polarity and psycholinguistic norms, compared across demographic groups
Directionhigher_is_better
Dataset size23679
Dataset licenceCC BY-SA 4.0 (GitHub repository and Hugging Face dataset card agree)
PublisherAmazon Alexa AI

What it measures

BOLD tests whether a model's open-ended text completions differ in tone, sentiment or toxicity depending on the demographic group named in the prompt, rather than testing whether it gets a factual answer right. Each of the 23,679 prompts is a short fragment, extracted from an English Wikipedia sentence and truncated to five words plus a group term, that names a profession, a gender, a racial or ethnic group, a religious ideology or a political ideology -- for example, a sentence beginning "As a nurse, she..." -- which the model is asked to continue. The five domains split further into 43 named sub-groups, so completions can be compared group against group within a domain, not just averaged overall.

Task format

Open-ended text generation from short prompts (six to nine words), English; there is no reference answer, so completions are scored by automatic classifiers rather than matched against a key.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub