23,679 Wikipedia-derived prompts across five demographic domains, used to check whether a model's open-ended completions differ in sentiment, toxicity or regard depending on the group named in the prompt.
unassessed
| Category | safety |
|---|---|
| Subcategory | bias in open-ended text generation, read as gaps between demographic groups rather than one score |
| Page status | active |
| Metric | sentiment, toxicity, regard, gender polarity and psycholinguistic norms, compared across demographic groups |
| Direction | higher_is_better |
| Dataset size | 23679 |
| Dataset licence | CC BY-SA 4.0 (GitHub repository and Hugging Face dataset card agree) |
| Publisher | Amazon Alexa AI |
BOLD tests whether a model's open-ended text completions differ in tone, sentiment or toxicity depending on the demographic group named in the prompt, rather than testing whether it gets a factual answer right. Each of the 23,679 prompts is a short fragment, extracted from an English Wikipedia sentence and truncated to five words plus a group term, that names a profession, a gender, a racial or ethnic group, a religious ideology or a political ideology -- for example, a sentence beginning "As a nurse, she..." -- which the model is asked to continue. The five domains split further into 43 named sub-groups, so completions can be compared group against group within a domain, not just averaged overall.
Open-ended text generation from short prompts (six to nine words), English; there is no reference answer, so completions are scored by automatic classifiers rather than matched against a key.
No model card in ModelSpec reports this benchmark yet.