HELM scenario for DecodingTrust stereotypes: agree/disagree with 1,152 statements across 16 topics, 24 groups, and 3 system prompts.
unassessed
| Category | safety |
|---|---|
| Subcategory | HELM wrap of DecodingTrust section 4 (1,152 stereotype statements × 3 system prompts) |
| Page status | unknown |
| Metric | decodingtrust_stereotype_bias |
| Direction | higher_is_better |
| Dataset size | 3456 |
| Dataset licence | CC-BY-SA-4.0 |
| Publisher | DecodingTrust authors (UIUC / Stanford / collaborators); HELM wrap by Stanford CRFM |
decodingtrust_stereotype_bias is HELM's wrap of DecodingTrust section 4. Each English user prompt states a stereotype and tells the model to append agree or disagree. There are 16 topics (HIV, terrorists, drug addiction, intelligence, greed, parenting, country, technology, weakness, driving, crime, drug dealing, jobs, leadership, STEM, hygiene) and 24 demographic groups spanning race/ethnicity, gender/orientation, nationality, age, religion, disability, and socioeconomic status. Three system-prompt types — benign, untargeted jailbreak, and targeted jailbreak — are stored as tags. Every statement is a stereotype; the paper treats agreement as the biased outcome.
HELM instruct generation, num_outputs 25, max_tokens 150, temperature 1. The scenario loads stereotype_bias_data.jsonl and does not take a task argument (the run function still declares unused task: str).
No model card in ModelSpec reports this benchmark yet.