DecodingTrust Stereotype Bias

HELM scenario for DecodingTrust stereotypes: agree/disagree with 1,152 statements across 16 topics, 24 groups, and 3 system prompts.

Also known as: DecodingTrust - Stereotype Bias, DecodingTrustStereotypeBiasScenario

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorysafety
SubcategoryHELM wrap of DecodingTrust section 4 (1,152 stereotype statements × 3 system prompts)
Page statusunknown
Metricdecodingtrust_stereotype_bias
Directionhigher_is_better
Dataset size3456
Dataset licenceCC-BY-SA-4.0
PublisherDecodingTrust authors (UIUC / Stanford / collaborators); HELM wrap by Stanford CRFM

What it measures

decodingtrust_stereotype_bias is HELM's wrap of DecodingTrust section 4. Each English user prompt states a stereotype and tells the model to append agree or disagree. There are 16 topics (HIV, terrorists, drug addiction, intelligence, greed, parenting, country, technology, weakness, driving, crime, drug dealing, jobs, leadership, STEM, hygiene) and 24 demographic groups spanning race/ethnicity, gender/orientation, nationality, age, religion, disability, and socioeconomic status. Three system-prompt types — benign, untargeted jailbreak, and targeted jailbreak — are stored as tags. Every statement is a stereotype; the paper treats agreement as the biased outcome.

Task format

HELM instruct generation, num_outputs 25, max_tokens 150, temperature 1. The scenario loads stereotype_bias_data.jsonl and does not take a task argument (the run function still declares unused task: str).

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub