DecodingTrust Toxicity Prompts

HELM scenario for DecodingTrust toxicity: continue RealToxicityPrompts-style prefixes and score PerspectiveAPI toxic_frac.

Also known as: DecodingTrust - Toxicity, DecodingTrustToxicityPromptsScenario

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorysafety
SubcategoryHELM wrap of DecodingTrust section 3 (RealToxicityPrompts continuations + GPT-filtered prompts)
Page statusunknown
Metrictoxic_frac
Directionlower_is_better
Dataset licenceCC-BY-SA-4.0
PublisherDecodingTrust authors (UIUC / Stanford / collaborators); HELM wrap by Stanford CRFM

What it measures

decodingtrust_toxicity_prompts is HELM's wrap of DecodingTrust section 3. The model continues an English prefix sampled from RealToxicityPrompts. HELM subjects are toxic, non-toxic, gpt-3.5, and gpt-4, each a jsonl under Data-Legacy user_prompts. The paper also studies 33 hand-written system prompts and jailbreaks; HELM's run spec is a completion adapter without those system strings. Scenario metadata and schema_decodingtrust.yaml still describe this group with privacy wording, which does not match the files.

Task format

HELM completion adapter: temperature 0.9, max_tokens 20, num_outputs 5. Comments in the run spec cite Gehman et al. (temperature 1, 20 tokens, k=25) and note that HELM caps k at 5 and does not use nucleus sampling.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub