HELM scenario for DecodingTrust toxicity: continue RealToxicityPrompts-style prefixes and score PerspectiveAPI toxic_frac.
unassessed
| Category | safety |
|---|---|
| Subcategory | HELM wrap of DecodingTrust section 3 (RealToxicityPrompts continuations + GPT-filtered prompts) |
| Page status | unknown |
| Metric | toxic_frac |
| Direction | lower_is_better |
| Dataset licence | CC-BY-SA-4.0 |
| Publisher | DecodingTrust authors (UIUC / Stanford / collaborators); HELM wrap by Stanford CRFM |
decodingtrust_toxicity_prompts is HELM's wrap of DecodingTrust section 3. The model continues an English prefix sampled from RealToxicityPrompts. HELM subjects are toxic, non-toxic, gpt-3.5, and gpt-4, each a jsonl under Data-Legacy user_prompts. The paper also studies 33 hand-written system prompts and jailbreaks; HELM's run spec is a completion adapter without those system strings. Scenario metadata and schema_decodingtrust.yaml still describe this group with privacy wording, which does not match the files.
HELM completion adapter: temperature 0.9, max_tokens 20, num_outputs 5. Comments in the run spec cite Gehman et al. (temperature 1, 20 tokens, k=25) and note that HELM caps k at 5 and does not use nucleus sampling.
No model card in ModelSpec reports this benchmark yet.