HELM scenario for DecodingTrust ethics: wrong/not-wrong labels on ETHICS and Jiminy Cricket, with jailbreak and evasive filters.
unassessed
| Category | safety |
|---|---|
| Subcategory | HELM wrap of DecodingTrust section 9 (ETHICS commonsense + Jiminy Cricket) |
| Page status | unknown |
| Metric | quasi_exact_match |
| Direction | higher_is_better |
| Dataset licence | CC-BY-SA-4.0 |
| Publisher | DecodingTrust authors (UIUC / Stanford / collaborators); HELM wrap by Stanford CRFM |
decodingtrust_machine_ethics is HELM's wrap of DecodingTrust section 9. The model reads an English morality vignette and must emit a closed label. Published HELM run entries use the ETHICS commonsense split (short 1–2 sentence items and long 1–6 paragraph posts) as wrong versus not wrong, and Jiminy Cricket text-adventure scenes as good, bad, or neutral. The paper also studies jailbreak prefixes and evasive add-ons that excuse harm. HELM's scenario class can load virtue, justice, deontology, utilitarianism, and Jiminy conditional-harm slices, but the shipped run-entry list does not.
HELM generation adapter, max_tokens 20. Run-entry strings set data_name, jailbreak_prompt, evasive_sentence, and sometimes max_train_instances (0, 3, 8, or 32) and max_eval_instances=200. Instructions and prefixes are per data_name (for example "Reaction: This is " on short commonsense).
No model card in ModelSpec reports this benchmark yet.