DecodingTrust Machine Ethics

HELM scenario for DecodingTrust ethics: wrong/not-wrong labels on ETHICS and Jiminy Cricket, with jailbreak and evasive filters.

Also known as: DecodingTrust - Ethics, DecodingTrustMachineEthicsScenario

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorysafety
SubcategoryHELM wrap of DecodingTrust section 9 (ETHICS commonsense + Jiminy Cricket)
Page statusunknown
Metricquasi_exact_match
Directionhigher_is_better
Dataset licenceCC-BY-SA-4.0
PublisherDecodingTrust authors (UIUC / Stanford / collaborators); HELM wrap by Stanford CRFM

What it measures

decodingtrust_machine_ethics is HELM's wrap of DecodingTrust section 9. The model reads an English morality vignette and must emit a closed label. Published HELM run entries use the ETHICS commonsense split (short 1–2 sentence items and long 1–6 paragraph posts) as wrong versus not wrong, and Jiminy Cricket text-adventure scenes as good, bad, or neutral. The paper also studies jailbreak prefixes and evasive add-ons that excuse harm. HELM's scenario class can load virtue, justice, deontology, utilitarianism, and Jiminy conditional-harm slices, but the shipped run-entry list does not.

Task format

HELM generation adapter, max_tokens 20. Run-entry strings set data_name, jailbreak_prompt, evasive_sentence, and sometimes max_train_instances (0, 3, 8, or 32) and max_eval_instances=200. Instructions and prefixes are per data_name (for example "Reaction: This is " on short commonsense).

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub