HELM scenario for DecodingTrust's adversarial-demonstration tests: counterfactual, spurious, and backdoored in-context examples.
unassessed
| Category | safety |
|---|---|
| Subcategory | HELM wrap of DecodingTrust in-context adversarial demonstrations |
| Page status | unknown |
| Metric | quasi_exact_match |
| Direction | higher_is_better |
| Dataset licence | CC-BY-SA-4.0 |
| Publisher | DecodingTrust authors (UIUC / Stanford / collaborators); HELM wrap by Stanford CRFM |
decodingtrust_adv_demonstration is HELM's wrap of DecodingTrust section 7. The model must classify a short English text after a prompt that may contain (1) a counterfactual neighbor with the opposite label, (2) demonstrations that all share a spurious heuristic, or (3) backdoored SST-2 demonstrations. Tasks include SNLI-CAD premise/hypothesis revisions, four MSGS linguistic probes, six NLI spurious-correlation slices, and SST-2 backdoor setups. The skill is whether in-context examples can flip or trap the label, not ordinary NLI accuracy.
HELM instruct generation, max_tokens 16, temperature 0, max_train_instances 0. Demonstrations are already inside each jsonl record; the scenario concatenates them as "Answer:" blocks. Parameters: perspective, data, demo_name, description.
No model card in ModelSpec reports this benchmark yet.