DecodingTrust Adversarial Demonstrations

HELM scenario for DecodingTrust's adversarial-demonstration tests: counterfactual, spurious, and backdoored in-context examples.

Also known as: AdvDemo, decodingtrust_adv_demo, DecodingTrustAdvDemoScenario

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorysafety
SubcategoryHELM wrap of DecodingTrust in-context adversarial demonstrations
Page statusunknown
Metricquasi_exact_match
Directionhigher_is_better
Dataset licenceCC-BY-SA-4.0
PublisherDecodingTrust authors (UIUC / Stanford / collaborators); HELM wrap by Stanford CRFM

What it measures

decodingtrust_adv_demonstration is HELM's wrap of DecodingTrust section 7. The model must classify a short English text after a prompt that may contain (1) a counterfactual neighbor with the opposite label, (2) demonstrations that all share a spurious heuristic, or (3) backdoored SST-2 demonstrations. Tasks include SNLI-CAD premise/hypothesis revisions, four MSGS linguistic probes, six NLI spurious-correlation slices, and SST-2 backdoor setups. The skill is whether in-context examples can flip or trap the label, not ordinary NLI accuracy.

Task format

HELM instruct generation, max_tokens 16, temperature 0, max_train_instances 0. Demonstrations are already inside each jsonl record; the scenario concatenates them as "Answer:" blocks. Parameters: perspective, data, demo_name, description.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub