DecodingTrust Adversarial Robustness (AdvGLUE++)

HELM scenario for DecodingTrust AdvGLUE++: GLUE items rewritten against Alpaca, Vicuna, and StableVicuna.

Also known as: AdvGLUE++, decodingtrust_adv_glue_plus_plus, DecodingTrustAdvRobustnessScenario

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorysafety
SubcategoryHELM wrap of DecodingTrust AdvGLUE++ GLUE adversarial classification
Page statusunknown
Metricquasi_exact_match
Directionhigher_is_better
Dataset size42017
Dataset licenceCC-BY-SA-4.0
PublisherDecodingTrust authors (UIUC / Stanford / collaborators); HELM wrap by Stanford CRFM

What it measures

decodingtrust_adv_robustness is HELM's wrap of DecodingTrust section 5.2. Each item is an English GLUE classification example whose text was perturbed to fool Alpaca-7B, Vicuna-13B or StableVicuna-13B, then transferred to the model under test. Tasks are SST-2, QQP, MNLI, MNLI-mismatched, QNLI and RTE. HELM's scenario class is named decodingtrust_adv_glue_plus_plus; metadata and the run spec use decodingtrust_adv_robustness. The original AdvGLUE (BERT-era) split is documented in the paper but is not implemented in HELM and is not loaded.

Task format

HELM instruct generation, max_tokens 16, temperature 0, one completion. The prompt is a task instruction plus labelled fields (sentence, premise/hypothesis, question pairs). Parameter glue_task selects one GLUE slice.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub