HELM scenario for DecodingTrust AdvGLUE++: GLUE items rewritten against Alpaca, Vicuna, and StableVicuna.
unassessed
| Category | safety |
|---|---|
| Subcategory | HELM wrap of DecodingTrust AdvGLUE++ GLUE adversarial classification |
| Page status | unknown |
| Metric | quasi_exact_match |
| Direction | higher_is_better |
| Dataset size | 42017 |
| Dataset licence | CC-BY-SA-4.0 |
| Publisher | DecodingTrust authors (UIUC / Stanford / collaborators); HELM wrap by Stanford CRFM |
decodingtrust_adv_robustness is HELM's wrap of DecodingTrust section 5.2. Each item is an English GLUE classification example whose text was perturbed to fool Alpaca-7B, Vicuna-13B or StableVicuna-13B, then transferred to the model under test. Tasks are SST-2, QQP, MNLI, MNLI-mismatched, QNLI and RTE. HELM's scenario class is named decodingtrust_adv_glue_plus_plus; metadata and the run spec use decodingtrust_adv_robustness. The original AdvGLUE (BERT-era) split is documented in the paper but is not implemented in HELM and is not loaded.
HELM instruct generation, max_tokens 16, temperature 0, one completion. The prompt is a task instruction plus labelled fields (sentence, premise/hypothesis, question pairs). Parameter glue_task selects one GLUE slice.
No model card in ModelSpec reports this benchmark yet.