DecodingTrust OoD Robustness

HELM scenario for DecodingTrust out-of-distribution robustness: style-shifted SST-2 sentiment and RealtimeQA knowledge with optional refusal.

Also known as: DecodingTrust - OoD Robustness, DecodingTrustOODRobustnessScenario, decodingtrust_ood

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorysafety
SubcategoryHELM wrap of DecodingTrust section 6 (SST-2 styles + RealtimeQA knowledge)
Page statusunknown
Metricquasi_exact_match (style); ood_knowledge_rr / ood_knowledge_macc (knowledge)
Directionhigher_is_better
Dataset licenceCC-BY-SA-4.0
PublisherDecodingTrust authors (UIUC / Stanford / collaborators); HELM wrap by Stanford CRFM

What it measures

decodingtrust_ood_robustness is HELM's wrap of DecodingTrust section 6. Style runs take an English SST-2 sentence after a style transform (Shakespeare word/sentence, Bible, romantic, tweet, augmentation, or the untransformed base) and require exactly positive or negative. Knowledge runs prepend "Today is {date}" to a RealtimeQA item (qa_2020 or qa_2023) and ask for a multiple-choice letter; an I-don't-know option can be appended. Few-shot style uses eight demonstrations; knowledge uses five, including MMLU-topic demos (global facts, machine learning, moral scenarios, US foreign policy) or 2021 QA.

Task format

Style: HELM instruct or 8-shot instruct, max_tokens 16, temperature 0. Knowledge: joint multiple-choice adapter, max_tokens 16, max_train_instances 5. Parameters: ood_type, task, demo_name, run_id, idk.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub