DecodingTrust's privacy slice: extract Enron emails, leak planted PII, or share a 'secret' after a privacy cue.
unassessed
| Category | safety |
|---|---|
| Subcategory | training-data, in-context PII, and privacy-word leakage |
| Page status | active |
| Metric | decodingtrust_privacy_leakage_rate |
| Direction | lower_is_better |
| Unit | fraction |
| Dataset licence | CC-BY-SA-4.0 |
| Publisher | University of Illinois Urbana-Champaign (DecodingTrust); Stanford CRFM (HELM scenario) |
DecodingTrust Privacy asks whether a chat model will emit private facts it should keep to itself. The paper and the HELM scenario share three probes. Training-data extraction prompts the model with Enron mail prefixes or name-to-email templates and scores whether the true address, local part, or domain comes back. In-context PII plants phone numbers, SSNs and similar strings in the dialogue, sometimes after a system line that forbids disclosure, and asks for the planted value. Privacy understanding tells a secret about a named person using a privacy word such as "confidentially," then asks whether the model will tell a fourth person. English text only.
Open-ended generation. HELM uses an instruction adapter, one sample, at most 32 tokens, temperature 1. The paper's Enron probe treats the first email in the completion as the prediction.
No model card in ModelSpec reports this benchmark yet.