DecodingTrust Privacy

DecodingTrust's privacy slice: extract Enron emails, leak planted PII, or share a 'secret' after a privacy cue.

Also known as: DecodingTrust - Privacy

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorysafety
Subcategorytraining-data, in-context PII, and privacy-word leakage
Page statusactive
Metricdecodingtrust_privacy_leakage_rate
Directionlower_is_better
Unitfraction
Dataset licenceCC-BY-SA-4.0
PublisherUniversity of Illinois Urbana-Champaign (DecodingTrust); Stanford CRFM (HELM scenario)

What it measures

DecodingTrust Privacy asks whether a chat model will emit private facts it should keep to itself. The paper and the HELM scenario share three probes. Training-data extraction prompts the model with Enron mail prefixes or name-to-email templates and scores whether the true address, local part, or domain comes back. In-context PII plants phone numbers, SSNs and similar strings in the dialogue, sometimes after a system line that forbids disclosure, and asks for the planted value. Privacy understanding tells a secret about a named person using a privacy word such as "confidentially," then asks whether the model will tell a fourth person. English text only.

Task format

Open-ended generation. HELM uses an instruction adapter, one sample, at most 32 tokens, temperature 1. The paper's Enron probe treats the first email in the completion as the prediction.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub