154 model-generated, human-filtered yes/no datasets probing a model's persona, sycophancy and advanced-AI-risk tendencies rather than testing right-or-wrong knowledge.
unassessed
| Category | safety |
|---|---|
| Subcategory | model persona, sycophancy and advanced-AI-risk behavioural probes |
| Page status | active |
| Metric | % of answers matching the tested behaviour |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 154 |
| Dataset licence | CC BY 4.0 |
| Publisher | Anthropic (Surge AI and the Machine Intelligence Research Institute credited for human-generated comparison data) |
Model-Written Evaluations (MWE) is not one test but a method plus its output: Anthropic used language models to write large sets of yes/no and A/B questions designed to reveal how a model behaves along a given trait, then had crowdworkers filter and validate the results. The released collection spans four areas: persona (does the model's stated personality, politics, religion, ethics, or desire to pursue goals like power or self-preservation match a given description), sycophancy (does the model echo a user's stated opinion on philosophy, NLP research, or politics rather than giving an independent answer), advanced AI risk (does the model express tendencies such as corrigibility, coordination with other AI instances, or awareness of its own situation), and Winogenerated, a model-generated, human-validated expansion of the Winogender gender-bias schema. Every item is a forced-choice question with one answer marked as "matching" the behaviour under test and one marked as "not matching"; there is no objectively correct answer.
Binary or A/B forced-choice questions. The model (or its next-token probabilities) is scored on whether it selects the answer_matching_behavior or answer_not_matching_behavior option for each item; most implementations compare the log-likelihood the model assigns to each labelled continuation rather than requiring free-text generation.
No model card in ModelSpec reports this benchmark yet.