Model-Written Evaluations

154 model-generated, human-filtered yes/no datasets probing a model's persona, sycophancy and advanced-AI-risk tendencies rather than testing right-or-wrong knowledge.

Also known as: Discovering Language Model Behaviors with Model-Written Evaluations, Anthropic Model-Written Evals, MWE

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorysafety
Subcategorymodel persona, sycophancy and advanced-AI-risk behavioural probes
Page statusactive
Metric% of answers matching the tested behaviour
Directionhigher_is_better
Unit%
Dataset size154
Dataset licenceCC BY 4.0
PublisherAnthropic (Surge AI and the Machine Intelligence Research Institute credited for human-generated comparison data)

What it measures

Model-Written Evaluations (MWE) is not one test but a method plus its output: Anthropic used language models to write large sets of yes/no and A/B questions designed to reveal how a model behaves along a given trait, then had crowdworkers filter and validate the results. The released collection spans four areas: persona (does the model's stated personality, politics, religion, ethics, or desire to pursue goals like power or self-preservation match a given description), sycophancy (does the model echo a user's stated opinion on philosophy, NLP research, or politics rather than giving an independent answer), advanced AI risk (does the model express tendencies such as corrigibility, coordination with other AI instances, or awareness of its own situation), and Winogenerated, a model-generated, human-validated expansion of the Winogender gender-bias schema. Every item is a forced-choice question with one answer marked as "matching" the behaviour under test and one marked as "not matching"; there is no objectively correct answer.

Task format

Binary or A/B forced-choice questions. The model (or its next-token probabilities) is scored on whether it selects the answer_matching_behavior or answer_not_matching_behavior option for each item; most implementations compare the log-likelihood the model assigns to each labelled continuation rather than requiring free-text generation.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub