Convince Me

A programmatic BIG-bench task that scores how far a model sways copies of itself toward 455 true and false TruthfulQA statements.

Also known as: BIG-bench convinceme, ConvinceMe

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorysafety
SubcategoryBIG-bench programmatic self-play persuasiveness on 455 TruthfulQA-derived statements
Page statusunknown
Metricoverall_difference
Directionhigher_is_better
Dataset size455
Dataset licenceApache-2.0
PublisherGoogle (BIG-bench collaboration)

What it measures

convinceme prompts one instance of a model (the debater) to argue that a statement is true, then polls other instances of the same model (the jury) before and after. Statements are true and false answers drawn from five TruthfulQA categories, not the TruthfulQA questions. The skill is ability to produce convincing arguments, including for falsehoods, not whether the model volunteers a lie unprompted. Authors: Cedrick Argueta, Vinay Ramasesh, and Jaime Fernández Fisac. English text. Programmatic task.py.

Task format

Five questionnaire items per statement (two Likert, three yes/no or true/false), ten jury samples by default, five generated persuasion paragraphs. Preferred score overall_difference (signed Jensen-Shannon change, negated so higher is better). Secondary overall_alpha_avg (Cronbach's alpha). Canary GUID embedded.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub