A programmatic BIG-bench task that scores how far a model sways copies of itself toward 455 true and false TruthfulQA statements.
unassessed
| Category | safety |
|---|---|
| Subcategory | BIG-bench programmatic self-play persuasiveness on 455 TruthfulQA-derived statements |
| Page status | unknown |
| Metric | overall_difference |
| Direction | higher_is_better |
| Dataset size | 455 |
| Dataset licence | Apache-2.0 |
| Publisher | Google (BIG-bench collaboration) |
convinceme prompts one instance of a model (the debater) to argue that a statement is true, then polls other instances of the same model (the jury) before and after. Statements are true and false answers drawn from five TruthfulQA categories, not the TruthfulQA questions. The skill is ability to produce convincing arguments, including for falsehoods, not whether the model volunteers a lie unprompted. Authors: Cedrick Argueta, Vinay Ramasesh, and Jaime Fernández Fisac. English text. Programmatic task.py.
Five questionnaire items per statement (two Likert, three yes/no or true/false), ten jury samples by default, five generated persuasion paragraphs. Preferred score overall_difference (signed Jensen-Shannon change, negated so higher is better). Secondary overall_alpha_avg (Cronbach's alpha). Canary GUID embedded.
No model card in ModelSpec reports this benchmark yet.