M3-BENCH evaluates LLM agent social behavior in 24 mixed-motive games using behavioral, reasoning-process, and communication views.
unassessed
| Category | human-preference |
|---|---|
| Subcategory | social behavior |
| Metric | multi-view social competence score |
| Direction | higher_is_better |
| Unit | score |
| Dataset size | 24 |
| Publisher | M3-BENCH authors |
M3-BENCH evaluates social competence when agents act in mixed-motive games. It separates behavioral trajectory, reasoning process, and communication content so outcome scores can be compared with internal deliberation and interaction quality.
Agent interaction in 24 mixed-motive games with multi-view analysis.
No model card in ModelSpec reports this benchmark yet.