M3-BENCH

M3-BENCH evaluates LLM agent social behavior in 24 mixed-motive games using behavioral, reasoning-process, and communication views.

Also known as: M3-Bench

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryhuman-preference
Subcategorysocial behavior
Metricmulti-view social competence score
Directionhigher_is_better
Unitscore
Dataset size24
PublisherM3-BENCH authors

What it measures

M3-BENCH evaluates social competence when agents act in mixed-motive games. It separates behavioral trajectory, reasoning process, and communication content so outcome scores can be compared with internal deliberation and interaction quality.

Task format

Agent interaction in 24 mixed-motive games with multi-view analysis.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub