RoleBench

A 168,093-sample role-play benchmark over 100 characters; OpenCompass scores English and Chinese splits with ROUGE against RoleGPT-style references.

Also known as: RoleLLM, RoleBench (RoleLLM)

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorygeneration
Subcategorycharacter-level role-play: instruction and role generalization, English and Chinese
Page statusactive
MetricROUGE-L (OpenCompass RougeEvaluator / JiebaRougeEvaluator); paper also reports GPT-3.5 judgments
Directionhigher_is_better
Dataset size168093
Dataset licenceApache-2.0
PublisherInteractiveNLP-Team (RoleLLM)

What it measures

RoleBench tests whether a model can stay in character. Each item names a role, gives a short profile, and asks a question. The model must answer in that voice without breaking character. The paper splits evaluation two ways: instruction generalization (held-out questions for seen roles) and role generalization (held-out English roles). Profiles cover 95 English characters and 5 Chinese ones (100 roles). References come from RoleGPT (GPT-4 role prompting) and Context-Instruct, not from human scripts of the same replies. OpenCompass runs three of those splits as generation tasks.

Task format

Chat-style generation: a system prompt that assigns the role and description, then a user question. OpenCompass is zero-shot with max_out_len 512. English uses RougeEvaluator; Chinese uses JiebaRougeEvaluator.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub