A HELM scenario built from the 80 open-ended questions LMSYS used to evaluate Vicuna in 2023, scored in HELM by human critique ratings rather than the original GPT-4-judge protocol.
unassessed
| Category | instruction-following |
|---|---|
| Subcategory | open-ended instruction following (helpfulness rating on diverse prompts) |
| Page status | active |
| Metric | Human critique rating of response helpfulness (HELM); pairwise GPT-4-judge 1-10 score (original LMSYS protocol) |
| Direction | higher_is_better |
| Unit | rating |
| Dataset size | 80 |
| Dataset licence | Apache-2.0 |
| Publisher | LMSYS Org (source questions); Stanford CRFM (HELM implementation) |
The Vicuna scenario prompts a model with one of 80 open-ended instructions spanning nine categories (generic, knowledge, roleplay, common-sense, Fermi estimation, counterfactual, coding, math, and writing) and evaluates the quality of its free-form response. It measures general instruction-following and response helpfulness across a deliberately varied set of everyday and reasoning-style prompts, rather than a single narrow skill.
Zero-shot, open-ended generation: the model receives one instruction and produces a free-text response with no fixed answer to match. In HELM, responses are then rated by human annotators using a critique-based helpfulness metric, configurable by number of respondents. In the original LMSYS release, responses from multiple chatbots were instead compared pairwise and scored on a 1-10 scale by GPT-4 acting as a judge.
No model card in ModelSpec reports this benchmark yet.