Vicuna Questions

A HELM scenario built from the 80 open-ended questions LMSYS used to evaluate Vicuna in 2023, scored in HELM by human critique ratings rather than the original GPT-4-judge protocol.

Also known as: Vicuna-80

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryinstruction-following
Subcategoryopen-ended instruction following (helpfulness rating on diverse prompts)
Page statusactive
MetricHuman critique rating of response helpfulness (HELM); pairwise GPT-4-judge 1-10 score (original LMSYS protocol)
Directionhigher_is_better
Unitrating
Dataset size80
Dataset licenceApache-2.0
PublisherLMSYS Org (source questions); Stanford CRFM (HELM implementation)

What it measures

The Vicuna scenario prompts a model with one of 80 open-ended instructions spanning nine categories (generic, knowledge, roleplay, common-sense, Fermi estimation, counterfactual, coding, math, and writing) and evaluates the quality of its free-form response. It measures general instruction-following and response helpfulness across a deliberately varied set of everyday and reasoning-style prompts, rather than a single narrow skill.

Task format

Zero-shot, open-ended generation: the model receives one instruction and produces a free-text response with no fixed answer to match. In HELM, responses are then rated by human annotators using a critique-based helpfulness metric, configurable by number of respondents. In the original LMSYS release, responses from multiple chatbots were instead compared pairwise and scored on a 1-10 scale by GPT-4 acting as a judge.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub