Anthropic HH-RLHF (HELM Instruct)

HELM Instruct scores first human utterances from Anthropic HH-RLHF with a 1-5 Helpfulness critique; it does not train on or rank the chosen/rejected pairs.

Also known as: HH-RLHF, Anthropic RLHF dataset, anthropic_hh_rlhf:subset=hh, anthropic_hh_rlhf:subset=red_team

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryinstruction-following
SubcategoryHELM Instruct critique of first-turn prompts from Anthropic HH-RLHF
Page statusunknown
MetricHelpfulness
Directionhigher_is_better
Unit1-5
Dataset licenceMIT
PublisherAnthropic (dataset); Stanford CRFM (HELM Instruct scenario)

What it measures

anthropic_hh_rlhf, as this id, is Stanford CRFM's HELM Instruct scenario over Anthropic's public HH-RLHF dialogues. The model sees only the first human utterance of a conversation and must write a free-form English reply. HELM does not present the assistant turns, does not use chosen versus rejected labels as the score, and does not run RLHF. Two subsets exist: `hh` (helpfulness/harmlessness preference dialogues) and `red_team` (red-team transcripts). The skill is following that first request, under a human critique, not preference-model accuracy.

Task format

Zero-shot generation via get_instruct_adapter_spec (max_tokens 512, temperature 0.7, max_train_instances 0). References are empty. Scoring is InstructionFollowingCritiqueMetric with caller-chosen num_respondents. Official classic run entries use num_respondents=1.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub