Grammar / Best ChatGPT Prompts (HELM Instruct)

HELM Instruct scenario that expands a CFG of Gridfiti ChatGPT prompts and scores free-form replies with a 1-5 Helpfulness critique.

Also known as: Best ChatGPT Prompts, HELM grammar, grammar:path, best_chatgpt_prompts

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryinstruction-following
SubcategoryHELM Instruct CFG expansion of Gridfiti ChatGPT prompts, scored by Helpfulness
Page statusunknown
MetricHelpfulness
Directionhigher_is_better
Unit1-5
Dataset size320
Dataset licenceApache-2.0
PublisherStanford CRFM (HELM Instruct); prompt list attributed to Gridfiti Staff

What it measures

grammar, as this id, is not linguistic acceptability and not [glue_cola](glue_cola.md) or [blimp](blimp.md). It is Stanford HELM Instruct's GrammarScenario: a context-free grammar expands into English user prompts, the model writes a free-form reply, and a critique metric rates Helpfulness. The checked-in grammar best_chatgpt_prompts.yaml restates GRIDFITI's 2023 "best ChatGPT prompts" list, with slots such as language, city, and event. The HELM schema file says the group should have been named best_chatgpt_prompts, but results use grammar.

Task format

Zero-shot generation via get_instruct_adapter_spec (max_tokens 512, temperature 0.7, max_train_instances 0). References are empty. Run spec grammar:path=<yaml>,tags=<csv> keeps only derivations whose collected tags include every requested tag. Scoring is InstructionFollowingCritiqueMetric.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub