HELM MELT synthetic reasoning (natural language)

HELM's Vietnamese synthetic-reasoning (natural) task: deduce attributes from generated Vietnamese rules and facts, scored by set-overlap F1.

Also known as: melt_synthetic_reasoning_natural, MELT SRN

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryreasoning
SubcategoryVietnamese synthetic rule-fact deduction
Page statusunknown
Metricf1_set_match
Directionhigher_is_better
Dataset size5000
PublisherStanford CRFM (HELM MELT scenarios)

What it measures

MELT SRN generates Vietnamese rule-and-fact puzzles and asks the model to list what else must be true of a subject. Each item is a set of "if … then …" rules, one fact about a person, animal or plant, and a prompt asking what can be determined about that subject. The target is the set of attributes implied in one reasoning step. Easy mode always uses the specific subject (for example a name rather than "một người"). Hard mode substitutes more specific synonyms in the fact. The puzzles are synthetic Vietnamese, not mined text.

Task format

Open generation. HELM instruction "Hãy giải quyết vấn đề sau.", input noun "Quy luật", three in-context examples, max 20 tokens. Difficulty is a run-spec argument: easy or hard. Main metric is f1_set_match on the generated test split.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub