HELM MELT knowledge (ZaloE2E and ViMMRC)

HELM's Vietnamese knowledge track: closed-book answers on ZaloE2E and multiple-choice reading on ViMMRC, both scored by quasi-exact match.

Also known as: melt_knowledge_zalo, melt_knowledge_vimmrc, ZaloE2E, ViMMRC

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryknowledge
SubcategoryVietnamese closed-book QA and multiple-choice reading
Page statusunknown
Metricquasi_exact_match on each scenario's test split (no combined score)
Directionhigher_is_better
Dataset size1114
Dataset licenceura-hcmut/zalo_e2eqa card: MIT. ura-hcmut/ViMMRC card: CC-BY-NC-ND-4.0. The two scenarios do not share a licence.
PublisherStanford CRFM (HELM); dataset mirrors ura-hcmut; ZaloE2E from Zalo AI Challenge 2022

What it measures

MELT knowledge is two Vietnamese question sets behind one HELM filename, not one quiz. ZaloE2E is closed-book QA: the model sees a Vietnamese question and must write an answer without a passage. ViMMRC is multiple-choice reading: the model sees a Vietnamese article, a question, and lettered options, and must pick the gold option. HELM scores both with quasi-exact match on the test split. There is no official average of the two. The schema file's parent blurb "medical domain" does not describe these tasks.

Task format

ZaloE2E: open generation, instruction to answer from commonsense and to say "không có đáp án" if unknown, max 128 tokens. ViMMRC: joint multiple-choice (`ADAPT_MULTIPLE_CHOICE_JOINT`) with instruction "Sau đây là các câu hỏi trắc nghiệm (có đáp án)." An optional `randomize_order` flag shuffles choices; default run entries set it false.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub