HELM wrap of HealthQA-BR: 5,632 Portuguese multiple-choice items from Brazilian medical licensing and residency exams.
unassessed
| Category | domain |
|---|---|
| Subcategory | Brazilian public-health multiple-choice exams (Revalida and Enare) |
| Page status | active |
| Metric | exact_match |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 5632 |
| Dataset licence | CC-BY-4.0 |
| Publisher | Andrew Maranhão Ventura D'addario (dataset); Stanford CRFM (HELM scenario) |
healthqa_br is Stanford CRFM HELM's scenario over HealthQA-BR, a Portuguese multiple-choice set of Brazilian public-health exams. The model reads a question plus lettered options and must select the correct alternative. Items come from Revalida (revalidation of foreign medical diplomas) and Enare (national residency, medical and multiprofessional). Unlike USMLE-style English medical quizzes, the set covers medicine and allied professions in the Sistema Único de Saúde (nursing, dentistry, psychology, social work, pharmacy, physiotherapy, and others). This id is HELM's wrap, not the authors' own letter-only generation script, and not [healthbench](healthbench.md) or [medqa](medqa.md).
HELM multiple-choice joint adaptation (ADAPT_MULTIPLE_CHOICE_JOINT). Portuguese instructions include one worked insulin/pancreas example and ask for a letter only. Input noun Pergunta, output noun Resposta. Default adapter max_train_instances is 5 and max_tokens is 1, but the scenario labels every instance TEST_SPLIT, so the realized shot count is not established from these files. The paper instead used zero-shot generation of a single letter A–E.
No model card in ModelSpec reports this benchmark yet.