HealthQA-BR (HELM)

HELM wrap of HealthQA-BR: 5,632 Portuguese multiple-choice items from Brazilian medical licensing and residency exams.

Also known as: HealthQA-BR, healthqa-br, Larxel/healthqa-br

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorydomain
SubcategoryBrazilian public-health multiple-choice exams (Revalida and Enare)
Page statusactive
Metricexact_match
Directionhigher_is_better
Unit%
Dataset size5632
Dataset licenceCC-BY-4.0
PublisherAndrew Maranhão Ventura D'addario (dataset); Stanford CRFM (HELM scenario)

What it measures

healthqa_br is Stanford CRFM HELM's scenario over HealthQA-BR, a Portuguese multiple-choice set of Brazilian public-health exams. The model reads a question plus lettered options and must select the correct alternative. Items come from Revalida (revalidation of foreign medical diplomas) and Enare (national residency, medical and multiprofessional). Unlike USMLE-style English medical quizzes, the set covers medicine and allied professions in the Sistema Único de Saúde (nursing, dentistry, psychology, social work, pharmacy, physiotherapy, and others). This id is HELM's wrap, not the authors' own letter-only generation script, and not [healthbench](healthbench.md) or [medqa](medqa.md).

Task format

HELM multiple-choice joint adaptation (ADAPT_MULTIPLE_CHOICE_JOINT). Portuguese instructions include one worked insulin/pancreas example and ask for a letter only. Input noun Pergunta, output noun Resposta. Default adapter max_train_instances is 5 and max_tokens is 1, but the scenario labels every instance TEST_SPLIT, so the realized shot count is not established from these files. The paper instead used zero-shot generation of a single letter A–E.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub