5,621 closed multiple-choice healthcare questions from Spain's 2020-2024 specialised exams, in English and Spanish, plus a 2,769-item open-ended English variant scored by a new metric.
unassessed
| Category | domain |
|---|---|
| Subcategory | Spanish specialised healthcare training exam question answering (closed and open-ended) |
| Page status | active |
| Metric | accuracy (closed version); Relaxed Perplexity and other text-generation metrics (open version) |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 5621 |
| Dataset licence | Apache License 2.0, per the Hugging Face dataset card |
| Publisher | Barcelona Supercomputing Center (BSC), High Performance Artificial Intelligence (HPAI) research group |
CareQA tests healthcare knowledge across six professional categories -- medicine, nursing, biology, chemistry, psychology and pharmacology -- sourced from Spain's official specialised healthcare training exams for the 2020-2024 editions. The paper's own text calls the source "the Spanish Specialised Healthcare Training (MIR) exams," while the Hugging Face dataset card instead names the source "Specialized Healthcare Training (FSE) examinations"; MIR (Medico Interno Residente) is properly the medicine-only track of Spain's broader FSE exam system, which also covers the other five professions the dataset draws from, so the paper's own shorthand undersells how many professional tracks are actually represented. Original items are in Spanish; an English version was produced by GPT-4 translation. CareQA ships in two forms: a closed multiple-choice version (English and Spanish) and an open-ended free-response version (English only), built by rephrasing closed questions with Qwen2.5-72B-Instruct and validating the results by hand. It was built to re-check the health of existing medical benchmarks and to study how open-ended and closed-ended healthcare evaluation relate to each other, and has separately been used to evaluate the Aloe family of healthcare LLMs from the same research group.
Two forms: (1) four-option multiple-choice, single letter answer, in English or Spanish; (2) open-ended free-text response to the same underlying clinical/scientific question, rephrased away from multiple-choice framing, English only, graded by automatic text-similarity metrics rather than exact match.
No model card in ModelSpec reports this benchmark yet.