104 real, historical (2017) consumer health questions from the TREC LiveQA medical track; despite the "live" name this is a static, fixed test set, not a continuously refreshed one.
unassessed
| Category | domain |
|---|---|
| Subcategory | consumer health / medical question answering, free-form generation graded by an LLM judge against reference answers |
| Page status | unknown |
| Metric | LLM-as-judge score (0, 0.3, 0.7 or 1 per response, averaged across the test set as live_qa_score), plus HELM's standard open-ended-generation reference metrics |
| Direction | higher_is_better |
| Unit | points |
| Dataset size | 104 |
| Dataset licence | CC BY 4.0, per the GitHub repository. |
| Publisher | U.S. National Library of Medicine (NLM); Emory University; TREC (Text REtrieval Conference) |
Despite what "live" suggests, live_qa is not a continuously refreshed dataset -- it is HELM's implementation of a single, fixed historical benchmark: the TREC-2017 LiveQA Medical Task, a one-time consumer-health question-answering track organised with the U.S. National Library of Medicine (NLM). The name describes the provenance of the questions, not the freshness of the benchmark: the underlying questions are genuinely live in the sense that they were real, real-time queries submitted by members of the public to NLM's consumer health inquiry service, rather than written specifically for the benchmark. Once the TREC 2017 track concluded, the organisers released a fixed 104-question test set (alongside a separately released, larger training set) with reference answers vetted by medical experts, and that fixed archive is exactly what HELM's live_qa scenario downloads and evaluates against today -- it has not been expanded, refreshed, or replaced with newer questions since.
A real consumer health question (for example, "What is the relationship between Noonan syndrome and polycystic renal disease?") in; a free-form generated answer out, produced zero-shot with no in-context examples.
No model card in ModelSpec reports this benchmark yet.