LiveQA (TREC-2017 Medical Task)

104 real, historical (2017) consumer health questions from the TREC LiveQA medical track; despite the "live" name this is a static, fixed test set, not a continuously refreshed one.

Also known as: TREC-2017 LiveQA: Medical Question Answering Task, LiveQA'17 Medical

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorydomain
Subcategoryconsumer health / medical question answering, free-form generation graded by an LLM judge against reference answers
Page statusunknown
MetricLLM-as-judge score (0, 0.3, 0.7 or 1 per response, averaged across the test set as live_qa_score), plus HELM's standard open-ended-generation reference metrics
Directionhigher_is_better
Unitpoints
Dataset size104
Dataset licenceCC BY 4.0, per the GitHub repository.
PublisherU.S. National Library of Medicine (NLM); Emory University; TREC (Text REtrieval Conference)

What it measures

Despite what "live" suggests, live_qa is not a continuously refreshed dataset -- it is HELM's implementation of a single, fixed historical benchmark: the TREC-2017 LiveQA Medical Task, a one-time consumer-health question-answering track organised with the U.S. National Library of Medicine (NLM). The name describes the provenance of the questions, not the freshness of the benchmark: the underlying questions are genuinely live in the sense that they were real, real-time queries submitted by members of the public to NLM's consumer health inquiry service, rather than written specifically for the benchmark. Once the TREC 2017 track concluded, the organisers released a fixed 104-question test set (alongside a separately released, larger training set) with reference answers vetted by medical experts, and that fixed archive is exactly what HELM's live_qa scenario downloads and evaluates against today -- it has not been expanded, refreshed, or replaced with newer questions since.

Task format

A real consumer health question (for example, "What is the relationship between Noonan syndrome and polycystic renal disease?") in; a free-form generated answer out, produced zero-shot with no in-context examples.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub