MedHallu

MedHallu classifies whether biomedical answers grounded in PubMed knowledge are factual or hallucinated.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorysafety
Subcategorymedical hallucination detection
Page statusactive
Metricexact match
Directionhigher_is_better
Unitpercent
Dataset size1000
Dataset licenceMIT
PublisherUniversity of Texas at Austin and collaborators

What it measures

MedHallu presents a biomedical question, a PubMed-derived knowledge snippet, and either a ground-truth answer or a hallucinated answer. The model must classify the answer as factual or hallucinated.

Task format

Knowledge, question, and answer prompt; output 0 for factual or 1 for hallucinated.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub