MedHallu classifies whether biomedical answers grounded in PubMed knowledge are factual or hallucinated.
unassessed
| Category | safety |
|---|---|
| Subcategory | medical hallucination detection |
| Page status | active |
| Metric | exact match |
| Direction | higher_is_better |
| Unit | percent |
| Dataset size | 1000 |
| Dataset licence | MIT |
| Publisher | University of Texas at Austin and collaborators |
MedHallu presents a biomedical question, a PubMed-derived knowledge snippet, and either a ground-truth answer or a hallucinated answer. The model must classify the answer as factual or hallucinated.
Knowledge, question, and answer prompt; output 0 for factual or 1 for hallucinated.
No model card in ModelSpec reports this benchmark yet.