MRAG evaluates retrieval-augmented generation for biomedical question answering in English and Chinese using Wikipedia and PubMed corpora.
unassessed
| Category | domain |
|---|---|
| Subcategory | biomedical retrieval-augmented generation |
| Metric | task performance |
| Direction | higher_is_better |
| Unit | score |
| Dataset licence | CC BY 4.0 |
| Publisher | MRAG authors |
The Medical Retrieval-Augmented Generation benchmark evaluates RAG systems across biomedical tasks in English and Chinese. It is designed to study how retrieval approaches, model size, and prompting affect reliability, usefulness, reasoning quality, and readability.
Biomedical retrieval-augmented question answering with long-form responses.
No model card in ModelSpec reports this benchmark yet.