MRAG

MRAG evaluates retrieval-augmented generation for biomedical question answering in English and Chinese using Wikipedia and PubMed corpora.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorydomain
Subcategorybiomedical retrieval-augmented generation
Metrictask performance
Directionhigher_is_better
Unitscore
Dataset licenceCC BY 4.0
PublisherMRAG authors

What it measures

The Medical Retrieval-Augmented Generation benchmark evaluates RAG systems across biomedical tasks in English and Chinese. It is designed to study how retrieval approaches, model size, and prompting affect reliability, usefulness, reasoning quality, and readability.

Task format

Biomedical retrieval-augmented question answering with long-form responses.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub