LIT-RAGBench evaluates retrieval-augmented generation generators on 114 human-constructed Japanese questions and a curated English version across integration, reasoning, logic, table and abstention capabilities.
unassessed
| Category | reasoning |
|---|---|
| Subcategory | retrieval-augmented generation answer evaluation |
| Page status | active |
| Metric | LLM-as-a-Judge accuracy |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 114 |
| Publisher | LIT-RAGBench authors |
LIT-RAGBench tests whether a language model can use supplied retrieved context to generate an answer, including recognizing when it should abstain. Its five categories are Integration, Reasoning, Logic, Table, and Abstention, with Japanese questions and an English machine-translated and human-curated version.
Context-grounded question answering with category-specific answer generation and abstention cases.
No model card in ModelSpec reports this benchmark yet.