LIT-RAGBench

LIT-RAGBench evaluates retrieval-augmented generation generators on 114 human-constructed Japanese questions and a curated English version across integration, reasoning, logic, table and abstention capabilities.

Also known as: LIT-RAGBench: Benchmarking Generator Capabilities of Large Language Models in Retrieval-Augmented Generation

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryreasoning
Subcategoryretrieval-augmented generation answer evaluation
Page statusactive
MetricLLM-as-a-Judge accuracy
Directionhigher_is_better
Unit%
Dataset size114
PublisherLIT-RAGBench authors

What it measures

LIT-RAGBench tests whether a language model can use supplied retrieved context to generate an answer, including recognizing when it should abstain. Its five categories are Integration, Reasoning, Logic, Table, and Abstention, with Japanese questions and an English machine-translated and human-curated version.

Task format

Context-grounded question answering with category-specific answer generation and abstention cases.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub