japanese_leaderboard is an lm-evaluation-harness group running eight independent Japanese NLP tasks -- JAQKET, four JGLUE tasks, MGSM, XL-Sum and XWinograd -- with no combined score computed.
unassessed
| Category | composite |
|---|---|
| Subcategory | lm-evaluation-harness task group running eight independently-authored Japanese NLP tasks, reported without a combined score |
| Page status | active |
| Metric | no combined metric: each of the eight component tasks reports its own metric (exact_match, acc, or a custom ROUGE-2 aggregation) independently |
| Direction | higher_is_better |
| Unit | % |
| Dataset licence | Varies by component dataset (JGLUE, JAQKET, MGSM, XL-Sum and XWinograd each carry their own terms); not established as a single licence for this page. |
| Publisher | lm-evaluation-harness (EleutherAI); the task group's README credits its prompts to Stability AI's Japanese fork of the harness |
japanese_leaderboard is not one benchmark but an lm-evaluation-harness task group: running it executes eight separately-authored Japanese evaluation tasks and reports each one's own metric, not a single blended score. Four come from JGLUE, Japan's general language-understanding benchmark suite (JCommonsenseQA, JNLI, JSQuAD, MARC-ja); the other four are JAQKET v2 (Wikipedia-grounded quiz question answering), the Japanese subset of MGSM (grade-school maths word problems), the Japanese subset of XL-Sum (news summarisation), and the Japanese subset of XWinograd (commonsense pronoun resolution). A model's results under this group name are eight separate numbers on different scales -- exact match, classification accuracy and ROUGE-2 among them -- not one measurement of "Japanese ability."
A group of eight independently-formatted tasks. JCommonsenseQA, JNLI, MARC-ja and XWinograd are multiple-choice or classification tasks scored by comparing log-likelihoods over labelled options. JAQKET v2 and JSQuAD are extractive question answering scored by exact match against generated text. MGSM is a generated chain-of-reasoning math answer scored by extracting a final number. XL-Sum is free-form generated summarisation scored by ROUGE-2. Few-shot counts also vary by component, from zero (XWinograd) to five (MGSM); no single prompt format or shot count applies across the group.
No model card in ModelSpec reports this benchmark yet.