Japanese Leaderboard (lm-evaluation-harness)

japanese_leaderboard is an lm-evaluation-harness group running eight independent Japanese NLP tasks -- JAQKET, four JGLUE tasks, MGSM, XL-Sum and XWinograd -- with no combined score computed.

Also known as: ja_leaderboard

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorycomposite
Subcategorylm-evaluation-harness task group running eight independently-authored Japanese NLP tasks, reported without a combined score
Page statusactive
Metricno combined metric: each of the eight component tasks reports its own metric (exact_match, acc, or a custom ROUGE-2 aggregation) independently
Directionhigher_is_better
Unit%
Dataset licenceVaries by component dataset (JGLUE, JAQKET, MGSM, XL-Sum and XWinograd each carry their own terms); not established as a single licence for this page.
Publisherlm-evaluation-harness (EleutherAI); the task group's README credits its prompts to Stability AI's Japanese fork of the harness

What it measures

japanese_leaderboard is not one benchmark but an lm-evaluation-harness task group: running it executes eight separately-authored Japanese evaluation tasks and reports each one's own metric, not a single blended score. Four come from JGLUE, Japan's general language-understanding benchmark suite (JCommonsenseQA, JNLI, JSQuAD, MARC-ja); the other four are JAQKET v2 (Wikipedia-grounded quiz question answering), the Japanese subset of MGSM (grade-school maths word problems), the Japanese subset of XL-Sum (news summarisation), and the Japanese subset of XWinograd (commonsense pronoun resolution). A model's results under this group name are eight separate numbers on different scales -- exact match, classification accuracy and ROUGE-2 among them -- not one measurement of "Japanese ability."

Task format

A group of eight independently-formatted tasks. JCommonsenseQA, JNLI, MARC-ja and XWinograd are multiple-choice or classification tasks scored by comparing log-likelihoods over labelled options. JAQKET v2 and JSQuAD are extractive question answering scored by exact match against generated text. MGSM is a generated chain-of-reasoning math answer scored by extracting a final number. XL-Sum is free-form generated summarisation scored by ROUGE-2. Few-shot counts also vary by component, from zero (XWinograd) to five (MGSM); no single prompt format or shot count applies across the group.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub