P-MMEval

A Qwen/Tongyi parallel multilingual suite that extends eight existing tasks across the same ten languages so cross-lingual gaps are not confounded with different item sets.

Also known as: PMMEval, P MMEval

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorycomposite
Subcategoryparallel multilingual suite: translation, NLI, commonsense, code, math, logic, knowledge, instruction following
Page statusactive
Metrictask-dependent: BLEU (FLORES), pass@1 (HumanEval-XL), accuracy (XNLI, MHellaSwag, MGSM, MLogiQA, MMMLU, MIFEval)
Directionhigher_is_better
Dataset licenceApache-2.0
PublisherTongyi Lab, Alibaba Group

What it measures

P-MMEval (Parallel Multilingual Multitask Evaluation) scores the same underlying items in up to ten languages: English, Chinese, Arabic, Spanish, French, Japanese, Korean, Portuguese, Thai and Vietnamese. The paper body says those ten span seven language families; the Hub card says eight. Where a source benchmark lacked a language, the authors add expert-reviewed translations so coverage is consistent and samples are parallel. The eight source tasks mix generation (FLORES-200 English-to-X, HumanEval-XL) with understanding and specialised skills (XNLI, MHellaSwag, MGSM, MLogiQA, MMMLU, MIFEval).

Task format

Varies by task: free-text translation (FLORES), code generation (HumanEval-XL), numeric or short-answer math (MGSM), four-option multiple choice (MLogiQA, MMMLU, MHellaSwag), three-way NLI (XNLI), and instruction-following checks (MIFEval). OpenCompass runs each language as its own generate-until task under the pmmeval_gen umbrella.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub