A Qwen/Tongyi parallel multilingual suite that extends eight existing tasks across the same ten languages so cross-lingual gaps are not confounded with different item sets.
unassessed
| Category | composite |
|---|---|
| Subcategory | parallel multilingual suite: translation, NLI, commonsense, code, math, logic, knowledge, instruction following |
| Page status | active |
| Metric | task-dependent: BLEU (FLORES), pass@1 (HumanEval-XL), accuracy (XNLI, MHellaSwag, MGSM, MLogiQA, MMMLU, MIFEval) |
| Direction | higher_is_better |
| Dataset licence | Apache-2.0 |
| Publisher | Tongyi Lab, Alibaba Group |
P-MMEval (Parallel Multilingual Multitask Evaluation) scores the same underlying items in up to ten languages: English, Chinese, Arabic, Spanish, French, Japanese, Korean, Portuguese, Thai and Vietnamese. The paper body says those ten span seven language families; the Hub card says eight. Where a source benchmark lacked a language, the authors add expert-reviewed translations so coverage is consistent and samples are parallel. The eight source tasks mix generation (FLORES-200 English-to-X, HumanEval-XL) with understanding and specialised skills (XNLI, MHellaSwag, MGSM, MLogiQA, MMMLU, MIFEval).
Varies by task: free-text translation (FLORES), code generation (HumanEval-XL), numeric or short-answer math (MGSM), four-option multiple choice (MLogiQA, MMMLU, MHellaSwag), three-way NLI (XNLI), and instruction-following checks (MIFEval). OpenCompass runs each language as its own generate-until task under the pmmeval_gen umbrella.
No model card in ModelSpec reports this benchmark yet.