MultiPL-E: Python

The Python subset of MultiPL-E: the original, untranslated HumanEval and MBPP problems, used as the harness's reference language.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorycoding
Subcategorymultilingual code generation
Page statusactive
Metricpass@1
Directionhigher_is_better
Unit%
Dataset licenceMIT
PublisherNortheastern University Programming Research Lab (nuprl)

What it measures

This subset is Python itself: the original HumanEval and MBPP function-completion problems that every other MultiPL-E language is translated from, run through the same execution harness as the translated languages. It exists so a model's in-language Python score can be compared directly against its scores on the translated languages, using one consistent harness and prompt style.

Task format

Function completion in Python: given the original signature, docstring and doctests, the model generates a function body, which is executed and checked against the original unit tests.

Models reporting this benchmark

These figures come from the model cards, which carry one collection date per card and no per-score attribution. They are shown as reported, not as verified evidence.
ModelProviderScoreCard as of
Claude Opus 4Anthropic95.52026-04
Claude Opus 4.6Anthropic95.52026-04
Claude Sonnet 4Anthropic93.22026-04
Claude Sonnet 4.5Anthropic93.22026-04
Claude Sonnet 4.5 (latest)Anthropic93.22026-04
DeepSeek R1DeepSeek92.82026-04
DeepSeek R1 0528DeepSeek92.82026-04
DeepSeek R1 0528 NVFP4 v2NVIDIA92.82026-04
DeepSeek ReasonerDeepSeek92.82026-04
GPT-4.1OpenAI92.52026-04
Gemini 2.5 ProGoogle DeepMind91.82026-04
GPT-4oOpenAI90.22026-04
GPT-4o (2024-05-13)OpenAI90.22026-04
GPT-4o (2024-08-06)OpenAI90.22026-04
GPT-4o (2024-11-20)OpenAI90.22026-04
GPT-4o miniOpenAI90.22026-04
Qwen2.5 Coder 32B InstructAlibaba / Qwen Team88.52026-04
Qwen2.5 Coder 32B Instruct AWQAlibaba / Qwen Team88.52026-04
Gemma 4 31BGoogle DeepMind84.12026-04
gemma 4 31B itGoogle DeepMind84.12026-04
gemma 4 31B it GGUFUnsloth84.12026-04
Gemma 4 31B IT NVFP4NVIDIA84.12026-04
Mistral Large (latest)Mistral AI82.82026-04
Mistral Large 2.1Mistral AI82.82026-04
Mistral Large 3Mistral AI82.82026-04
Gemma 4 26BGoogle DeepMind82.52026-04
Qwen2.5 Coder 14B InstructAlibaba / Qwen Team82.12026-04
Codestral (latest)Mistral AI81.52026-04
Llama 3.3 70B Instruct NVFP4NVIDIA80.22026-04
Llama-3.3-70B-InstructMeta80.22026-04
phi 4Microsoft79.22026-04
Phi 4 mini instructMicrosoft79.22026-04
Llama 3.1 70BMeta78.52026-04
Llama 3.1 70B InstructMeta78.52026-04
Qwen2.5 Coder 7B InstructAlibaba / Qwen Team75.82026-04
Qwen2.5 Coder 7B Instruct GPTQ Int4Alibaba / Qwen Team75.82026-04
CodeLlama 34B Instruct hfMeta68.52026-04

Data

This page as JSON · Edit on GitHub