Aider Polyglot Benchmark

225 hard Exercism exercises across six languages, scoring whether a model can write and then fix its own code from failing test output.

Also known as: Aider's polyglot benchmark, polyglot-benchmark

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorycoding
Subcategorymulti-language code editing
Page statusactive
Metricpass rate (2 attempts)
Directionhigher_is_better
Unit%
Dataset size225
Dataset licenceExercise content is Exercism's, used under Exercism's per-track open-source licences
PublisherAider (Paul Gauthier)

What it measures

The model is given an Exercism programming exercise description and must produce a working solution, expressed as an edit to a starting file, in one of six languages: C++, Go, Java, JavaScript, Python or Rust. Unlike a single-shot coding benchmark, it also measures whether the model can act on its own failing test output to fix a first attempt.

Task format

Agentic code editing inside the aider CLI tool: the model receives the exercise prompt and must emit an edit in one of aider's supported edit formats, which aider applies to a real file and tests with the language's unit test suite.

Models reporting this benchmark

These figures come from the model cards, which carry one collection date per card and no per-score attribution. They are shown as reported, not as verified evidence.
ModelProviderScoreCard as of
Gemini 2.5 Pro Preview 06-05Google DeepMind83.12026-04
Claude Opus 4Anthropic82.12026-04
Claude Opus 4.6Anthropic82.12026-04
o3OpenAI80.52026-04
o3-deep-researchOpenAI80.52026-04
o3-miniOpenAI80.52026-04
o3-proOpenAI80.52026-04
Claude Opus 4.1Anthropic78.82026-04
Claude Opus 4.1 (latest)Anthropic78.82026-04
Gemini 2.5 Pro Preview 05-06Google DeepMind76.92026-04
Claude Sonnet 4Anthropic75.52026-04
Claude Sonnet 4.5Anthropic75.52026-04
Claude Sonnet 4.5 (latest)Anthropic75.52026-04
Gemini 2.5 Pro Preview TTSGoogle DeepMind72.92026-04
Qwen3-Coder 480B-A35B InstructAlibaba / Qwen Team72.52026-04
GPT-5OpenAI72.12026-04
GPT-5 Chat (latest)OpenAI72.12026-04
GPT-5 MiniOpenAI72.12026-04
GPT-5 NanoOpenAI72.12026-04
GPT-5 ProOpenAI72.12026-04
GPT-5-CodexOpenAI72.12026-04
GPT-5.1OpenAI72.12026-04
GPT-5.1 ChatOpenAI72.12026-04
GPT-5.1 CodexOpenAI72.12026-04
GPT-5.1 Codex MaxOpenAI72.12026-04
GPT-5.1 Codex miniOpenAI72.12026-04
GPT-5.2OpenAI72.12026-04
GPT-5.2 ChatOpenAI72.12026-04
GPT-5.2 CodexOpenAI72.12026-04
GPT-5.2 ProOpenAI72.12026-04
GPT-5.3 Chat (latest)OpenAI72.12026-04
GPT-5.3 CodexOpenAI72.12026-04
GPT-5.3 Codex SparkOpenAI72.12026-04
GPT-5.4OpenAI72.12026-04
GPT-5.4 miniOpenAI72.12026-04
GPT-5.4 nanoOpenAI72.12026-04
GPT-5.4 ProOpenAI72.12026-04
Claude Opus 4 (latest)Anthropic72.02026-04
o4-miniOpenAI72.02026-04
o4-mini-deep-researchOpenAI72.02026-04
Gemini 2.5 ProGoogle DeepMind71.22026-04
Grok 4xAI70.82026-04
Grok 4 FastxAI70.82026-04
Grok 4 Fast (Non-Reasoning)xAI70.82026-04
Grok 4.1 FastxAI70.82026-04
Grok 4.1 Fast (Non-Reasoning)xAI70.82026-04
Grok 4.20 (Non-Reasoning)xAI70.82026-04
Grok 4.20 (Reasoning)xAI70.82026-04
Grok 4.20 Multi-AgentxAI70.82026-04
GPT-4.1OpenAI68.52026-04
Qwen 3 235B InstructCerebras68.22026-04
Qwen3 235B-A22BAlibaba / Qwen Team68.22026-04
DeepSeek R1DeepSeek65.82026-04
DeepSeek R1 0528DeepSeek65.82026-04
DeepSeek R1 0528 NVFP4 v2NVIDIA65.82026-04
DeepSeek ReasonerDeepSeek65.82026-04
Claude Sonnet 3.7Anthropic64.92026-04
o1OpenAI61.72026-04
o1-previewOpenAI61.72026-04
Claude Sonnet 4 (latest)Anthropic61.32026-04
Codestral (latest)Mistral AI60.52026-04
DeepSeek ChatDeepSeek58.52026-04
DeepSeek V3DeepSeek58.52026-04
DeepSeek V3 0324DeepSeek58.52026-04
DeepSeek V3.1DeepSeek58.52026-04
DeepSeek V3.2DeepSeek58.52026-04
DeepSeek V3.2 ExpDeepSeek58.52026-04
GPT-4oOpenAI58.22026-04
GPT-4o (2024-05-13)OpenAI58.22026-04
GPT-4o (2024-08-06)OpenAI58.22026-04
GPT-4o (2024-11-20)OpenAI58.22026-04
GPT-4o miniOpenAI58.22026-04
Gemini 2.5 FlashGoogle DeepMind55.12026-04
Gemini 2.5 Flash ImageGoogle DeepMind55.12026-04
Gemini 2.5 Flash Image (Preview)Google DeepMind55.12026-04
Gemini 2.5 Flash LiteGoogle DeepMind55.12026-04
Gemini 2.5 Flash Lite Preview 06-17Google DeepMind55.12026-04
Gemini 2.5 Flash Lite Preview 09-25Google DeepMind55.12026-04
Gemini 2.5 Flash Preview 05-20Google DeepMind55.12026-04
Gemini 2.5 Flash Preview 09-25Google DeepMind55.12026-04
Gemini 2.5 Flash Preview TTSGoogle DeepMind55.12026-04
Gemma 4 31BGoogle DeepMind54.82026-04
gemma 4 31B itGoogle DeepMind54.82026-04
gemma 4 31B it GGUFUnsloth54.82026-04
Gemma 4 31B IT NVFP4NVIDIA54.82026-04
Grok 3xAI53.32026-04
Grok 3 FastxAI53.32026-04
Grok 3 Fast LatestxAI53.32026-04
Grok 3 LatestxAI53.32026-04
Mistral Large (latest)Mistral AI52.82026-04
Mistral Large 2.1Mistral AI52.82026-04
Mistral Large 3Mistral AI52.82026-04
GPT-4OpenAI52.42026-04
Gemma 4 26BGoogle DeepMind52.12026-04
Claude Sonnet 3.5Anthropic51.62026-04
Claude Sonnet 3.5 v2Anthropic51.62026-04
Grok 3 MinixAI49.32026-04
Grok 3 Mini FastxAI49.32026-04
Grok 3 Mini Fast LatestxAI49.32026-04
Grok 3 Mini LatestxAI49.32026-04
Gemini 2.5 Flash Preview 04-17Google DeepMind47.12026-04
Llama 3.3 70B Instruct NVFP4NVIDIA45.22026-04
Llama-3.3-70B-InstructMeta45.22026-04
Qwen3 32BAlibaba / Qwen Team40.02026-04
Qwen3 32B AWQAlibaba / Qwen Team40.02026-04
Qwen3 32B NVFP4NVIDIA40.02026-04
o1-miniOpenAI32.92026-04
GPT-4.1 miniOpenAI32.42026-04
Claude Haiku 3.5Anthropic28.02026-04
Claude Haiku 3.5 (latest)Anthropic28.02026-04
Qwen2.5 Coder 32B InstructAlibaba / Qwen Team16.42026-04
Qwen2.5 Coder 32B Instruct AWQAlibaba / Qwen Team16.42026-04
Llama 4 Maverick 17B 128E InstructMeta15.62026-04
Llama-4-Maverick-17B-128E-Instruct-FP8Meta15.62026-04
GPT-4.1 nanoOpenAI8.92026-04
Gemma 3 27BGoogle DeepMind4.92026-04

Data

This page as JSON · Edit on GitHub