LiveCodeBench

LiveCodeBench scores code generation, self-repair, test-output prediction and code execution on dated competitive-programming problems, filterable by a model's training cutoff.

Also known as: LCB, livecodebench

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorycoding
Subcategorycompetitive programming, contamination-resistant via dated problems
Page statusactive
MetricPass@1 (and Pass@5 for code generation)
Directionhigher_is_better
Unit%
Dataset licenceCC BY 4.0
PublisherUC Berkeley, MIT and Cornell University

What it measures

LiveCodeBench evaluates a model on competitive-programming problems collected continuously from LeetCode, AtCoder and Codeforces, going beyond plain code generation to also test self-repair (fixing a wrong solution given feedback), test output prediction (predicting what a given piece of code outputs) and code execution (simulating running code by hand). Because every problem is tagged with its original release date, a reader can score a model only on problems released after that model's training cutoff, directly addressing whether high scores reflect memorised solutions rather than genuine problem-solving.

Task format

For code generation, the model is given a natural-language problem statement (as posed on the source contest site) and must produce a working solution, evaluated against the contest's own or reconstructed test cases. The other three scenarios reuse the same problem pool but change what the model is asked to produce: a corrected solution given a failing one and error feedback (self-repair), the printed output of a given program on given input (test output prediction), or the result of executing a given snippet by reasoning about it directly (code execution).

Models reporting this benchmark

These figures come from the model cards, which carry one collection date per card and no per-score attribution. They are shown as reported, not as verified evidence.
ModelProviderScoreCard as of
Gemini 3 Pro PreviewGoogle DeepMind91.72026-04
GPT-5.2 ProOpenAI88.92026-04
GPT-5.1OpenAI86.82026-04
GPT-5.1 ChatOpenAI86.82026-04
o4-miniOpenAI85.92026-04
o4-mini-deep-researchOpenAI85.92026-04
GPT-5.1 CodexOpenAI84.92026-04
GPT-5.1 Codex MaxOpenAI84.92026-04
GPT-5OpenAI84.62026-04
GPT-5 Chat (latest)OpenAI84.62026-04
GPT-5 ProOpenAI84.62026-04
GPT-5.2OpenAI84.62026-04
GPT-5.2 ChatOpenAI84.62026-04
GPT-5.2 CodexOpenAI84.62026-04
GPT-5.3 Chat (latest)OpenAI84.62026-04
GPT-5.3 CodexOpenAI84.62026-04
GPT-5.3 Codex SparkOpenAI84.62026-04
GPT-5.4OpenAI84.62026-04
GPT-5.4 miniOpenAI84.62026-04
GPT-5.4 nanoOpenAI84.62026-04
GPT-5.4 ProOpenAI84.62026-04
GPT-5-CodexOpenAI84.02026-04
GPT-5 MiniOpenAI83.82026-04
GPT-5.1 Codex miniOpenAI83.62026-04
Grok 4 FastxAI83.22026-04
Grok 4 Fast (Non-Reasoning)xAI83.22026-04
MiniMax-M2MiniMax82.62026-04
MiniMax-M2.5MiniMax82.62026-04
MiniMax-M2.7MiniMax82.62026-04
Grok 4xAI81.92026-04
Grok 4.1 FastxAI81.92026-04
Grok 4.1 Fast (Non-Reasoning)xAI81.92026-04
Grok 4.20 (Non-Reasoning)xAI81.92026-04
Grok 4.20 (Reasoning)xAI81.92026-04
Grok 4.20 Multi-AgentxAI81.92026-04
MiniMax-M2.1MiniMax81.02026-04
Gemini 3 Flash PreviewGoogle DeepMind79.72026-04
Grok 3xAI79.42026-04
Grok 3 FastxAI79.42026-04
Grok 3 Fast LatestxAI79.42026-04
Grok 3 LatestxAI79.42026-04
DeepSeek R1 0528DeepSeek77.02026-04
DeepSeek R1 0528 NVFP4 v2NVIDIA77.02026-04
Claude Opus 4.5Anthropic73.82026-04
Claude Opus 4.5 (latest)Anthropic73.82026-04
o3-miniOpenAI71.72026-04
Grok 3 MinixAI69.62026-04
Grok 3 Mini FastxAI69.62026-04
Grok 3 Mini Fast LatestxAI69.62026-04
Grok 3 Mini LatestxAI69.62026-04
o1OpenAI67.92026-04
o1-previewOpenAI67.92026-04
Claude Sonnet 4 (latest)Anthropic65.52026-04
Claude Opus 4.1Anthropic65.42026-04
Claude Opus 4.1 (latest)Anthropic65.42026-04
Qwen3 30B A3B Instruct 2507Alibaba / Qwen Team62.62026-04
Qwen3 30B A3B NVFP4NVIDIA62.62026-04
Qwen3 30B-A3BAlibaba / Qwen Team62.62026-04
Claude Opus 4Anthropic62.42026-04
Claude Opus 4.6Anthropic62.42026-04
Qwen3 235B-A22BAlibaba / Qwen Team62.22026-04
DeepSeek ChatDeepSeek59.32026-04
DeepSeek V3.2DeepSeek59.32026-04
Claude Sonnet 4Anthropic59.02026-04
Claude Sonnet 4.5Anthropic59.02026-04
Claude Sonnet 4.5 (latest)Anthropic59.02026-04
Qwen3-Coder 480B-A35B InstructAlibaba / Qwen Team58.52026-04
o3OpenAI58.32026-04
o3-deep-researchOpenAI58.32026-04
Gemini 2.5 ProGoogle DeepMind57.82026-04
Gemini 2.5 Pro Preview 05-06Google DeepMind57.82026-04
Gemini 2.5 Pro Preview 06-05Google DeepMind57.82026-04
Gemini 2.5 Pro Preview TTSGoogle DeepMind57.82026-04
DeepSeek V3.1DeepSeek57.72026-04
DeepSeek R1 Distill Llama 70BDeepSeek57.52026-04
DeepSeek R1 Distill Qwen 32BDeepSeek57.22026-04
Qwen2.5 72B InstructAlibaba / Qwen Team55.52026-04
DeepSeek V3.2 ExpDeepSeek55.42026-04
Qwen3 32BAlibaba / Qwen Team54.62026-04
Qwen3 32B AWQAlibaba / Qwen Team54.62026-04
Qwen3 32B NVFP4NVIDIA54.62026-04
Claude Opus 4 (latest)Anthropic54.22026-04
DeepSeek R1 Distill Qwen 14BDeepSeek53.12026-04
DeepSeek R1DeepSeek52.12026-04
DeepSeek ReasonerDeepSeek52.12026-04
Magistral SmallMistral AI51.32026-04
Magistral Small 2506Mistral AI51.32026-04
Claude Haiku 4.5Anthropic51.12026-04
Claude Haiku 4.5 (latest)Anthropic51.12026-04
Magistral Medium (latest)Mistral AI50.32026-04
Gemini 2.5 FlashGoogle DeepMind49.52026-04
Gemini 2.5 Flash ImageGoogle DeepMind49.52026-04
Gemini 2.5 Flash Image (Preview)Google DeepMind49.52026-04
Gemini 2.5 Flash Preview 04-17Google DeepMind49.52026-04
Gemini 2.5 Flash Preview 05-20Google DeepMind49.52026-04
Gemini 2.5 Flash Preview 09-25Google DeepMind49.52026-04
Gemini 2.5 Flash Preview TTSGoogle DeepMind49.52026-04
DeepSeek V3 0324DeepSeek49.22026-04
GPT-4OpenAI48.32026-04
GPT-4.1OpenAI48.32026-04
GPT-4.1 miniOpenAI48.32026-04
GPT-4.1 nanoOpenAI48.32026-04
Claude Sonnet 3.7Anthropic47.32026-04
GPT-5 NanoOpenAI47.02026-04
DeepSeek V3DeepSeek40.52026-04
Llama 4 Maverick 17B 128E InstructMeta39.72026-04
Llama-4-Maverick-17B-128E-Instruct-FP8Meta39.72026-04
Claude Sonnet 3.5Anthropic38.12026-04
Claude Sonnet 3.5 v2Anthropic38.12026-04
Gemini 2.0 FlashGoogle DeepMind35.12026-04
Gemini 2.0 Flash LiteGoogle DeepMind35.12026-04
Mistral Large (latest)Mistral AI34.42026-04
Mistral Large 2.1Mistral AI34.42026-04
Mistral Large 3Mistral AI34.42026-04
Gemini 2.5 Flash LiteGoogle DeepMind33.72026-04
Gemini 2.5 Flash Lite Preview 06-17Google DeepMind33.72026-04
Gemini 2.5 Flash Lite Preview 09-25Google DeepMind33.72026-04
GPT-4oOpenAI31.72026-04
GPT-4o (2024-05-13)OpenAI31.72026-04
GPT-4o (2024-08-06)OpenAI31.72026-04
GPT-4o (2024-11-20)OpenAI31.72026-04
Codestral (latest)Mistral AI31.42026-04
Llama 3.1 405BMeta30.52026-04
Llama 3.1 405B FP8Meta30.52026-04
Llama 3.1 405B InstructMeta30.52026-04
Llama 3.1 405B Instruct FP8Meta30.52026-04
Llama 4 Scout 17B 16EMeta29.92026-04
Llama 4 Scout 17B 16E InstructMeta29.92026-04
Llama-4-Scout-17B-16E-Instruct-FP8Meta29.92026-04
Gemma 3 27BGoogle DeepMind29.72026-04
Qwen2.5 Coder 32B InstructAlibaba / Qwen Team29.52026-04
Qwen2.5 Coder 32B Instruct AWQAlibaba / Qwen Team29.52026-04
GPT-4 TurboOpenAI29.12026-04
Claude Haiku 3.5Anthropic28.82026-04
Claude Haiku 3.5 (latest)Anthropic28.82026-04
Llama 3.3 70B Instruct NVFP4NVIDIA28.82026-04
Llama-3.3-70B-InstructMeta28.82026-04
Gemma 3 12BGoogle DeepMind24.62026-04
GPT-4o miniOpenAI23.42026-04
Llama 3.1 70BMeta23.22026-04
Llama 3.1 70B InstructMeta23.22026-04
phi 4Microsoft23.12026-04
Phi 4 mini instructMicrosoft23.12026-04
Gemma 3 4BGoogle DeepMind12.62026-04

Data

This page as JSON · Edit on GitHub