LiveCodeBench scores code generation, self-repair, test-output prediction and code execution on dated competitive-programming problems, filterable by a model's training cutoff.
unassessed
| Category | coding |
|---|---|
| Subcategory | competitive programming, contamination-resistant via dated problems |
| Page status | active |
| Metric | Pass@1 (and Pass@5 for code generation) |
| Direction | higher_is_better |
| Unit | % |
| Dataset licence | CC BY 4.0 |
| Publisher | UC Berkeley, MIT and Cornell University |
LiveCodeBench evaluates a model on competitive-programming problems collected continuously from LeetCode, AtCoder and Codeforces, going beyond plain code generation to also test self-repair (fixing a wrong solution given feedback), test output prediction (predicting what a given piece of code outputs) and code execution (simulating running code by hand). Because every problem is tagged with its original release date, a reader can score a model only on problems released after that model's training cutoff, directly addressing whether high scores reflect memorised solutions rather than genuine problem-solving.
For code generation, the model is given a natural-language problem statement (as posed on the source contest site) and must produce a working solution, evaluated against the contest's own or reconstructed test cases. The other three scenarios reuse the same problem pool but change what the model is asked to produce: a corrected solution given a failing one and error feedback (self-repair), the printed output of a given program on given input (test output prediction), or the result of executing a given snippet by reasoning about it directly (code execution).
| Model | Provider | Score | Card as of |
|---|---|---|---|
| Gemini 3 Pro Preview | Google DeepMind | 91.7 | 2026-04 |
| GPT-5.2 Pro | OpenAI | 88.9 | 2026-04 |
| GPT-5.1 | OpenAI | 86.8 | 2026-04 |
| GPT-5.1 Chat | OpenAI | 86.8 | 2026-04 |
| o4-mini | OpenAI | 85.9 | 2026-04 |
| o4-mini-deep-research | OpenAI | 85.9 | 2026-04 |
| GPT-5.1 Codex | OpenAI | 84.9 | 2026-04 |
| GPT-5.1 Codex Max | OpenAI | 84.9 | 2026-04 |
| GPT-5 | OpenAI | 84.6 | 2026-04 |
| GPT-5 Chat (latest) | OpenAI | 84.6 | 2026-04 |
| GPT-5 Pro | OpenAI | 84.6 | 2026-04 |
| GPT-5.2 | OpenAI | 84.6 | 2026-04 |
| GPT-5.2 Chat | OpenAI | 84.6 | 2026-04 |
| GPT-5.2 Codex | OpenAI | 84.6 | 2026-04 |
| GPT-5.3 Chat (latest) | OpenAI | 84.6 | 2026-04 |
| GPT-5.3 Codex | OpenAI | 84.6 | 2026-04 |
| GPT-5.3 Codex Spark | OpenAI | 84.6 | 2026-04 |
| GPT-5.4 | OpenAI | 84.6 | 2026-04 |
| GPT-5.4 mini | OpenAI | 84.6 | 2026-04 |
| GPT-5.4 nano | OpenAI | 84.6 | 2026-04 |
| GPT-5.4 Pro | OpenAI | 84.6 | 2026-04 |
| GPT-5-Codex | OpenAI | 84.0 | 2026-04 |
| GPT-5 Mini | OpenAI | 83.8 | 2026-04 |
| GPT-5.1 Codex mini | OpenAI | 83.6 | 2026-04 |
| Grok 4 Fast | xAI | 83.2 | 2026-04 |
| Grok 4 Fast (Non-Reasoning) | xAI | 83.2 | 2026-04 |
| MiniMax-M2 | MiniMax | 82.6 | 2026-04 |
| MiniMax-M2.5 | MiniMax | 82.6 | 2026-04 |
| MiniMax-M2.7 | MiniMax | 82.6 | 2026-04 |
| Grok 4 | xAI | 81.9 | 2026-04 |
| Grok 4.1 Fast | xAI | 81.9 | 2026-04 |
| Grok 4.1 Fast (Non-Reasoning) | xAI | 81.9 | 2026-04 |
| Grok 4.20 (Non-Reasoning) | xAI | 81.9 | 2026-04 |
| Grok 4.20 (Reasoning) | xAI | 81.9 | 2026-04 |
| Grok 4.20 Multi-Agent | xAI | 81.9 | 2026-04 |
| MiniMax-M2.1 | MiniMax | 81.0 | 2026-04 |
| Gemini 3 Flash Preview | Google DeepMind | 79.7 | 2026-04 |
| Grok 3 | xAI | 79.4 | 2026-04 |
| Grok 3 Fast | xAI | 79.4 | 2026-04 |
| Grok 3 Fast Latest | xAI | 79.4 | 2026-04 |
| Grok 3 Latest | xAI | 79.4 | 2026-04 |
| DeepSeek R1 0528 | DeepSeek | 77.0 | 2026-04 |
| DeepSeek R1 0528 NVFP4 v2 | NVIDIA | 77.0 | 2026-04 |
| Claude Opus 4.5 | Anthropic | 73.8 | 2026-04 |
| Claude Opus 4.5 (latest) | Anthropic | 73.8 | 2026-04 |
| o3-mini | OpenAI | 71.7 | 2026-04 |
| Grok 3 Mini | xAI | 69.6 | 2026-04 |
| Grok 3 Mini Fast | xAI | 69.6 | 2026-04 |
| Grok 3 Mini Fast Latest | xAI | 69.6 | 2026-04 |
| Grok 3 Mini Latest | xAI | 69.6 | 2026-04 |
| o1 | OpenAI | 67.9 | 2026-04 |
| o1-preview | OpenAI | 67.9 | 2026-04 |
| Claude Sonnet 4 (latest) | Anthropic | 65.5 | 2026-04 |
| Claude Opus 4.1 | Anthropic | 65.4 | 2026-04 |
| Claude Opus 4.1 (latest) | Anthropic | 65.4 | 2026-04 |
| Qwen3 30B A3B Instruct 2507 | Alibaba / Qwen Team | 62.6 | 2026-04 |
| Qwen3 30B A3B NVFP4 | NVIDIA | 62.6 | 2026-04 |
| Qwen3 30B-A3B | Alibaba / Qwen Team | 62.6 | 2026-04 |
| Claude Opus 4 | Anthropic | 62.4 | 2026-04 |
| Claude Opus 4.6 | Anthropic | 62.4 | 2026-04 |
| Qwen3 235B-A22B | Alibaba / Qwen Team | 62.2 | 2026-04 |
| DeepSeek Chat | DeepSeek | 59.3 | 2026-04 |
| DeepSeek V3.2 | DeepSeek | 59.3 | 2026-04 |
| Claude Sonnet 4 | Anthropic | 59.0 | 2026-04 |
| Claude Sonnet 4.5 | Anthropic | 59.0 | 2026-04 |
| Claude Sonnet 4.5 (latest) | Anthropic | 59.0 | 2026-04 |
| Qwen3-Coder 480B-A35B Instruct | Alibaba / Qwen Team | 58.5 | 2026-04 |
| o3 | OpenAI | 58.3 | 2026-04 |
| o3-deep-research | OpenAI | 58.3 | 2026-04 |
| Gemini 2.5 Pro | Google DeepMind | 57.8 | 2026-04 |
| Gemini 2.5 Pro Preview 05-06 | Google DeepMind | 57.8 | 2026-04 |
| Gemini 2.5 Pro Preview 06-05 | Google DeepMind | 57.8 | 2026-04 |
| Gemini 2.5 Pro Preview TTS | Google DeepMind | 57.8 | 2026-04 |
| DeepSeek V3.1 | DeepSeek | 57.7 | 2026-04 |
| DeepSeek R1 Distill Llama 70B | DeepSeek | 57.5 | 2026-04 |
| DeepSeek R1 Distill Qwen 32B | DeepSeek | 57.2 | 2026-04 |
| Qwen2.5 72B Instruct | Alibaba / Qwen Team | 55.5 | 2026-04 |
| DeepSeek V3.2 Exp | DeepSeek | 55.4 | 2026-04 |
| Qwen3 32B | Alibaba / Qwen Team | 54.6 | 2026-04 |
| Qwen3 32B AWQ | Alibaba / Qwen Team | 54.6 | 2026-04 |
| Qwen3 32B NVFP4 | NVIDIA | 54.6 | 2026-04 |
| Claude Opus 4 (latest) | Anthropic | 54.2 | 2026-04 |
| DeepSeek R1 Distill Qwen 14B | DeepSeek | 53.1 | 2026-04 |
| DeepSeek R1 | DeepSeek | 52.1 | 2026-04 |
| DeepSeek Reasoner | DeepSeek | 52.1 | 2026-04 |
| Magistral Small | Mistral AI | 51.3 | 2026-04 |
| Magistral Small 2506 | Mistral AI | 51.3 | 2026-04 |
| Claude Haiku 4.5 | Anthropic | 51.1 | 2026-04 |
| Claude Haiku 4.5 (latest) | Anthropic | 51.1 | 2026-04 |
| Magistral Medium (latest) | Mistral AI | 50.3 | 2026-04 |
| Gemini 2.5 Flash | Google DeepMind | 49.5 | 2026-04 |
| Gemini 2.5 Flash Image | Google DeepMind | 49.5 | 2026-04 |
| Gemini 2.5 Flash Image (Preview) | Google DeepMind | 49.5 | 2026-04 |
| Gemini 2.5 Flash Preview 04-17 | Google DeepMind | 49.5 | 2026-04 |
| Gemini 2.5 Flash Preview 05-20 | Google DeepMind | 49.5 | 2026-04 |
| Gemini 2.5 Flash Preview 09-25 | Google DeepMind | 49.5 | 2026-04 |
| Gemini 2.5 Flash Preview TTS | Google DeepMind | 49.5 | 2026-04 |
| DeepSeek V3 0324 | DeepSeek | 49.2 | 2026-04 |
| GPT-4 | OpenAI | 48.3 | 2026-04 |
| GPT-4.1 | OpenAI | 48.3 | 2026-04 |
| GPT-4.1 mini | OpenAI | 48.3 | 2026-04 |
| GPT-4.1 nano | OpenAI | 48.3 | 2026-04 |
| Claude Sonnet 3.7 | Anthropic | 47.3 | 2026-04 |
| GPT-5 Nano | OpenAI | 47.0 | 2026-04 |
| DeepSeek V3 | DeepSeek | 40.5 | 2026-04 |
| Llama 4 Maverick 17B 128E Instruct | Meta | 39.7 | 2026-04 |
| Llama-4-Maverick-17B-128E-Instruct-FP8 | Meta | 39.7 | 2026-04 |
| Claude Sonnet 3.5 | Anthropic | 38.1 | 2026-04 |
| Claude Sonnet 3.5 v2 | Anthropic | 38.1 | 2026-04 |
| Gemini 2.0 Flash | Google DeepMind | 35.1 | 2026-04 |
| Gemini 2.0 Flash Lite | Google DeepMind | 35.1 | 2026-04 |
| Mistral Large (latest) | Mistral AI | 34.4 | 2026-04 |
| Mistral Large 2.1 | Mistral AI | 34.4 | 2026-04 |
| Mistral Large 3 | Mistral AI | 34.4 | 2026-04 |
| Gemini 2.5 Flash Lite | Google DeepMind | 33.7 | 2026-04 |
| Gemini 2.5 Flash Lite Preview 06-17 | Google DeepMind | 33.7 | 2026-04 |
| Gemini 2.5 Flash Lite Preview 09-25 | Google DeepMind | 33.7 | 2026-04 |
| GPT-4o | OpenAI | 31.7 | 2026-04 |
| GPT-4o (2024-05-13) | OpenAI | 31.7 | 2026-04 |
| GPT-4o (2024-08-06) | OpenAI | 31.7 | 2026-04 |
| GPT-4o (2024-11-20) | OpenAI | 31.7 | 2026-04 |
| Codestral (latest) | Mistral AI | 31.4 | 2026-04 |
| Llama 3.1 405B | Meta | 30.5 | 2026-04 |
| Llama 3.1 405B FP8 | Meta | 30.5 | 2026-04 |
| Llama 3.1 405B Instruct | Meta | 30.5 | 2026-04 |
| Llama 3.1 405B Instruct FP8 | Meta | 30.5 | 2026-04 |
| Llama 4 Scout 17B 16E | Meta | 29.9 | 2026-04 |
| Llama 4 Scout 17B 16E Instruct | Meta | 29.9 | 2026-04 |
| Llama-4-Scout-17B-16E-Instruct-FP8 | Meta | 29.9 | 2026-04 |
| Gemma 3 27B | Google DeepMind | 29.7 | 2026-04 |
| Qwen2.5 Coder 32B Instruct | Alibaba / Qwen Team | 29.5 | 2026-04 |
| Qwen2.5 Coder 32B Instruct AWQ | Alibaba / Qwen Team | 29.5 | 2026-04 |
| GPT-4 Turbo | OpenAI | 29.1 | 2026-04 |
| Claude Haiku 3.5 | Anthropic | 28.8 | 2026-04 |
| Claude Haiku 3.5 (latest) | Anthropic | 28.8 | 2026-04 |
| Llama 3.3 70B Instruct NVFP4 | NVIDIA | 28.8 | 2026-04 |
| Llama-3.3-70B-Instruct | Meta | 28.8 | 2026-04 |
| Gemma 3 12B | Google DeepMind | 24.6 | 2026-04 |
| GPT-4o mini | OpenAI | 23.4 | 2026-04 |
| Llama 3.1 70B | Meta | 23.2 | 2026-04 |
| Llama 3.1 70B Instruct | Meta | 23.2 | 2026-04 |
| phi 4 | Microsoft | 23.1 | 2026-04 |
| Phi 4 mini instruct | Microsoft | 23.1 | 2026-04 |
| Gemma 3 4B | Google DeepMind | 12.6 | 2026-04 |