An Arena leaderboard built only from votes on prompts an automatic classifier scored as complex and demanding across several hardness criteria.
unassessed
| Category | human-preference |
|---|---|
| Subcategory | pairwise human preference, algorithmically hard prompts |
| Page status | active |
| Metric | Elo |
| Direction | higher_is_better |
| Publisher | LMSYS (Large Model Systems Organization), UC Berkeley Sky Computing Lab (founding org); operates today as Arena (formerly LMArena) |
arena_elo_hard_prompts restricts the same anonymous pairwise-vote pool behind the Arena text leaderboard to prompts an automatic classifier scored as demanding. LMSYS defined seven hardness criteria — specificity, domain knowledge, complexity, problem-solving, creativity, technical accuracy and real-world application — and used Llama-3-70B-Instruct to label whether each of over a million Arena prompts met each criterion. Prompts meeting six or more of the seven (about 20% of the labelled pool) form the Hard Prompts category, published as separate English and Overall (multilingual) leaderboards. A de-duplication step also down-samples very common, low-signal prompts (chiefly greetings) before the leaderboard is built.
Anonymous, randomized side-by-side text chat restricted to prompts scoring 6 or more of 7 hardness criteria; a user votes for the preferred response.
| Model | Provider | Score | Card as of |
|---|---|---|---|
| Claude Opus 4.6 | Anthropic | 1534.7 | 2026-04 |
| GPT-5.4 | OpenAI | 1507.2 | 2026-04 |
| Grok 4.20 (Reasoning) | xAI | 1505.0 | 2026-04 |
| Claude Opus 4.5 | Anthropic | 1500.0 | 2026-04 |
| Claude Sonnet 4.6 | Anthropic | 1498.7 | 2026-04 |
| GPT-5.2 | OpenAI | 1495.4 | 2026-04 |
| GLM 5.1 | Z.ai (Zhipu AI) | 1494.3 | 2026-04 |
| Grok 4.20 Multi-Agent | xAI | 1491.6 | 2026-04 |
| Claude Sonnet 4.5 | Anthropic | 1485.0 | 2026-04 |
| GPT-5.4 mini | OpenAI | 1483.9 | 2026-04 |
| Grok 4 | xAI | 1482.8 | 2026-04 |
| GPT-5.3 Codex Spark | OpenAI | 1482.5 | 2026-04 |
| Claude Opus 4.1 | Anthropic | 1479.9 | 2026-04 |
| GLM 5 | Zhipu AI | 1475.9 | undated |
| GPT-5.1 | OpenAI | 1474.6 | 2026-04 |
| Gemma 4 31B | Google DeepMind | 1474.2 | 2026-04 |
| GLM 4.7 | Zhipu AI | 1463.5 | undated |
| gemma 4 26B A4B it | Google DeepMind | 1461.1 | 2026-04 |
| Gemini 2.5 Pro | Google DeepMind | 1460.5 | 2026-04 |
| GPT-4o | OpenAI | 1456.0 | 2026-04 |
| Claude Opus 4 | Anthropic | 1455.5 | 2026-04 |
| GPT-5 | OpenAI | 1447.9 | 2026-04 |
| DeepSeek V3.2 Exp | DeepSeek | 1447.7 | 2026-04 |
| DeepSeek V3.2 | DeepSeek | 1446.6 | 2026-04 |
| GLM 4.6 | Zhipu AI | 1442.6 | undated |
| Grok 4.1 Fast | xAI | 1441.8 | 2026-04 |
| Qwen3 235B-A22B | Alibaba / Qwen Team | 1440.4 | 2026-04 |
| o3 | OpenAI | 1439.7 | 2026-04 |
| Claude Haiku 4.5 | Anthropic | 1436.6 | 2026-04 |
| DeepSeek R1 0528 | DeepSeek | 1433.6 | 2026-04 |
| DeepSeek V3.1 | DeepSeek | 1433.3 | 2026-04 |
| GLM 4.5 | Zhipu AI | 1432.7 | undated |
| Grok 4 Fast (Non-Reasoning) | xAI | 1432.7 | 2026-04 |
| Mistral Large 3 | Mistral AI | 1431.1 | 2026-04 |
| GPT-4.1 | OpenAI | 1430.8 | 2026-04 |
| Claude Sonnet 4 | Anthropic | 1430.5 | 2026-04 |
| Mistral Medium 3.1 | Mistral AI | 1429.3 | 2026-04 |
| Grok 3 | xAI | 1426.2 | 2026-04 |
| Gemini 2.5 Flash Preview 09-25 | Google DeepMind | 1420.7 | 2026-04 |
| Gemini 2.5 Flash | Google DeepMind | 1420.0 | 2026-04 |
| DeepSeek R1 | DeepSeek | 1418.1 | 2026-04 |
| o1 | OpenAI | 1417.4 | 2026-04 |
| GPT-5.4 nano | OpenAI | 1417.2 | 2026-04 |
| Claude Sonnet 3.7 | Anthropic | 1415.7 | 2026-04 |
| Qwen3-Coder 480B-A35B Instruct | Alibaba / Qwen Team | 1413.6 | 2026-04 |
| DeepSeek V3 0324 | DeepSeek | 1408.3 | 2026-04 |
| o4-mini | OpenAI | 1405.2 | 2026-04 |
| GPT-4.1 mini | OpenAI | 1402.3 | 2026-04 |
| GPT-5 Mini | OpenAI | 1401.8 | 2026-04 |
| Mistral Medium 3 | Mistral AI | 1401.1 | 2026-04 |
| o3-mini | OpenAI | 1401.1 | 2026-04 |
| Claude Sonnet 3.5 v2 | Anthropic | 1396.7 | 2026-04 |
| o1-preview | OpenAI | 1396.1 | 2026-04 |
| GLM 4.5 Air | Zhipu AI | 1391.2 | undated |
| Gemini 2.5 Flash Lite Preview 09-25 | Google DeepMind | 1390.7 | 2026-04 |
| GLM 4.7 Flash | Zhipu AI | 1387.9 | undated |
| GLM 4.6V | Zhipu AI | 1384.1 | undated |
| Gemini 2.5 Flash Lite | Google DeepMind | 1380.9 | 2026-04 |
| Grok 3 Mini | xAI | 1377.8 | 2026-04 |
| GLM 4.5V | Zhipu AI | 1375.5 | undated |
| Mistral Small 3.2 | Mistral AI | 1374.3 | 2026-04 |
| Command A | Cohere | 1367.9 | 2026-04 |
| Qwen3 32B | Alibaba / Qwen Team | 1367.6 | 2026-04 |
| Gemma 3 27B | Google DeepMind | 1364.9 | 2026-04 |
| o1-mini | OpenAI | 1360.5 | 2026-04 |
| Gemini 2.0 Flash | Google DeepMind | 1360.4 | 2026-04 |
| Claude Sonnet 3.5 | Anthropic | 1358.7 | 2026-04 |
| GPT-5 Nano | OpenAI | 1353.3 | 2026-04 |
| DeepSeek V3 | DeepSeek | 1350.3 | 2026-04 |
| Gemini 1.5 Pro | Google DeepMind | 1350.2 | 2026-04 |
| Gemini 2.0 Flash Lite | Google DeepMind | 1347.9 | 2026-04 |
| Qwen3 30B-A3B | Alibaba / Qwen Team | 1345.6 | 2026-04 |
| Claude Haiku 3.5 | Anthropic | 1343.3 | 2026-04 |
| Llama 3.1 405B Instruct | Meta | 1340.1 | 2026-04 |
| Llama 4 Maverick 17B 128E Instruct | Meta | 1338.0 | 2026-04 |
| GPT-4o (2024-05-13) | OpenAI | 1336.8 | 2026-04 |
| Llama 3.1 405B Instruct FP8 | Meta | 1334.7 | 2026-04 |
| GPT-4.1 nano | OpenAI | 1332.8 | 2026-04 |
| Gemma 3 12B | Google DeepMind | 1331.5 | 2026-04 |
| Magistral Medium (latest) | Mistral AI | 1331.2 | 2026-04 |
| Llama 4 Scout 17B 16E Instruct | Meta | 1329.0 | 2026-04 |
| Claude Opus 3 | Anthropic | 1327.0 | 2026-04 |
| GPT-4o (2024-08-06) | OpenAI | 1326.3 | 2026-04 |
| Grok 2 | xAI | 1325.7 | 2026-04 |
| Llama-3.3-70B-Instruct | Meta | 1320.1 | 2026-04 |
| Mistral Small 3.1 24B Instruct 2503 | Mistral AI | 1318.7 | 2026-04 |
| Qwen2.5 72B Instruct | Alibaba / Qwen Team | 1317.3 | 2026-04 |
| GPT-4 Turbo | OpenAI | 1315.4 | 2026-04 |
| Mistral Large 2.1 | Mistral AI | 1312.6 | 2026-04 |
| GPT-4o mini | OpenAI | 1311.1 | 2026-04 |
| Gemini 1.5 Flash | Google DeepMind | 1302.4 | 2026-04 |
| Llama 3.1 70B Instruct | Meta | 1297.6 | 2026-04 |
| Mistral Small 24B Instruct 2501 | Mistral AI | 1284.3 | 2026-04 |
| Gemma 3 4B | Google DeepMind | 1283.4 | 2026-04 |
| gemma 2 27B it | Google DeepMind | 1280.9 | 2026-04 |
| Claude Sonnet 3 | Anthropic | 1280.5 | 2026-04 |
| phi 4 | Microsoft | 1277.2 | 2026-04 |
| Claude Haiku 3 | Anthropic | 1263.2 | 2026-04 |
| Command R+ | Cohere | 1259.3 | 2026-04 |
| Gemini 1.5 Flash-8B | Google DeepMind | 1258.2 | 2026-04 |
| gemma 2 9B it | Google DeepMind | 1256.0 | 2026-04 |
| Command R | Cohere | 1254.6 | 2026-04 |
| Llama 3.1 8B Instruct | Meta | 1221.8 | 2026-04 |