Arena Elo — Hard Prompts

An Arena leaderboard built only from votes on prompts an automatic classifier scored as complex and demanding across several hardness criteria.

Also known as: Chatbot Arena Hard Prompts, LMArena Hard Prompts category

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryhuman-preference
Subcategorypairwise human preference, algorithmically hard prompts
Page statusactive
MetricElo
Directionhigher_is_better
PublisherLMSYS (Large Model Systems Organization), UC Berkeley Sky Computing Lab (founding org); operates today as Arena (formerly LMArena)

What it measures

arena_elo_hard_prompts restricts the same anonymous pairwise-vote pool behind the Arena text leaderboard to prompts an automatic classifier scored as demanding. LMSYS defined seven hardness criteria — specificity, domain knowledge, complexity, problem-solving, creativity, technical accuracy and real-world application — and used Llama-3-70B-Instruct to label whether each of over a million Arena prompts met each criterion. Prompts meeting six or more of the seven (about 20% of the labelled pool) form the Hard Prompts category, published as separate English and Overall (multilingual) leaderboards. A de-duplication step also down-samples very common, low-signal prompts (chiefly greetings) before the leaderboard is built.

Task format

Anonymous, randomized side-by-side text chat restricted to prompts scoring 6 or more of 7 hardness criteria; a user votes for the preferred response.

Models reporting this benchmark

These figures come from the model cards, which carry one collection date per card and no per-score attribution. They are shown as reported, not as verified evidence.
ModelProviderScoreCard as of
Claude Opus 4.6Anthropic1534.72026-04
GPT-5.4OpenAI1507.22026-04
Grok 4.20 (Reasoning)xAI1505.02026-04
Claude Opus 4.5Anthropic1500.02026-04
Claude Sonnet 4.6Anthropic1498.72026-04
GPT-5.2OpenAI1495.42026-04
GLM 5.1Z.ai (Zhipu AI)1494.32026-04
Grok 4.20 Multi-AgentxAI1491.62026-04
Claude Sonnet 4.5Anthropic1485.02026-04
GPT-5.4 miniOpenAI1483.92026-04
Grok 4xAI1482.82026-04
GPT-5.3 Codex SparkOpenAI1482.52026-04
Claude Opus 4.1Anthropic1479.92026-04
GLM 5Zhipu AI1475.9undated
GPT-5.1OpenAI1474.62026-04
Gemma 4 31BGoogle DeepMind1474.22026-04
GLM 4.7Zhipu AI1463.5undated
gemma 4 26B A4B itGoogle DeepMind1461.12026-04
Gemini 2.5 ProGoogle DeepMind1460.52026-04
GPT-4oOpenAI1456.02026-04
Claude Opus 4Anthropic1455.52026-04
GPT-5OpenAI1447.92026-04
DeepSeek V3.2 ExpDeepSeek1447.72026-04
DeepSeek V3.2DeepSeek1446.62026-04
GLM 4.6Zhipu AI1442.6undated
Grok 4.1 FastxAI1441.82026-04
Qwen3 235B-A22BAlibaba / Qwen Team1440.42026-04
o3OpenAI1439.72026-04
Claude Haiku 4.5Anthropic1436.62026-04
DeepSeek R1 0528DeepSeek1433.62026-04
DeepSeek V3.1DeepSeek1433.32026-04
GLM 4.5Zhipu AI1432.7undated
Grok 4 Fast (Non-Reasoning)xAI1432.72026-04
Mistral Large 3Mistral AI1431.12026-04
GPT-4.1OpenAI1430.82026-04
Claude Sonnet 4Anthropic1430.52026-04
Mistral Medium 3.1Mistral AI1429.32026-04
Grok 3xAI1426.22026-04
Gemini 2.5 Flash Preview 09-25Google DeepMind1420.72026-04
Gemini 2.5 FlashGoogle DeepMind1420.02026-04
DeepSeek R1DeepSeek1418.12026-04
o1OpenAI1417.42026-04
GPT-5.4 nanoOpenAI1417.22026-04
Claude Sonnet 3.7Anthropic1415.72026-04
Qwen3-Coder 480B-A35B InstructAlibaba / Qwen Team1413.62026-04
DeepSeek V3 0324DeepSeek1408.32026-04
o4-miniOpenAI1405.22026-04
GPT-4.1 miniOpenAI1402.32026-04
GPT-5 MiniOpenAI1401.82026-04
Mistral Medium 3Mistral AI1401.12026-04
o3-miniOpenAI1401.12026-04
Claude Sonnet 3.5 v2Anthropic1396.72026-04
o1-previewOpenAI1396.12026-04
GLM 4.5 AirZhipu AI1391.2undated
Gemini 2.5 Flash Lite Preview 09-25Google DeepMind1390.72026-04
GLM 4.7 FlashZhipu AI1387.9undated
GLM 4.6VZhipu AI1384.1undated
Gemini 2.5 Flash LiteGoogle DeepMind1380.92026-04
Grok 3 MinixAI1377.82026-04
GLM 4.5VZhipu AI1375.5undated
Mistral Small 3.2Mistral AI1374.32026-04
Command ACohere1367.92026-04
Qwen3 32BAlibaba / Qwen Team1367.62026-04
Gemma 3 27BGoogle DeepMind1364.92026-04
o1-miniOpenAI1360.52026-04
Gemini 2.0 FlashGoogle DeepMind1360.42026-04
Claude Sonnet 3.5Anthropic1358.72026-04
GPT-5 NanoOpenAI1353.32026-04
DeepSeek V3DeepSeek1350.32026-04
Gemini 1.5 ProGoogle DeepMind1350.22026-04
Gemini 2.0 Flash LiteGoogle DeepMind1347.92026-04
Qwen3 30B-A3BAlibaba / Qwen Team1345.62026-04
Claude Haiku 3.5Anthropic1343.32026-04
Llama 3.1 405B InstructMeta1340.12026-04
Llama 4 Maverick 17B 128E InstructMeta1338.02026-04
GPT-4o (2024-05-13)OpenAI1336.82026-04
Llama 3.1 405B Instruct FP8Meta1334.72026-04
GPT-4.1 nanoOpenAI1332.82026-04
Gemma 3 12BGoogle DeepMind1331.52026-04
Magistral Medium (latest)Mistral AI1331.22026-04
Llama 4 Scout 17B 16E InstructMeta1329.02026-04
Claude Opus 3Anthropic1327.02026-04
GPT-4o (2024-08-06)OpenAI1326.32026-04
Grok 2xAI1325.72026-04
Llama-3.3-70B-InstructMeta1320.12026-04
Mistral Small 3.1 24B Instruct 2503Mistral AI1318.72026-04
Qwen2.5 72B InstructAlibaba / Qwen Team1317.32026-04
GPT-4 TurboOpenAI1315.42026-04
Mistral Large 2.1Mistral AI1312.62026-04
GPT-4o miniOpenAI1311.12026-04
Gemini 1.5 FlashGoogle DeepMind1302.42026-04
Llama 3.1 70B InstructMeta1297.62026-04
Mistral Small 24B Instruct 2501Mistral AI1284.32026-04
Gemma 3 4BGoogle DeepMind1283.42026-04
gemma 2 27B itGoogle DeepMind1280.92026-04
Claude Sonnet 3Anthropic1280.52026-04
phi 4Microsoft1277.22026-04
Claude Haiku 3Anthropic1263.22026-04
Command R+Cohere1259.32026-04
Gemini 1.5 Flash-8BGoogle DeepMind1258.22026-04
gemma 2 9B itGoogle DeepMind1256.02026-04
Command RCohere1254.62026-04
Llama 3.1 8B InstructMeta1221.82026-04

Data

This page as JSON · Edit on GitHub