1,024 hard tasks mined from over a million real chatbot conversations, scored automatically by an LLM judge against a task-specific checklist.
unassessed
| Category | human-preference |
|---|---|
| Subcategory | real-user task quality |
| Page status | active |
| Metric | WB-Score and WB-Reward |
| Direction | higher_is_better |
| Unit | points |
| Dataset size | 1024 |
| Dataset licence | CC BY 4.0 |
| Publisher | Allen Institute for AI (AI2) |
WildBench draws its questions from real conversations logged by AI2's WildChat project rather than writing them by hand, then keeps only the harder, more distinguishing ones. Tasks span writing assistance, coding, math, data analysis, role play and planning, and over a fifth of the conversations run to three or more turns, so the benchmark exercises both single-turn quality and multi-turn coherence. The goal is to approximate what a broad population of real users actually asks chat models to do, rather than a curated academic question set.
Open-ended chat completion (single- or multi-turn) over a real user task, graded after the fact by an LLM judge using a per-task checklist rather than a fixed answer key.
| Model | Provider | Score | Card as of |
|---|---|---|---|
| Claude Opus 4 | Anthropic | 82.5 | 2026-04 |
| Claude Opus 4.6 | Anthropic | 82.5 | 2026-04 |
| GPT-4 | OpenAI | 80.5 | 2026-04 |
| GPT-4.1 | OpenAI | 80.5 | 2026-04 |
| GPT-4.1 mini | OpenAI | 80.5 | 2026-04 |
| GPT-4.1 nano | OpenAI | 80.5 | 2026-04 |
| Gemini 2.5 Pro | Google DeepMind | 79.8 | 2026-04 |
| Gemini 2.5 Pro Preview 05-06 | Google DeepMind | 79.8 | 2026-04 |
| Gemini 2.5 Pro Preview 06-05 | Google DeepMind | 79.8 | 2026-04 |
| Gemini 2.5 Pro Preview TTS | Google DeepMind | 79.8 | 2026-04 |
| Claude Sonnet 4 | Anthropic | 78.2 | 2026-04 |
| Claude Sonnet 4.5 | Anthropic | 78.2 | 2026-04 |
| Claude Sonnet 4.5 (latest) | Anthropic | 78.2 | 2026-04 |
| GPT-4o | OpenAI | 75.8 | 2026-04 |
| GPT-4o (2024-05-13) | OpenAI | 75.8 | 2026-04 |
| GPT-4o (2024-08-06) | OpenAI | 75.8 | 2026-04 |
| GPT-4o (2024-11-20) | OpenAI | 75.8 | 2026-04 |
| GPT-4o mini | OpenAI | 75.8 | 2026-04 |
| Qwen 3 235B Instruct | Cerebras | 73.8 | 2026-04 |
| Qwen3 235B-A22B | Alibaba / Qwen Team | 73.8 | 2026-04 |
| DeepSeek R1 | DeepSeek | 72.5 | 2026-04 |
| DeepSeek R1 0528 | DeepSeek | 72.5 | 2026-04 |
| DeepSeek R1 0528 NVFP4 v2 | NVIDIA | 72.5 | 2026-04 |
| DeepSeek Reasoner | DeepSeek | 72.5 | 2026-04 |
| DeepSeek Chat | DeepSeek | 70.2 | 2026-04 |
| DeepSeek V2 | DeepSeek | 70.2 | 2026-04 |
| DeepSeek V2 Lite | DeepSeek | 70.2 | 2026-04 |
| DeepSeek V2 Lite Chat | DeepSeek | 70.2 | 2026-04 |
| DeepSeek V3 | DeepSeek | 70.2 | 2026-04 |
| DeepSeek V3 0324 | DeepSeek | 70.2 | 2026-04 |
| DeepSeek V3.1 | DeepSeek | 70.2 | 2026-04 |
| DeepSeek V3.2 | DeepSeek | 70.2 | 2026-04 |
| DeepSeek V3.2 Exp | DeepSeek | 70.2 | 2026-04 |
| Llama 3.3 70B Instruct NVFP4 | NVIDIA | 68.2 | 2026-04 |
| Llama-3.3-70B-Instruct | Meta | 68.2 | 2026-04 |
| Mistral Large (latest) | Mistral AI | 65.8 | 2026-04 |
| Mistral Large 2.1 | Mistral AI | 65.8 | 2026-04 |
| Mistral Large 3 | Mistral AI | 65.8 | 2026-04 |
| Qwen3 32B | Alibaba / Qwen Team | 65.5 | 2026-04 |
| Qwen3 32B AWQ | Alibaba / Qwen Team | 65.5 | 2026-04 |
| Qwen3 32B NVFP4 | NVIDIA | 65.5 | 2026-04 |
| Command R+ | Cohere | 60.5 | 2026-04 |
| phi 4 | Microsoft | 58.2 | 2026-04 |
| Phi 4 mini instruct | Microsoft | 58.2 | 2026-04 |
| Phi 4 multimodal instruct | Microsoft | 58.2 | 2026-04 |