WildBench

1,024 hard tasks mined from over a million real chatbot conversations, scored automatically by an LLM judge against a task-specific checklist.

Also known as: WildBench v2

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryhuman-preference
Subcategoryreal-user task quality
Page statusactive
MetricWB-Score and WB-Reward
Directionhigher_is_better
Unitpoints
Dataset size1024
Dataset licenceCC BY 4.0
PublisherAllen Institute for AI (AI2)

What it measures

WildBench draws its questions from real conversations logged by AI2's WildChat project rather than writing them by hand, then keeps only the harder, more distinguishing ones. Tasks span writing assistance, coding, math, data analysis, role play and planning, and over a fifth of the conversations run to three or more turns, so the benchmark exercises both single-turn quality and multi-turn coherence. The goal is to approximate what a broad population of real users actually asks chat models to do, rather than a curated academic question set.

Task format

Open-ended chat completion (single- or multi-turn) over a real user task, graded after the fact by an LLM judge using a per-task checklist rather than a fixed answer key.

Models reporting this benchmark

These figures come from the model cards, which carry one collection date per card and no per-score attribution. They are shown as reported, not as verified evidence.
ModelProviderScoreCard as of
Claude Opus 4Anthropic82.52026-04
Claude Opus 4.6Anthropic82.52026-04
GPT-4OpenAI80.52026-04
GPT-4.1OpenAI80.52026-04
GPT-4.1 miniOpenAI80.52026-04
GPT-4.1 nanoOpenAI80.52026-04
Gemini 2.5 ProGoogle DeepMind79.82026-04
Gemini 2.5 Pro Preview 05-06Google DeepMind79.82026-04
Gemini 2.5 Pro Preview 06-05Google DeepMind79.82026-04
Gemini 2.5 Pro Preview TTSGoogle DeepMind79.82026-04
Claude Sonnet 4Anthropic78.22026-04
Claude Sonnet 4.5Anthropic78.22026-04
Claude Sonnet 4.5 (latest)Anthropic78.22026-04
GPT-4oOpenAI75.82026-04
GPT-4o (2024-05-13)OpenAI75.82026-04
GPT-4o (2024-08-06)OpenAI75.82026-04
GPT-4o (2024-11-20)OpenAI75.82026-04
GPT-4o miniOpenAI75.82026-04
Qwen 3 235B InstructCerebras73.82026-04
Qwen3 235B-A22BAlibaba / Qwen Team73.82026-04
DeepSeek R1DeepSeek72.52026-04
DeepSeek R1 0528DeepSeek72.52026-04
DeepSeek R1 0528 NVFP4 v2NVIDIA72.52026-04
DeepSeek ReasonerDeepSeek72.52026-04
DeepSeek ChatDeepSeek70.22026-04
DeepSeek V2DeepSeek70.22026-04
DeepSeek V2 LiteDeepSeek70.22026-04
DeepSeek V2 Lite ChatDeepSeek70.22026-04
DeepSeek V3DeepSeek70.22026-04
DeepSeek V3 0324DeepSeek70.22026-04
DeepSeek V3.1DeepSeek70.22026-04
DeepSeek V3.2DeepSeek70.22026-04
DeepSeek V3.2 ExpDeepSeek70.22026-04
Llama 3.3 70B Instruct NVFP4NVIDIA68.22026-04
Llama-3.3-70B-InstructMeta68.22026-04
Mistral Large (latest)Mistral AI65.82026-04
Mistral Large 2.1Mistral AI65.82026-04
Mistral Large 3Mistral AI65.82026-04
Qwen3 32BAlibaba / Qwen Team65.52026-04
Qwen3 32B AWQAlibaba / Qwen Team65.52026-04
Qwen3 32B NVFP4NVIDIA65.52026-04
Command R+Cohere60.52026-04
phi 4Microsoft58.22026-04
Phi 4 mini instructMicrosoft58.22026-04
Phi 4 multimodal instructMicrosoft58.22026-04

Data

This page as JSON · Edit on GitHub