Eighty multi-turn chat questions graded by an LLM judge as a fast, repeatable stand-in for human conversational preference.
unassessed
| Category | human-preference |
|---|---|
| Subcategory | multi-turn chat quality |
| Page status | active |
| Metric | LLM judge score |
| Direction | higher_is_better |
| Unit | points |
| Dataset size | 80 |
| Publisher | LMSYS Org |
MT-Bench gives a chat model 80 open-ended questions spread across eight categories - writing, role-play, extraction, reasoning, math, coding, and two knowledge categories covering STEM and humanities or social science. Each question carries one scripted follow-up, so the model has to hold context across two turns rather than answer a single isolated prompt. The benchmark targets general chat quality and instruction-following on subjective, everyday requests, not narrow factual recall.
Open-ended two-turn chat completion, graded after generation by a separate judge model rather than by an exact-match answer key.
| Model | Provider | Score | Card as of |
|---|---|---|---|
| Claude Mythos Preview | Anthropic | 9.4 | 2026-04 |
| Claude Opus 4 | Anthropic | 9.4 | 2026-04 |
| Claude Opus 4.6 | Anthropic | 9.4 | 2026-04 |
| GPT-4 | OpenAI | 9.3 | 2026-04 |
| GPT-4.1 | OpenAI | 9.3 | 2026-04 |
| GPT-4.1 mini | OpenAI | 9.3 | 2026-04 |
| GPT-4.1 nano | OpenAI | 9.3 | 2026-04 |
| Claude Sonnet 4 | Anthropic | 9.2 | 2026-04 |
| Claude Sonnet 4.5 | Anthropic | 9.2 | 2026-04 |
| Claude Sonnet 4.5 (latest) | Anthropic | 9.2 | 2026-04 |
| Gemini 2.5 Pro | Google DeepMind | 9.2 | 2026-04 |
| Gemini 2.5 Pro Preview 05-06 | Google DeepMind | 9.2 | 2026-04 |
| Gemini 2.5 Pro Preview 06-05 | Google DeepMind | 9.2 | 2026-04 |
| Gemini 2.5 Pro Preview TTS | Google DeepMind | 9.2 | 2026-04 |
| GPT-4o | OpenAI | 9.1 | 2026-04 |
| GPT-4o (2024-05-13) | OpenAI | 9.1 | 2026-04 |
| GPT-4o (2024-08-06) | OpenAI | 9.1 | 2026-04 |
| GPT-4o (2024-11-20) | OpenAI | 9.1 | 2026-04 |
| GPT-4o mini | OpenAI | 9.1 | 2026-04 |
| DeepSeek R1 | DeepSeek | 9.0 | 2026-04 |
| DeepSeek R1 0528 | DeepSeek | 9.0 | 2026-04 |
| DeepSeek R1 0528 NVFP4 v2 | NVIDIA | 9.0 | 2026-04 |
| DeepSeek Reasoner | DeepSeek | 9.0 | 2026-04 |
| Qwen 3 235B Instruct | Cerebras | 8.9 | 2026-04 |
| Qwen3 235B-A22B | Alibaba / Qwen Team | 8.9 | 2026-04 |
| DeepSeek Chat | DeepSeek | 8.8 | 2026-04 |
| DeepSeek V2 | DeepSeek | 8.8 | 2026-04 |
| DeepSeek V2 Lite | DeepSeek | 8.8 | 2026-04 |
| DeepSeek V2 Lite Chat | DeepSeek | 8.8 | 2026-04 |
| DeepSeek V3 | DeepSeek | 8.8 | 2026-04 |
| DeepSeek V3 0324 | DeepSeek | 8.8 | 2026-04 |
| DeepSeek V3.1 | DeepSeek | 8.8 | 2026-04 |
| DeepSeek V3.2 | DeepSeek | 8.8 | 2026-04 |
| DeepSeek V3.2 Exp | DeepSeek | 8.8 | 2026-04 |
| Llama 3.3 70B Instruct NVFP4 | NVIDIA | 8.6 | 2026-04 |
| Llama-3.3-70B-Instruct | Meta | 8.6 | 2026-04 |
| Mistral Large (latest) | Mistral AI | 8.5 | 2026-04 |
| Mistral Large 2.1 | Mistral AI | 8.5 | 2026-04 |
| Mistral Large 3 | Mistral AI | 8.5 | 2026-04 |
| Qwen3 32B | Alibaba / Qwen Team | 8.5 | 2026-04 |
| Qwen3 32B AWQ | Alibaba / Qwen Team | 8.5 | 2026-04 |
| Qwen3 32B NVFP4 | NVIDIA | 8.5 | 2026-04 |
| Command R+ | Cohere | 8.3 | 2026-04 |
| phi 4 | Microsoft | 8.2 | 2026-04 |
| Phi 4 mini instruct | Microsoft | 8.2 | 2026-04 |
| Phi 4 multimodal instruct | Microsoft | 8.2 | 2026-04 |