MT-Bench

Eighty multi-turn chat questions graded by an LLM judge as a fast, repeatable stand-in for human conversational preference.

Also known as: MT-bench

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryhuman-preference
Subcategorymulti-turn chat quality
Page statusactive
MetricLLM judge score
Directionhigher_is_better
Unitpoints
Dataset size80
PublisherLMSYS Org

What it measures

MT-Bench gives a chat model 80 open-ended questions spread across eight categories - writing, role-play, extraction, reasoning, math, coding, and two knowledge categories covering STEM and humanities or social science. Each question carries one scripted follow-up, so the model has to hold context across two turns rather than answer a single isolated prompt. The benchmark targets general chat quality and instruction-following on subjective, everyday requests, not narrow factual recall.

Task format

Open-ended two-turn chat completion, graded after generation by a separate judge model rather than by an exact-match answer key.

Models reporting this benchmark

These figures come from the model cards, which carry one collection date per card and no per-score attribution. They are shown as reported, not as verified evidence.
ModelProviderScoreCard as of
Claude Mythos PreviewAnthropic9.42026-04
Claude Opus 4Anthropic9.42026-04
Claude Opus 4.6Anthropic9.42026-04
GPT-4OpenAI9.32026-04
GPT-4.1OpenAI9.32026-04
GPT-4.1 miniOpenAI9.32026-04
GPT-4.1 nanoOpenAI9.32026-04
Claude Sonnet 4Anthropic9.22026-04
Claude Sonnet 4.5Anthropic9.22026-04
Claude Sonnet 4.5 (latest)Anthropic9.22026-04
Gemini 2.5 ProGoogle DeepMind9.22026-04
Gemini 2.5 Pro Preview 05-06Google DeepMind9.22026-04
Gemini 2.5 Pro Preview 06-05Google DeepMind9.22026-04
Gemini 2.5 Pro Preview TTSGoogle DeepMind9.22026-04
GPT-4oOpenAI9.12026-04
GPT-4o (2024-05-13)OpenAI9.12026-04
GPT-4o (2024-08-06)OpenAI9.12026-04
GPT-4o (2024-11-20)OpenAI9.12026-04
GPT-4o miniOpenAI9.12026-04
DeepSeek R1DeepSeek9.02026-04
DeepSeek R1 0528DeepSeek9.02026-04
DeepSeek R1 0528 NVFP4 v2NVIDIA9.02026-04
DeepSeek ReasonerDeepSeek9.02026-04
Qwen 3 235B InstructCerebras8.92026-04
Qwen3 235B-A22BAlibaba / Qwen Team8.92026-04
DeepSeek ChatDeepSeek8.82026-04
DeepSeek V2DeepSeek8.82026-04
DeepSeek V2 LiteDeepSeek8.82026-04
DeepSeek V2 Lite ChatDeepSeek8.82026-04
DeepSeek V3DeepSeek8.82026-04
DeepSeek V3 0324DeepSeek8.82026-04
DeepSeek V3.1DeepSeek8.82026-04
DeepSeek V3.2DeepSeek8.82026-04
DeepSeek V3.2 ExpDeepSeek8.82026-04
Llama 3.3 70B Instruct NVFP4NVIDIA8.62026-04
Llama-3.3-70B-InstructMeta8.62026-04
Mistral Large (latest)Mistral AI8.52026-04
Mistral Large 2.1Mistral AI8.52026-04
Mistral Large 3Mistral AI8.52026-04
Qwen3 32BAlibaba / Qwen Team8.52026-04
Qwen3 32B AWQAlibaba / Qwen Team8.52026-04
Qwen3 32B NVFP4NVIDIA8.52026-04
Command R+Cohere8.32026-04
phi 4Microsoft8.22026-04
Phi 4 mini instructMicrosoft8.22026-04
Phi 4 multimodal instructMicrosoft8.22026-04

Data

This page as JSON · Edit on GitHub