Arena Elo — Style Control

A style-adjusted Arena ranking that regresses out response length and markdown formatting so ratings lean more on substance than presentation.

Also known as: Chatbot Arena Style Control, LMArena Style Control, Arena Score, style-controlled

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryhuman-preference
Subcategorypairwise human preference, style-adjusted
Page statusactive
MetricElo
Directionhigher_is_better
PublisherLMSYS (Large Model Systems Organization), UC Berkeley Sky Computing Lab (founding org); operates today as Arena (formerly LMArena)

What it measures

arena_elo_style_control applies a style adjustment to the same votes and the same Bradley-Terry fit used elsewhere in the arena_elo family, rather than drawing on a different vote pool. LMSYS adds response length and markdown formatting (header, bold and list counts, each expressed as a normalised difference between the two compared responses) as extra regressors in the logistic regression that produces Arena ratings, so the resulting model coefficients are adjusted for — controlled for — those style effects instead of reflecting them. It can be layered onto Overall or onto a category such as Hard Prompts.

Task format

Same anonymous, randomized side-by-side text chat as the underlying category; ratings are recomputed with response-length and markdown-count features included in the regression.

Models reporting this benchmark

These figures come from the model cards, which carry one collection date per card and no per-score attribution. They are shown as reported, not as verified evidence.
ModelProviderScoreCard as of
Claude Opus 4.6Anthropic1502.82026-04
Grok 4.20 (Reasoning)xAI1490.42026-04
GPT-5.4OpenAI1483.62026-04
GPT-5.2OpenAI1477.02026-04
Grok 4.20 Multi-AgentxAI1474.52026-04
Claude Opus 4.5Anthropic1473.12026-04
Grok 4xAI1470.72026-04
GLM 5.1Z.ai (Zhipu AI)1467.42026-04
Claude Sonnet 4.6Anthropic1462.22026-04
GPT-5.4 miniOpenAI1457.82026-04
GPT-5.3 Codex SparkOpenAI1456.42026-04
GLM 5Zhipu AI1455.6undated
GPT-5.1OpenAI1454.62026-04
Claude Sonnet 4.5Anthropic1451.92026-04
Gemma 4 31BGoogle DeepMind1451.22026-04
Claude Opus 4.1Anthropic1448.62026-04
Gemini 2.5 ProGoogle DeepMind1448.22026-04
GPT-4oOpenAI1443.02026-04
GLM 4.7Zhipu AI1442.7undated
gemma 4 26B A4B itGoogle DeepMind1437.92026-04
GPT-5OpenAI1433.42026-04
o3OpenAI1431.22026-04
Grok 4.1 FastxAI1431.12026-04
GLM 4.6Zhipu AI1425.8undated
DeepSeek V3.2DeepSeek1424.42026-04
Claude Opus 4Anthropic1423.72026-04
DeepSeek V3.2 ExpDeepSeek1422.82026-04
DeepSeek R1 0528DeepSeek1421.72026-04
Grok 4 Fast (Non-Reasoning)xAI1420.72026-04
DeepSeek V3.1DeepSeek1417.92026-04
Qwen3 235B-A22BAlibaba / Qwen Team1415.82026-04
Mistral Large 3Mistral AI1415.02026-04
GPT-4.1OpenAI1413.02026-04
Grok 3xAI1411.72026-04
Gemini 2.5 FlashGoogle DeepMind1411.02026-04
GLM 4.5Zhipu AI1410.9undated
Mistral Medium 3.1Mistral AI1410.32026-04
Claude Haiku 4.5Anthropic1407.22026-04
Gemini 2.5 Flash Preview 09-25Google DeepMind1405.12026-04
GPT-5.4 nanoOpenAI1402.72026-04
o1OpenAI1401.52026-04
Claude Sonnet 4Anthropic1398.52026-04
DeepSeek R1DeepSeek1397.52026-04
DeepSeek V3 0324DeepSeek1394.62026-04
o4-miniOpenAI1389.72026-04
GPT-5 MiniOpenAI1389.52026-04
o1-previewOpenAI1387.62026-04
Qwen3-Coder 480B-A35B InstructAlibaba / Qwen Team1387.42026-04
Claude Sonnet 3.7Anthropic1386.32026-04
Mistral Medium 3Mistral AI1386.22026-04
GPT-4.1 miniOpenAI1382.12026-04
Gemini 2.5 Flash Lite Preview 09-25Google DeepMind1380.02026-04
GLM 4.6VZhipu AI1377.9undated
Gemini 2.5 Flash LiteGoogle DeepMind1374.32026-04
GLM 4.5 AirZhipu AI1372.8undated
Claude Sonnet 3.5 v2Anthropic1371.42026-04
GLM 4.7 FlashZhipu AI1368.7undated
Gemma 3 27BGoogle DeepMind1365.12026-04
o3-miniOpenAI1363.22026-04
Grok 3 MinixAI1362.72026-04
Gemini 2.0 FlashGoogle DeepMind1360.02026-04
DeepSeek V3DeepSeek1358.22026-04
Mistral Small 3.2Mistral AI1357.02026-04
Command ACohere1353.42026-04
GLM 4.5VZhipu AI1353.3undated
Gemini 2.0 Flash LiteGoogle DeepMind1352.82026-04
Gemini 1.5 ProGoogle DeepMind1350.62026-04
Qwen3 32BAlibaba / Qwen Team1347.02026-04
GPT-4o (2024-05-13)OpenAI1345.12026-04
Gemma 3 12BGoogle DeepMind1341.42026-04
Claude Sonnet 3.5Anthropic1341.32026-04
o1-miniOpenAI1336.62026-04
GPT-5 NanoOpenAI1336.42026-04
Grok 2xAI1334.82026-04
GPT-4o (2024-08-06)OpenAI1334.32026-04
Llama 3.1 405B InstructMeta1334.22026-04
Llama 3.1 405B Instruct FP8Meta1332.42026-04
Qwen3 30B-A3BAlibaba / Qwen Team1327.32026-04
Llama 4 Maverick 17B 128E InstructMeta1326.62026-04
GPT-4 TurboOpenAI1323.42026-04
Claude Haiku 3.5Anthropic1322.42026-04
Llama 4 Scout 17B 16E InstructMeta1321.92026-04
GPT-4.1 nanoOpenAI1321.42026-04
Claude Opus 3Anthropic1320.72026-04
Llama-3.3-70B-InstructMeta1318.02026-04
GPT-4o miniOpenAI1317.22026-04
Gemini 1.5 FlashGoogle DeepMind1309.12026-04
Mistral Large 2.1Mistral AI1304.72026-04
Magistral Medium (latest)Mistral AI1303.02026-04
Gemma 3 4BGoogle DeepMind1302.82026-04
Mistral Small 3.1 24B Instruct 2503Mistral AI1302.72026-04
Qwen2.5 72B InstructAlibaba / Qwen Team1302.32026-04
Llama 3.1 70B InstructMeta1292.82026-04
gemma 2 27B itGoogle DeepMind1287.62026-04
Claude Sonnet 3Anthropic1279.82026-04
Command R+Cohere1275.52026-04
Mistral Small 24B Instruct 2501Mistral AI1273.52026-04
gemma 2 9B itGoogle DeepMind1265.02026-04
Claude Haiku 3Anthropic1259.92026-04
Gemini 1.5 Flash-8BGoogle DeepMind1258.12026-04
phi 4Microsoft1255.42026-04
Command RCohere1249.12026-04
Llama 3.1 8B InstructMeta1211.02026-04

Data

This page as JSON · Edit on GitHub