Arena Elo — Vision

A separate Elo leaderboard for vision-language models, built only from Arena battles whose prompt included an image.

Also known as: Chatbot Arena Vision leaderboard, Multimodal Arena, LMArena Vision

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryhuman-preference
Subcategorypairwise human preference, image+text
Page statusactive
MetricElo
Directionhigher_is_better
PublisherLMSYS (Large Model Systems Organization), UC Berkeley Sky Computing Lab (founding org); operates today as Arena (formerly LMArena)

What it measures

arena_elo_vision is a separate Arena leaderboard computed only from anonymous battles whose prompt included an image: two vision-language models are compared head-to-head on the same image-plus-text input, and a voter picks the preferred response before identities are revealed. It launched in June 2024 as the Multimodal ("Vision") Arena, drawing over 17,000 votes across more than 60 languages in its first two weeks, on tasks the team observed spanning captioning, visual math questions, document understanding and meme explanation. It is rated with the same Bradley-Terry pipeline as the Text arena, but over its own, separate vote pool.

Task format

Anonymous, randomized side-by-side chat where the user's prompt includes at least one image; a user votes for the preferred response.

Models reporting this benchmark

These figures come from the model cards, which carry one collection date per card and no per-score attribution. They are shown as reported, not as verified evidence.
ModelProviderScoreCard as of
GPT-4oOpenAI1310.02026-04
GPT-4o (2024-05-13)OpenAI1310.02026-04
GPT-4o (2024-08-06)OpenAI1310.02026-04
GPT-4o (2024-11-20)OpenAI1310.02026-04
Claude Opus 4.6Anthropic1295.02026-04
GPT-5.2OpenAI1277.92026-04
Claude Sonnet 4.6Anthropic1271.12026-04
GPT-5.1OpenAI1249.12026-04
Gemini 2.5 ProGoogle DeepMind1246.02026-04
Grok 4.20 (Reasoning)xAI1242.82026-04
GPT-5.4 miniOpenAI1241.22026-04
Grok 4.20 Multi-AgentxAI1228.42026-04
Gemini 2.5 Flash Preview 09-25Google DeepMind1226.12026-04
GPT-5OpenAI1225.22026-04
o3OpenAI1217.62026-04
Qwen3 235B-A22BAlibaba / Qwen Team1214.62026-04
GPT-4.1OpenAI1213.52026-04
Gemini 2.5 FlashGoogle DeepMind1213.42026-04
Claude Opus 4Anthropic1207.52026-04
Claude Sonnet 4Anthropic1207.02026-04
GPT-4.1 miniOpenAI1202.12026-04
o4-miniOpenAI1201.82026-04
Claude Sonnet 3.7Anthropic1196.12026-04
o1OpenAI1192.92026-04
GPT-5.4 nanoOpenAI1191.82026-04
Gemini 2.5 Flash LiteGoogle DeepMind1188.02026-04
Grok 4xAI1182.02026-04
GPT-5 MiniOpenAI1181.92026-04
Gemini 1.5 ProGoogle DeepMind1178.62026-04
Grok 4.1 FastxAI1173.72026-04
Gemini 2.5 Flash Lite Preview 09-25Google DeepMind1173.22026-04
Gemini 2.0 FlashGoogle DeepMind1170.62026-04
GLM 4.6VZhipu AI1163.1undated
Claude Sonnet 3.5 v2Anthropic1160.42026-04
Gemma 3 27BGoogle DeepMind1157.32026-04
Mistral Medium 3.1Mistral AI1157.32026-04
GLM 4.5VZhipu AI1156.6undated
Mistral Medium 3Mistral AI1156.62026-04
Llama 4 Maverick 17B 128E InstructMeta1146.52026-04
GPT-5 NanoOpenAI1145.62026-04
Claude Sonnet 3.5Anthropic1145.52026-04
Gemini 1.5 FlashGoogle DeepMind1140.12026-04
Mistral Small 3.2Mistral AI1140.12026-04
Gemini 2.0 Flash LiteGoogle DeepMind1134.22026-04
Mistral Small 3.1 24B Instruct 2503Mistral AI1128.12026-04
Llama 4 Scout 17B 16E InstructMeta1127.22026-04
Claude Haiku 3.5Anthropic1126.82026-04
Qwen2.5 72B InstructAlibaba / Qwen Team1122.02026-04
GPT-4 TurboOpenAI1112.52026-04
GPT-4o miniOpenAI1097.82026-04
GPT-4.1 nanoOpenAI1088.92026-04
Gemini 1.5 Flash-8BGoogle DeepMind1070.62026-04
Claude Opus 3Anthropic1062.62026-04
Claude Sonnet 3Anthropic1017.72026-04
Claude Haiku 3Anthropic1000.82026-04

Data

This page as JSON · Edit on GitHub