HELM Safety

Stanford CRFM leaderboard that scores a model by averaging its results across five existing safety benchmarks into one 0-1 number.

Also known as: HELM-Safety

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorysafety
Page statusactive
MetricMean score (unweighted average of five normalized per-scenario scores)
Directionhigher_is_better
Unit0-1 scale
PublisherStanford Center for Research on Foundation Models (CRFM)

What it measures

HELM Safety is a Stanford CRFM leaderboard that scores a model's refusal and bias behaviour across five existing safety datasets: BBQ, SimpleSafetyTests, HarmBench, AnthropicRedTeam, and XSTest. Together they probe six risk categories CRFM drew from AI developers' acceptable-use policies: violence, fraud, discrimination, sexual content, harassment, and deception. A model sees single-turn prompts ranging from overtly unsafe requests, through red-teamed jailbreak attempts meant to bypass guardrails, to sensitive-sounding but benign questions; BBQ instead asks multiple-choice questions testing whether the model leans on a social stereotype when context does not support one.

Task format

Single-turn prompts; graded by exact-match accuracy (BBQ) or a two-model LLM-judge harmfulness/helpfulness rating (the other four scenarios).

Models reporting this benchmark

These figures come from the model cards, which carry one collection date per card and no per-score attribution. They are shown as reported, not as verified evidence.
ModelProviderScoreCard as of
Llama Guard 3 8BMeta95.22026-04
Llama Guard 3 8B INT8Meta95.22026-04
shieldgemma 27BGoogle DeepMind94.52026-04
shieldgemma 9BGoogle DeepMind93.22026-04
Llama Guard 3 1BMeta92.82026-04
Claude Opus 4Anthropic92.52026-04
Claude Opus 4.6Anthropic92.52026-04
Claude Sonnet 4Anthropic91.82026-04
Claude Sonnet 4.5Anthropic91.82026-04
Claude Sonnet 4.5 (latest)Anthropic91.82026-04
Claude Sonnet 3.5Anthropic91.52026-04
Claude Sonnet 3.5 v2Anthropic91.52026-04
Claude Opus 3Anthropic90.82026-04
GPT-4OpenAI90.52026-04
GPT-4.1OpenAI90.52026-04
GPT-4.1 miniOpenAI90.52026-04
GPT-4.1 nanoOpenAI90.52026-04
Claude Haiku 3.5Anthropic89.82026-04
Claude Haiku 3.5 (latest)Anthropic89.82026-04
Claude Sonnet 3Anthropic89.52026-04
GPT-4 TurboOpenAI89.52026-04
GPT-4oOpenAI89.22026-04
GPT-4o (2024-05-13)OpenAI89.22026-04
GPT-4o (2024-08-06)OpenAI89.22026-04
GPT-4o (2024-11-20)OpenAI89.22026-04
GPT-4o miniOpenAI89.22026-04
Gemini 2.5 ProGoogle DeepMind88.52026-04
Gemini 2.5 Pro Preview 05-06Google DeepMind88.52026-04
Gemini 2.5 Pro Preview 06-05Google DeepMind88.52026-04
Gemini 2.5 Pro Preview TTSGoogle DeepMind88.52026-04
Claude Haiku 3Anthropic88.22026-04
Gemini 2.0 FlashGoogle DeepMind87.82026-04
Gemini 1.5 ProGoogle DeepMind87.52026-04
Pixtral Large (latest)Mistral AI86.82026-04
Llama 4 Maverick 17B 128E InstructMeta86.52026-04
Llama-4-Maverick-17B-128E-Instruct-FP8Meta86.52026-04
Gemini 1.5 FlashGoogle DeepMind86.22026-04
Gemini 1.5 Flash-8BGoogle DeepMind86.22026-04
Gemini 2.0 Flash LiteGoogle DeepMind85.52026-04
Llama 4 Scout 17B 16EMeta85.22026-04
Llama 4 Scout 17B 16E InstructMeta85.22026-04
Llama-4-Scout-17B-16E-Instruct-FP8Meta85.22026-04
Llama 3.2 90B VisionMeta84.52026-04
Llama 3.2 90B Vision InstructMeta84.52026-04
Llama 3.3 70B Instruct NVFP4NVIDIA84.22026-04
Llama-3.3-70B-InstructMeta84.22026-04
Qwen2.5-VL 72B InstructAlibaba / Qwen Team84.22026-04
Command R+Cohere83.52026-04
Pixtral 12BMistral AI83.52026-04
Mistral Large (latest)Mistral AI82.82026-04
Mistral Large 2.1Mistral AI82.82026-04
Mistral Large 3Mistral AI82.82026-04
Llama 3.1 70BMeta82.52026-04
Llama 3.1 70B InstructMeta82.52026-04
Llama 3.2 11B VisionMeta82.52026-04
Llama 3.2 11B Vision InstructMeta82.52026-04
Qwen2.5 72B InstructAlibaba / Qwen Team80.52026-04
Llama 3.1 8BMeta78.52026-04
Llama 3.1 8B InstructMeta78.52026-04
Llama 3.1 8B InstructUnsloth78.52026-04
Llama 3.1 8B Instruct FP8NVIDIA78.52026-04
Llama 3.1 8B Instruct NVFP4NVIDIA78.52026-04
DeepSeek ChatDeepSeek78.22026-04
DeepSeek V2DeepSeek78.22026-04
DeepSeek V2 LiteDeepSeek78.22026-04
DeepSeek V2 Lite ChatDeepSeek78.22026-04
DeepSeek V3DeepSeek78.22026-04
DeepSeek V3 0324DeepSeek78.22026-04
DeepSeek V3.1DeepSeek78.22026-04
DeepSeek V3.2DeepSeek78.22026-04
DeepSeek V3.2 ExpDeepSeek78.22026-04

Data

This page as JSON · Edit on GitHub