A machine-generated dataset of 274,000 toxic and benign statements about 13 minority groups, used to test whether models can catch subtle, implicit hate speech.
unassessed
| Category | safety |
|---|---|
| Subcategory | implicit hate speech detection |
| Page status | active |
| Metric | classification accuracy / toxicity rate |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 274000 |
| Dataset licence | CC BY 4.0 |
| Publisher | Microsoft Research |
ToxiGen targets implicit and adversarial toxicity - text that is hateful in effect without using slurs or other easily-flagged surface markers - directed at 13 minority groups. It was built by prompting a large language model with a demonstration-based framework and an adversarial classifier-in-the-loop decoding method, producing both toxic and superficially similar benign statements that are hard to tell apart on wording alone. As a model benchmark, it is typically used to test how well a model or a classifier trained on its outputs can separate implicitly toxic text from benign text about the same groups, or how a chat model behaves when asked to complete or discuss ToxiGen-style prompts.
Binary classification of statements as toxic or benign (as a classifier benchmark), or completion/discussion of ToxiGen-style prompts scored for toxic output (as a generation safety check).
| Model | Provider | Score | Card as of |
|---|---|---|---|
| Llama Guard 3 8B | Meta | 97.5 | 2026-04 |
| Llama Guard 3 8B INT8 | Meta | 97.5 | 2026-04 |
| shieldgemma 27B | Google DeepMind | 96.8 | 2026-04 |
| shieldgemma 9B | Google DeepMind | 96.1 | 2026-04 |
| Llama Guard 3 1B | Meta | 95.8 | 2026-04 |
| Claude Opus 4 | Anthropic | 95.1 | 2026-04 |
| Claude Opus 4.6 | Anthropic | 95.1 | 2026-04 |
| Claude Sonnet 3.5 | Anthropic | 94.5 | 2026-04 |
| Claude Sonnet 3.5 v2 | Anthropic | 94.5 | 2026-04 |
| Claude Sonnet 4 | Anthropic | 94.5 | 2026-04 |
| Claude Sonnet 4.5 | Anthropic | 94.5 | 2026-04 |
| Claude Sonnet 4.5 (latest) | Anthropic | 94.5 | 2026-04 |
| Claude Opus 3 | Anthropic | 93.8 | 2026-04 |
| GPT-4 | OpenAI | 93.5 | 2026-04 |
| GPT-4.1 | OpenAI | 93.5 | 2026-04 |
| GPT-4.1 mini | OpenAI | 93.5 | 2026-04 |
| GPT-4.1 nano | OpenAI | 93.5 | 2026-04 |
| GPT-4 Turbo | OpenAI | 93.2 | 2026-04 |
| Claude Sonnet 3 | Anthropic | 92.8 | 2026-04 |
| GPT-4o | OpenAI | 92.8 | 2026-04 |
| GPT-4o (2024-05-13) | OpenAI | 92.8 | 2026-04 |
| GPT-4o (2024-08-06) | OpenAI | 92.8 | 2026-04 |
| GPT-4o (2024-11-20) | OpenAI | 92.8 | 2026-04 |
| GPT-4o mini | OpenAI | 92.8 | 2026-04 |
| Claude Haiku 3.5 | Anthropic | 92.5 | 2026-04 |
| Claude Haiku 3.5 (latest) | Anthropic | 92.5 | 2026-04 |
| Gemini 2.0 Flash | Google DeepMind | 91.8 | 2026-04 |
| Claude Haiku 3 | Anthropic | 91.5 | 2026-04 |
| Gemini 2.5 Pro | Google DeepMind | 91.5 | 2026-04 |
| Gemini 2.5 Pro Preview 05-06 | Google DeepMind | 91.5 | 2026-04 |
| Gemini 2.5 Pro Preview 06-05 | Google DeepMind | 91.5 | 2026-04 |
| Gemini 2.5 Pro Preview TTS | Google DeepMind | 91.5 | 2026-04 |
| Gemini 1.5 Pro | Google DeepMind | 91.2 | 2026-04 |
| Gemini 1.5 Flash | Google DeepMind | 90.5 | 2026-04 |
| Gemini 1.5 Flash-8B | Google DeepMind | 90.5 | 2026-04 |
| Pixtral Large (latest) | Mistral AI | 90.5 | 2026-04 |
| Llama 4 Maverick 17B 128E Instruct | Meta | 89.8 | 2026-04 |
| Llama-4-Maverick-17B-128E-Instruct-FP8 | Meta | 89.8 | 2026-04 |
| Gemini 2.0 Flash Lite | Google DeepMind | 89.5 | 2026-04 |
| Llama 4 Scout 17B 16E | Meta | 88.5 | 2026-04 |
| Llama 4 Scout 17B 16E Instruct | Meta | 88.5 | 2026-04 |
| Llama-4-Scout-17B-16E-Instruct-FP8 | Meta | 88.5 | 2026-04 |
| Llama 3.3 70B Instruct NVFP4 | NVIDIA | 88.2 | 2026-04 |
| Llama-3.3-70B-Instruct | Meta | 88.2 | 2026-04 |
| Qwen2.5-VL 72B Instruct | Alibaba / Qwen Team | 87.8 | 2026-04 |
| Command R+ | Cohere | 87.5 | 2026-04 |
| Llama 3.2 90B Vision | Meta | 87.5 | 2026-04 |
| Llama 3.2 90B Vision Instruct | Meta | 87.5 | 2026-04 |
| Pixtral 12B | Mistral AI | 86.8 | 2026-04 |
| Llama 3.1 70B | Meta | 86.5 | 2026-04 |
| Llama 3.1 70B Instruct | Meta | 86.5 | 2026-04 |
| Mistral Large (latest) | Mistral AI | 86.2 | 2026-04 |
| Mistral Large 2.1 | Mistral AI | 86.2 | 2026-04 |
| Mistral Large 3 | Mistral AI | 86.2 | 2026-04 |
| Llama 3.2 11B Vision | Meta | 85.8 | 2026-04 |
| Llama 3.2 11B Vision Instruct | Meta | 85.8 | 2026-04 |
| Qwen2.5 72B Instruct | Alibaba / Qwen Team | 84.8 | 2026-04 |
| DeepSeek Chat | DeepSeek | 82.5 | 2026-04 |
| DeepSeek V2 | DeepSeek | 82.5 | 2026-04 |
| DeepSeek V2 Lite | DeepSeek | 82.5 | 2026-04 |
| DeepSeek V2 Lite Chat | DeepSeek | 82.5 | 2026-04 |
| DeepSeek V3 | DeepSeek | 82.5 | 2026-04 |
| DeepSeek V3 0324 | DeepSeek | 82.5 | 2026-04 |
| DeepSeek V3.1 | DeepSeek | 82.5 | 2026-04 |
| DeepSeek V3.2 | DeepSeek | 82.5 | 2026-04 |
| DeepSeek V3.2 Exp | DeepSeek | 82.5 | 2026-04 |
| Llama 3.1 8B | Meta | 82.1 | 2026-04 |
| Llama 3.1 8B Instruct | Meta | 82.1 | 2026-04 |
| Llama 3.1 8B Instruct | Unsloth | 82.1 | 2026-04 |
| Llama 3.1 8B Instruct FP8 | NVIDIA | 82.1 | 2026-04 |
| Llama 3.1 8B Instruct NVFP4 | NVIDIA | 82.1 | 2026-04 |