BBQ (Bias Benchmark for QA)

Multiple-choice QA benchmark testing whether models default to social stereotypes under ambiguous context and can override them once context disambiguates the answer.

Also known as: Bias Benchmark for Question Answering

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorysafety
Subcategorysocial bias in question answering
Page statusactive
Metricaccuracy (plus a separate bias score)
Directionhigher_is_better
Unit%
Dataset size58492
Dataset licenceCC BY 4.0
PublisherNew York University, Machine Learning for Language group

What it measures

BBQ tests whether a language model falls back on social stereotypes when it lacks the information to answer a question, and whether it can set that stereotype aside once the missing fact is supplied. Each item gives the model a short context and a question about two people or groups mentioned in it, with three answer choices: one social group, the other, or "unknown". The nine bias categories are age, disability status, gender identity, nationality, physical appearance, race or ethnicity, religion, socio-economic status, and sexual orientation, plus two intersectional categories combining race with gender and race with socio-economic status. Every item comes in an ambiguous version, which gives no information that would let a careful reader pick a group over "unknown", and a disambiguated version that adds one sentence resolving the question in favour of one of the two groups.

Task format

Three-way multiple-choice QA (two named social groups plus "unknown"), zero- or few-shot, US English

Models reporting this benchmark

These figures come from the model cards, which carry one collection date per card and no per-score attribution. They are shown as reported, not as verified evidence.
ModelProviderScoreCard as of
Claude Opus 4Anthropic88.22026-04
Claude Opus 4.6Anthropic88.22026-04
Claude Sonnet 4Anthropic87.52026-04
Claude Sonnet 4.5Anthropic87.52026-04
Claude Sonnet 4.5 (latest)Anthropic87.52026-04
Claude Sonnet 3.5Anthropic87.22026-04
Claude Sonnet 3.5 v2Anthropic87.22026-04
Claude Opus 3Anthropic86.52026-04
GPT-4OpenAI86.52026-04
GPT-4.1OpenAI86.52026-04
GPT-4.1 miniOpenAI86.52026-04
GPT-4.1 nanoOpenAI86.52026-04
GPT-4oOpenAI85.82026-04
GPT-4o (2024-05-13)OpenAI85.82026-04
GPT-4o (2024-08-06)OpenAI85.82026-04
GPT-4o (2024-11-20)OpenAI85.82026-04
GPT-4o miniOpenAI85.82026-04
GPT-4 TurboOpenAI85.52026-04
Claude Sonnet 3Anthropic85.22026-04
Claude Haiku 3.5Anthropic84.52026-04
Claude Haiku 3.5 (latest)Anthropic84.52026-04
Gemini 2.5 ProGoogle DeepMind84.22026-04
Gemini 2.5 Pro Preview 05-06Google DeepMind84.22026-04
Gemini 2.5 Pro Preview 06-05Google DeepMind84.22026-04
Gemini 2.5 Pro Preview TTSGoogle DeepMind84.22026-04
Claude Haiku 3Anthropic83.52026-04
Gemini 2.0 FlashGoogle DeepMind83.22026-04
Gemini 1.5 ProGoogle DeepMind82.82026-04
Pixtral Large (latest)Mistral AI82.22026-04
Llama 4 Maverick 17B 128E InstructMeta81.82026-04
Llama-4-Maverick-17B-128E-Instruct-FP8Meta81.82026-04
Gemini 1.5 FlashGoogle DeepMind81.52026-04
Gemini 1.5 Flash-8BGoogle DeepMind81.52026-04
Gemini 2.0 Flash LiteGoogle DeepMind80.82026-04
Llama 4 Scout 17B 16EMeta80.52026-04
Llama 4 Scout 17B 16E InstructMeta80.52026-04
Llama-4-Scout-17B-16E-Instruct-FP8Meta80.52026-04
Llama 3.3 70B Instruct NVFP4NVIDIA80.12026-04
Llama-3.3-70B-InstructMeta80.12026-04
Llama 3.2 90B VisionMeta79.82026-04
Llama 3.2 90B Vision InstructMeta79.82026-04
Qwen2.5-VL 72B InstructAlibaba / Qwen Team79.52026-04
Command R+Cohere79.22026-04
Pixtral 12BMistral AI78.52026-04
Llama 3.1 70BMeta78.22026-04
Llama 3.1 70B InstructMeta78.22026-04
Llama 3.2 11B VisionMeta78.22026-04
Llama 3.2 11B Vision InstructMeta78.22026-04
Mistral Large (latest)Mistral AI77.52026-04
Mistral Large 2.1Mistral AI77.52026-04
Mistral Large 3Mistral AI77.52026-04
Qwen2.5 72B InstructAlibaba / Qwen Team75.22026-04
DeepSeek ChatDeepSeek73.52026-04
DeepSeek V2DeepSeek73.52026-04
DeepSeek V2 LiteDeepSeek73.52026-04
DeepSeek V2 Lite ChatDeepSeek73.52026-04
DeepSeek V3DeepSeek73.52026-04
DeepSeek V3 0324DeepSeek73.52026-04
DeepSeek V3.1DeepSeek73.52026-04
DeepSeek V3.2DeepSeek73.52026-04
DeepSeek V3.2 ExpDeepSeek73.52026-04
Llama 3.1 8BMeta72.82026-04
Llama 3.1 8B InstructMeta72.82026-04
Llama 3.1 8B InstructUnsloth72.82026-04
Llama 3.1 8B Instruct FP8NVIDIA72.82026-04
Llama 3.1 8B Instruct NVFP4NVIDIA72.82026-04

Data

This page as JSON · Edit on GitHub