AIR-Bench 2024

5,694 prompts judged against a 314-category safety taxonomy built from real government regulations and company policies; distinct from two other, unrelated benchmarks also named AIR-Bench.

Also known as: AIR-Bench, AIRBench 2024, AI Risk Benchmark

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorysafety
Subcategoryregulation-grounded risk-category safety benchmark
Page statusactive
MetricAIR score (judge-scored safety compliance)
Directionhigher_is_better
Unitscore (0-1)
Dataset size5694
Dataset licenceCC-BY-4.0

What it measures

AIR-Bench 2024 tests whether a model's responses align with safety expectations drawn directly from real government regulations and company usage policies, rather than from researcher intuition about what safety should mean. Each of its 5,694 prompts targets one of 314 fine-grained risk categories, organised into a four-level taxonomy built by decomposing 8 government regulations (EU, US and China) and 16 AI-company usage policies. A model is scored on how it responds to a risk-eliciting prompt: whether it refuses, redirects, or complies with a request that sits in one of these regulation-derived categories. This page documents AIR-Bench 2024, the AI-safety benchmark from arXiv 2407.17436 — see "Lineage" for the unrelated benchmarks that share its name.

Task format

Single-turn text prompt drawn from one of 314 risk categories; the model's free-text response is graded by an LLM judge against a category-specific rubric on a three-point scale.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub