5,694 prompts judged against a 314-category safety taxonomy built from real government regulations and company policies; distinct from two other, unrelated benchmarks also named AIR-Bench.
unassessed
| Category | safety |
|---|---|
| Subcategory | regulation-grounded risk-category safety benchmark |
| Page status | active |
| Metric | AIR score (judge-scored safety compliance) |
| Direction | higher_is_better |
| Unit | score (0-1) |
| Dataset size | 5694 |
| Dataset licence | CC-BY-4.0 |
AIR-Bench 2024 tests whether a model's responses align with safety expectations drawn directly from real government regulations and company usage policies, rather than from researcher intuition about what safety should mean. Each of its 5,694 prompts targets one of 314 fine-grained risk categories, organised into a four-level taxonomy built by decomposing 8 government regulations (EU, US and China) and 16 AI-company usage policies. A model is scored on how it responds to a risk-eliciting prompt: whether it refuses, redirects, or complies with a request that sits in one of these regulation-derived categories. This page documents AIR-Bench 2024, the AI-safety benchmark from arXiv 2407.17436 — see "Lineage" for the unrelated benchmarks that share its name.
Single-turn text prompt drawn from one of 314 risk categories; the model's free-text response is graded by an LLM judge against a category-specific rubric on a three-point scale.
No model card in ModelSpec reports this benchmark yet.