Inspect eval of animal-welfare moral reasoning across 13 dimensions; the public set is 115 questions, up from the paper's original 26.
unassessed
| Category | safety |
|---|---|
| Subcategory | LLM-graded moral reasoning about animal welfare (13 dimensions) |
| Page status | active |
| Metric | overall_mean (also dimension_normalized_avg and avg_by_dimension) |
| Direction | higher_is_better |
| Unit | 0-1 |
| Dataset size | 115 |
| Dataset licence | CC-BY-NC-4.0 |
| Publisher | Compassion Aligned Machine Learning (CaML) and Sentient Futures |
ANIMA asks a model to answer open-ended questions about animal welfare, then grades the reply on up to 13 ethical dimensions such as moral consideration, harm minimisation, sentience, prejudice, scope, evidence, and control questions. The paper's original 26 items are English. The public Hugging Face questions split is 115 rows and includes many non-English prompts. A refusal that never engages the scenario is meant to score poorly. It is not the separate 2025 "What do Large Language Models Say About Animals?" paper (arXiv:2503.04804), which uses another question set.
Open-ended generation (`inspect_ai.solver.generate`). Each question carries dimension tags and optional `{{variable}}` slots that the scorer expands. Default epochs is 5. Optional `languages` filter; null language in the file is treated as English. Scoring is model-graded per dimension, then weighted.
No model card in ModelSpec reports this benchmark yet.