Moral Stories

lm-eval ranks a crowd-written moral action against an immoral one, given a social norm, situation and intention from the 12k-story Moral Stories corpus.

Also known as: MoralStories

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorysafety
Subcategorybinary ranking of a moral versus immoral action given a social norm and context
Page statusunknown
Metricaccuracy (acc); length-normalised accuracy (acc_norm)
Directionhigher_is_better
Unit%
Dataset size12000
Dataset licenceMIT (demelin/moral_stories card and GitHub LICENSE; LabHC card states no licence field)
PublisherAllen Institute for AI; University of Edinburgh; University of Washington

What it measures

This id is EleutherAI lm-evaluation-harness task moral_stories, not the paper's original generation suite and not BIG-bench moral_permissibility. Each item is an English seven-part story. The harness concatenates the norm, situation and intention, then asks which of two action sentences is more likely: the crowd-written moral action or the immoral action. The labelled target is always the moral action. The original EMNLP 2021 work instead asked models to generate actions, consequences or norms under those constraints.

Task format

Two-way multiple_choice over moral_action versus immoral_action. Context is the capitalised norm, situation and intention. Zero extra few-shot examples in the YAML. English text.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub