MMLU: Moral Scenarios

Accuracy on MMLU's moral scenarios questions, one of 57 subject tests of academic and professional knowledge.

Also known as: moral_scenarios

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryknowledge
SubcategoryHumanities
Page statusactive
Metricaccuracy
Directionhigher_is_better
Unit%
Dataset size895
Dataset licenceMIT
PublisherUC Berkeley

What it measures

Short everyday scenarios describing an action, each paired with four judgements of whether that action is morally permissible or wrong under ordinary moral standards; it is MMLU's largest single subject test. Framed as four-option multiple-choice questions and scored zero-shot or few-shot by exact match against the labelled option, as one of the 57 subject subsets that make up the MMLU benchmark.

Task format

Four-option multiple-choice question answering (A-D), one correct answer, zero-shot or few-shot.

Models reporting this benchmark

These figures come from the model cards, which carry one collection date per card and no per-score attribution. They are shown as reported, not as verified evidence.
ModelProviderScoreCard as of
Nous Hermes 2 Yi 34BNous Research71.12024-07
Meta Llama 3 70BMeta70.82024-07
Meta Llama 3 70B InstructMeta70.82026-04
Meta Llama 3 70B InstructNous Research70.82026-04
Yi 34B Chat01.AI70.22024-07
Yi 1.5 34B Chat01.AI68.22026-04
Yi 1.5 34B Chat 16K01.AI68.22024-07
Yi 34B 200K01.AI67.62024-07
Yi 1.5 34B 32K01.AI67.22024-07
Mixtral 8x22B Instruct v0.1Mistral AI66.02026-04
Yi 1.5 34B01.AI64.72026-04
Phi 3 mini 4K instructMicrosoft58.52024-07
Phi 3 mini 128K instructMicrosoft58.42024-07
Nous Hermes 2 Mixtral 8x7B DPONous Research57.22024-07
Yi 1.5 9B Chat01.AI55.02024-07
Yi 1.5 9B Chat 16K01.AI54.62024-07
Yi 1.5 9B01.AI48.62024-07
Mixtral 8x7B Instruct v0.1Mistral AI46.02026-04
Yi 1.5 9B 32K01.AI45.82024-07
Hermes 2 Theta Llama 3 8BNous Research43.92024-07
Hermes 2 Pro Llama 3 8BNous Research43.72024-07
Meta Llama 3 8B InstructMeta43.72024-07
Meta Llama 3 8B InstructNous Research43.72024-07
Yi 6B01.AI42.52024-07
Yi 6B Chat01.AI42.52024-07
Meta Llama 3 8BMeta41.32024-07
Meta Llama 3 8BNous Research41.32024-07
gemma 7B itGoogle DeepMind40.32024-07
Mixtral 8x7B v0.1Mistral AI40.12026-04
Yi 9B01.AI40.02024-07
Mistral 7B v0.3Mistral AI39.82024-07
mistral 7B v0.3 bnb 4bitUnsloth39.82024-07
Yi 1.5 6B Chat01.AI38.22024-07
Nous Hermes 2 SOLAR 10.7BNous Research34.92024-07
Yi 1.5 6B01.AI33.22024-07
Mistral 7B Instruct v0.2Mistral AI31.22024-07
deepseek llm 7B baseDeepSeek30.62024-07
deepseek llm 7B chatDeepSeek30.62024-07
phi 2Microsoft29.92024-07
Qwen2 1.5B InstructAlibaba / Qwen Team29.72024-07
deepseek coder 6.7B baseDeepSeek28.92024-07
deepseek coder 1.3B baseDeepSeek27.32024-07
deepseek coder 1.3B instructDeepSeek27.32024-07
falcon 40BTII27.02024-07
deepseek coder 6.7B instructDeepSeek25.82024-07
Qwen2 0.5B InstructAlibaba / Qwen Team25.62024-07
gemma 2B itGoogle DeepMind25.32024-07
OLMo 1B hfAllen AI24.72024-07
gemma 2BGoogle DeepMind23.72024-07
chatglm2 6BZhipu AI23.62024-07

Data

This page as JSON · Edit on GitHub