Accuracy on MMLU's moral scenarios questions, one of 57 subject tests of academic and professional knowledge.
unassessed
| Category | knowledge |
|---|---|
| Subcategory | Humanities |
| Page status | active |
| Metric | accuracy |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 895 |
| Dataset licence | MIT |
| Publisher | UC Berkeley |
Short everyday scenarios describing an action, each paired with four judgements of whether that action is morally permissible or wrong under ordinary moral standards; it is MMLU's largest single subject test. Framed as four-option multiple-choice questions and scored zero-shot or few-shot by exact match against the labelled option, as one of the 57 subject subsets that make up the MMLU benchmark.
Four-option multiple-choice question answering (A-D), one correct answer, zero-shot or few-shot.
| Model | Provider | Score | Card as of |
|---|---|---|---|
| Nous Hermes 2 Yi 34B | Nous Research | 71.1 | 2024-07 |
| Meta Llama 3 70B | Meta | 70.8 | 2024-07 |
| Meta Llama 3 70B Instruct | Meta | 70.8 | 2026-04 |
| Meta Llama 3 70B Instruct | Nous Research | 70.8 | 2026-04 |
| Yi 34B Chat | 01.AI | 70.2 | 2024-07 |
| Yi 1.5 34B Chat | 01.AI | 68.2 | 2026-04 |
| Yi 1.5 34B Chat 16K | 01.AI | 68.2 | 2024-07 |
| Yi 34B 200K | 01.AI | 67.6 | 2024-07 |
| Yi 1.5 34B 32K | 01.AI | 67.2 | 2024-07 |
| Mixtral 8x22B Instruct v0.1 | Mistral AI | 66.0 | 2026-04 |
| Yi 1.5 34B | 01.AI | 64.7 | 2026-04 |
| Phi 3 mini 4K instruct | Microsoft | 58.5 | 2024-07 |
| Phi 3 mini 128K instruct | Microsoft | 58.4 | 2024-07 |
| Nous Hermes 2 Mixtral 8x7B DPO | Nous Research | 57.2 | 2024-07 |
| Yi 1.5 9B Chat | 01.AI | 55.0 | 2024-07 |
| Yi 1.5 9B Chat 16K | 01.AI | 54.6 | 2024-07 |
| Yi 1.5 9B | 01.AI | 48.6 | 2024-07 |
| Mixtral 8x7B Instruct v0.1 | Mistral AI | 46.0 | 2026-04 |
| Yi 1.5 9B 32K | 01.AI | 45.8 | 2024-07 |
| Hermes 2 Theta Llama 3 8B | Nous Research | 43.9 | 2024-07 |
| Hermes 2 Pro Llama 3 8B | Nous Research | 43.7 | 2024-07 |
| Meta Llama 3 8B Instruct | Meta | 43.7 | 2024-07 |
| Meta Llama 3 8B Instruct | Nous Research | 43.7 | 2024-07 |
| Yi 6B | 01.AI | 42.5 | 2024-07 |
| Yi 6B Chat | 01.AI | 42.5 | 2024-07 |
| Meta Llama 3 8B | Meta | 41.3 | 2024-07 |
| Meta Llama 3 8B | Nous Research | 41.3 | 2024-07 |
| gemma 7B it | Google DeepMind | 40.3 | 2024-07 |
| Mixtral 8x7B v0.1 | Mistral AI | 40.1 | 2026-04 |
| Yi 9B | 01.AI | 40.0 | 2024-07 |
| Mistral 7B v0.3 | Mistral AI | 39.8 | 2024-07 |
| mistral 7B v0.3 bnb 4bit | Unsloth | 39.8 | 2024-07 |
| Yi 1.5 6B Chat | 01.AI | 38.2 | 2024-07 |
| Nous Hermes 2 SOLAR 10.7B | Nous Research | 34.9 | 2024-07 |
| Yi 1.5 6B | 01.AI | 33.2 | 2024-07 |
| Mistral 7B Instruct v0.2 | Mistral AI | 31.2 | 2024-07 |
| deepseek llm 7B base | DeepSeek | 30.6 | 2024-07 |
| deepseek llm 7B chat | DeepSeek | 30.6 | 2024-07 |
| phi 2 | Microsoft | 29.9 | 2024-07 |
| Qwen2 1.5B Instruct | Alibaba / Qwen Team | 29.7 | 2024-07 |
| deepseek coder 6.7B base | DeepSeek | 28.9 | 2024-07 |
| deepseek coder 1.3B base | DeepSeek | 27.3 | 2024-07 |
| deepseek coder 1.3B instruct | DeepSeek | 27.3 | 2024-07 |
| falcon 40B | TII | 27.0 | 2024-07 |
| deepseek coder 6.7B instruct | DeepSeek | 25.8 | 2024-07 |
| Qwen2 0.5B Instruct | Alibaba / Qwen Team | 25.6 | 2024-07 |
| gemma 2B it | Google DeepMind | 25.3 | 2024-07 |
| OLMo 1B hf | Allen AI | 24.7 | 2024-07 |
| gemma 2B | Google DeepMind | 23.7 | 2024-07 |
| chatglm2 6B | Zhipu AI | 23.6 | 2024-07 |