MMLU: Professional Psychology

MMLU subject subset: Licensing-exam-style questions spanning clinical, developmental and social psychology, psychometrics and professional ethics.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryknowledge
Subcategorypsychology
Page statusactive
Metricaccuracy
Directionhigher_is_better
Unit%
Dataset size612
Dataset licenceMIT
PublisherUC Berkeley (original); Center for AI Safety (current host)

What it measures

Licensing-exam-style questions spanning clinical, developmental and social psychology, psychometrics, and professional ethics. Questions are four-option multiple-choice, drawn from the MMLU test set's "psychology" subcategory within the benchmark's "social sciences" top-level group, and are graded on the single correct labelled option.

Task format

Four-option multiple-choice questions, graded on the single correct labelled option; commonly evaluated 5-shot, consistent with the rest of MMLU.

Models reporting this benchmark

These figures come from the model cards, which carry one collection date per card and no per-score attribution. They are shown as reported, not as verified evidence.
ModelProviderScoreCard as of
Meta Llama 3 70BMeta85.32024-07
Meta Llama 3 70B InstructMeta85.32026-04
Meta Llama 3 70B InstructNous Research85.32026-04
Mixtral 8x22B Instruct v0.1Mistral AI84.02026-04
Yi 34B Chat01.AI82.72024-07
Nous Hermes 2 Yi 34BNous Research82.42024-07
Yi 34B 200K01.AI82.02024-07
Yi 1.5 34B 32K01.AI81.72024-07
Yi 1.5 34B01.AI80.92026-04
Nous Hermes 2 Mixtral 8x7B DPONous Research78.82024-07
Mixtral 8x7B v0.1Mistral AI78.42026-04
Yi 1.5 34B Chat01.AI78.32026-04
Yi 1.5 34B Chat 16K01.AI78.32024-07
Mixtral 8x7B Instruct v0.1Mistral AI76.52026-04
Phi 3 mini 4K instructMicrosoft75.82024-07
Phi 3 mini 128K instructMicrosoft75.32024-07
Yi 1.5 9B01.AI73.92024-07
Meta Llama 3 8BMeta72.42024-07
Meta Llama 3 8BNous Research72.42024-07
Yi 1.5 9B 32K01.AI71.72024-07
Yi 1.5 9B Chat 16K01.AI70.82024-07
Meta Llama 3 8B InstructMeta70.62024-07
Meta Llama 3 8B InstructNous Research70.62024-07
Yi 9B01.AI70.42024-07
Hermes 2 Theta Llama 3 8BNous Research70.32024-07
Hermes 2 Pro Llama 3 8BNous Research69.82024-07
Yi 1.5 9B Chat01.AI69.62024-07
gemma 7B itGoogle DeepMind68.82024-07
Nous Hermes 2 SOLAR 10.7BNous Research68.52024-07
Mistral 7B v0.3Mistral AI66.22024-07
mistral 7B v0.3 bnb 4bitUnsloth66.22024-07
Yi 6B01.AI66.02024-07
Yi 6B Chat01.AI66.02024-07
Yi 1.5 6B01.AI65.52024-07
Mistral 7B Instruct v0.2Mistral AI63.12024-07
Yi 1.5 6B Chat01.AI62.32024-07
falcon 40BTII56.52024-07
phi 2Microsoft56.42024-07
Qwen2 1.5B InstructAlibaba / Qwen Team55.42024-07
deepseek llm 7B baseDeepSeek52.02024-07
deepseek llm 7B chatDeepSeek52.02024-07
chatglm2 6BZhipu AI44.42024-07
Qwen2 0.5B InstructAlibaba / Qwen Team41.32024-07
gemma 2B itGoogle DeepMind38.22024-07
gemma 2BGoogle DeepMind37.42024-07
deepseek coder 6.7B instructDeepSeek35.12024-07
deepseek coder 6.7B baseDeepSeek31.72024-07
OLMo 1B hfAllen AI27.52024-07
deepseek coder 1.3B baseDeepSeek24.22024-07
deepseek coder 1.3B instructDeepSeek24.22024-07

Data

This page as JSON · Edit on GitHub