MMLU: Human Sexuality

Accuracy on MMLU's human sexuality questions, one of 57 subject tests of academic and professional knowledge.

Also known as: human_sexuality

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryknowledge
SubcategorySocial Sciences
Page statusactive
Metricaccuracy
Directionhigher_is_better
Unit%
Dataset size131
Dataset licenceMIT
PublisherUC Berkeley

What it measures

Human sexual biology, development, and behaviour, and the psychological and social aspects of sexuality, at the level of an introductory psychology or health-science course. Framed as four-option multiple-choice questions and scored zero-shot or few-shot by exact match against the labelled option, as one of the 57 subject subsets that make up the MMLU benchmark.

Task format

Four-option multiple-choice question answering (A-D), one correct answer, zero-shot or few-shot.

Models reporting this benchmark

These figures come from the model cards, which carry one collection date per card and no per-score attribution. They are shown as reported, not as verified evidence.
ModelProviderScoreCard as of
Nous Hermes 2 Yi 34BNous Research89.32024-07
Yi 34B Chat01.AI89.32024-07
Mixtral 8x22B Instruct v0.1Mistral AI88.52026-04
Yi 1.5 34B 32K01.AI88.52024-07
Nous Hermes 2 Mixtral 8x7B DPONous Research87.82024-07
Yi 1.5 34B01.AI87.82026-04
Yi 1.5 34B Chat01.AI87.82026-04
Yi 1.5 34B Chat 16K01.AI87.82024-07
Meta Llama 3 70BMeta87.02024-07
Meta Llama 3 70B InstructMeta87.02026-04
Meta Llama 3 70B InstructNous Research87.02026-04
Yi 34B 200K01.AI86.32024-07
Hermes 2 Theta Llama 3 8BNous Research80.92024-07
Mixtral 8x7B Instruct v0.1Mistral AI80.92026-04
Mixtral 8x7B v0.1Mistral AI80.92026-04
Yi 1.5 9B 32K01.AI80.22024-07
Yi 9B01.AI78.62024-07
Meta Llama 3 8B InstructMeta77.92024-07
Meta Llama 3 8B InstructNous Research77.92024-07
Nous Hermes 2 SOLAR 10.7BNous Research77.92024-07
Yi 1.5 9B01.AI77.92024-07
Meta Llama 3 8BMeta77.12024-07
Meta Llama 3 8BNous Research77.12024-07
Mistral 7B v0.3Mistral AI77.12024-07
mistral 7B v0.3 bnb 4bitUnsloth77.12024-07
Hermes 2 Pro Llama 3 8BNous Research76.32024-07
Phi 3 mini 4K instructMicrosoft76.32024-07
Phi 3 mini 128K instructMicrosoft75.62024-07
Yi 1.5 9B Chat01.AI74.82024-07
Yi 1.5 9B Chat 16K01.AI74.82024-07
Mistral 7B Instruct v0.2Mistral AI74.02024-07
Yi 6B01.AI74.02024-07
Yi 6B Chat01.AI74.02024-07
falcon 40BTII72.52024-07
gemma 7B itGoogle DeepMind72.52024-07
phi 2Microsoft70.22024-07
Yi 1.5 6B01.AI70.22024-07
Qwen2 1.5B InstructAlibaba / Qwen Team67.22024-07
Yi 1.5 6B Chat01.AI65.62024-07
deepseek llm 7B baseDeepSeek55.72024-07
deepseek llm 7B chatDeepSeek55.72024-07
chatglm2 6BZhipu AI48.12024-07
Qwen2 0.5B InstructAlibaba / Qwen Team48.12024-07
deepseek coder 6.7B baseDeepSeek46.62024-07
gemma 2BGoogle DeepMind45.82024-07
gemma 2B itGoogle DeepMind42.72024-07
deepseek coder 6.7B instructDeepSeek42.02024-07
deepseek coder 1.3B baseDeepSeek35.12024-07
deepseek coder 1.3B instructDeepSeek35.12024-07
OLMo 1B hfAllen AI21.42024-07

Data

This page as JSON · Edit on GitHub