MMLU: Miscellaneous

Accuracy on MMLU's miscellaneous questions, one of 57 subject tests of academic and professional knowledge.

Also known as: miscellaneous

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryknowledge
SubcategoryOther
Page statusactive
Metricaccuracy
Directionhigher_is_better
Unit%
Dataset size783
Dataset licenceMIT
PublisherUC Berkeley

What it measures

General-knowledge questions that did not fit any of MMLU's other 56 named subjects -- broad factual trivia spanning many domains -- which is also why this is one of the largest MMLU subject tests. Framed as four-option multiple-choice questions and scored zero-shot or few-shot by exact match against the labelled option, as one of the 57 subject subsets that make up the MMLU benchmark.

Task format

Four-option multiple-choice question answering (A-D), one correct answer, zero-shot or few-shot.

Models reporting this benchmark

These figures come from the model cards, which carry one collection date per card and no per-score attribution. They are shown as reported, not as verified evidence.
ModelProviderScoreCard as of
Meta Llama 3 70BMeta91.82024-07
Meta Llama 3 70B InstructMeta91.82026-04
Meta Llama 3 70B InstructNous Research91.82026-04
Nous Hermes 2 Yi 34BNous Research91.12024-07
Yi 1.5 34B 32K01.AI90.52024-07
Yi 34B 200K01.AI90.32024-07
Yi 1.5 34B01.AI90.22026-04
Mixtral 8x22B Instruct v0.1Mistral AI89.92026-04
Yi 1.5 34B Chat01.AI89.92026-04
Yi 1.5 34B Chat 16K01.AI89.92024-07
Yi 34B Chat01.AI89.92024-07
Mixtral 8x7B Instruct v0.1Mistral AI88.02026-04
Nous Hermes 2 Mixtral 8x7B DPONous Research88.02024-07
Mixtral 8x7B v0.1Mistral AI87.52026-04
Hermes 2 Theta Llama 3 8BNous Research84.42024-07
Yi 1.5 9B 32K01.AI84.42024-07
Yi 1.5 9B Chat 16K01.AI84.22024-07
Yi 9B01.AI83.92024-07
gemma 7B itGoogle DeepMind83.82024-07
Yi 1.5 9B Chat01.AI83.42024-07
Hermes 2 Pro Llama 3 8BNous Research83.12024-07
Meta Llama 3 8BMeta83.02024-07
Meta Llama 3 8BNous Research83.02024-07
Yi 1.5 9B01.AI83.02024-07
Nous Hermes 2 SOLAR 10.7BNous Research82.82024-07
Phi 3 mini 4K instructMicrosoft82.82024-07
Phi 3 mini 128K instructMicrosoft81.62024-07
Yi 6B01.AI80.72024-07
Yi 6B Chat01.AI80.72024-07
Yi 1.5 6B01.AI80.62024-07
Meta Llama 3 8B InstructMeta79.82024-07
Meta Llama 3 8B InstructNous Research79.82024-07
Mistral 7B v0.3Mistral AI79.82024-07
mistral 7B v0.3 bnb 4bitUnsloth79.82024-07
Yi 1.5 6B Chat01.AI79.72024-07
Mistral 7B Instruct v0.2Mistral AI78.02024-07
falcon 40BTII75.92024-07
deepseek llm 7B baseDeepSeek73.12024-07
deepseek llm 7B chatDeepSeek73.12024-07
Qwen2 1.5B InstructAlibaba / Qwen Team72.02024-07
phi 2Microsoft69.22024-07
chatglm2 6BZhipu AI59.62024-07
gemma 2BGoogle DeepMind54.72024-07
Qwen2 0.5B InstructAlibaba / Qwen Team52.12024-07
gemma 2B itGoogle DeepMind46.22024-07
deepseek coder 6.7B instructDeepSeek43.72024-07
deepseek coder 6.7B baseDeepSeek40.22024-07
deepseek coder 1.3B baseDeepSeek29.22024-07
deepseek coder 1.3B instructDeepSeek29.22024-07
OLMo 1B hfAllen AI29.02024-07

Data

This page as JSON · Edit on GitHub