MMLU: Formal Logic

MMLU subject subset: Propositional and predicate logic, validity and inference rules, as taught in a philosophy department.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryknowledge
Subcategoryphilosophy
Page statusactive
Metricaccuracy
Directionhigher_is_better
Unit%
Dataset size126
Dataset licenceMIT
PublisherUC Berkeley (original); Center for AI Safety (current host)

What it measures

Propositional and predicate logic, validity and inference rules, as taught in a philosophy department. Questions are four-option multiple-choice, drawn from the MMLU test set's "philosophy" subcategory within the benchmark's "humanities" top-level group, and are graded on the single correct labelled option.

Task format

Four-option multiple-choice questions, graded on the single correct labelled option; commonly evaluated 5-shot, consistent with the rest of MMLU.

Models reporting this benchmark

These figures come from the model cards, which carry one collection date per card and no per-score attribution. They are shown as reported, not as verified evidence.
ModelProviderScoreCard as of
Yi 1.5 34B01.AI67.52026-04
Yi 1.5 34B Chat01.AI66.72026-04
Yi 1.5 34B Chat 16K01.AI66.72024-07
Meta Llama 3 70BMeta62.72024-07
Meta Llama 3 70B InstructMeta62.72026-04
Meta Llama 3 70B InstructNous Research62.72026-04
Yi 1.5 34B 32K01.AI62.72024-07
Yi 1.5 9B Chat 16K01.AI61.12024-07
Mixtral 8x22B Instruct v0.1Mistral AI59.52026-04
Phi 3 mini 4K instructMicrosoft58.72024-07
Yi 1.5 6B Chat01.AI58.72024-07
Yi 9B01.AI58.72024-07
Nous Hermes 2 Yi 34BNous Research57.92024-07
Yi 1.5 9B Chat01.AI57.92024-07
Nous Hermes 2 Mixtral 8x7B DPONous Research57.12024-07
Phi 3 mini 128K instructMicrosoft57.12024-07
Yi 1.5 9B 32K01.AI57.12024-07
Mixtral 8x7B v0.1Mistral AI56.32026-04
Yi 34B Chat01.AI54.02024-07
Yi 1.5 9B01.AI53.22024-07
Mixtral 8x7B Instruct v0.1Mistral AI52.42026-04
Yi 34B 200K01.AI51.62024-07
gemma 7B itGoogle DeepMind50.02024-07
Yi 1.5 6B01.AI49.22024-07
Hermes 2 Theta Llama 3 8BNous Research48.42024-07
Meta Llama 3 8B InstructMeta48.42024-07
Meta Llama 3 8B InstructNous Research48.42024-07
Hermes 2 Pro Llama 3 8BNous Research47.62024-07
Meta Llama 3 8BMeta47.62024-07
Meta Llama 3 8BNous Research47.62024-07
Nous Hermes 2 SOLAR 10.7BNous Research45.22024-07
Yi 6B01.AI43.72024-07
Yi 6B Chat01.AI43.72024-07
Mistral 7B Instruct v0.2Mistral AI42.12024-07
Mistral 7B v0.3Mistral AI39.72024-07
mistral 7B v0.3 bnb 4bitUnsloth39.72024-07
Qwen2 1.5B InstructAlibaba / Qwen Team38.92024-07
deepseek coder 6.7B instructDeepSeek36.52024-07
phi 2Microsoft35.72024-07
chatglm2 6BZhipu AI34.92024-07
deepseek llm 7B baseDeepSeek33.32024-07
deepseek llm 7B chatDeepSeek33.32024-07
Qwen2 0.5B InstructAlibaba / Qwen Team33.32024-07
falcon 40BTII30.22024-07
deepseek coder 6.7B baseDeepSeek29.42024-07
gemma 2BGoogle DeepMind27.02024-07
gemma 2B itGoogle DeepMind25.42024-07
deepseek coder 1.3B baseDeepSeek21.42024-07
deepseek coder 1.3B instructDeepSeek21.42024-07
OLMo 1B hfAllen AI20.62024-07

Data

This page as JSON · Edit on GitHub