MMLU: Logical Fallacies

Accuracy on MMLU's logical fallacies questions, one of 57 subject tests of academic and professional knowledge.

Also known as: logical_fallacies

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryknowledge
SubcategoryHumanities
Page statusactive
Metricaccuracy
Directionhigher_is_better
Unit%
Dataset size163
Dataset licenceMIT
PublisherUC Berkeley

What it measures

Identification of informal fallacies in short arguments -- ad hominem, straw man, false dilemma, and the like -- at the level of an introductory critical-thinking or informal-logic course. Framed as four-option multiple-choice questions and scored zero-shot or few-shot by exact match against the labelled option, as one of the 57 subject subsets that make up the MMLU benchmark.

Task format

Four-option multiple-choice question answering (A-D), one correct answer, zero-shot or few-shot.

Models reporting this benchmark

These figures come from the model cards, which carry one collection date per card and no per-score attribution. They are shown as reported, not as verified evidence.
ModelProviderScoreCard as of
Yi 34B 200K01.AI88.32024-07
Mixtral 8x22B Instruct v0.1Mistral AI87.12026-04
Nous Hermes 2 Yi 34BNous Research87.12024-07
Yi 1.5 34B01.AI87.12026-04
Yi 1.5 34B 32K01.AI86.52024-07
Meta Llama 3 70BMeta85.92024-07
Meta Llama 3 70B InstructMeta85.92026-04
Meta Llama 3 70B InstructNous Research85.92026-04
Yi 1.5 34B Chat01.AI85.92026-04
Yi 1.5 34B Chat 16K01.AI85.92024-07
Yi 34B Chat01.AI85.32024-07
Yi 1.5 9B Chat 16K01.AI82.22024-07
Mixtral 8x7B Instruct v0.1Mistral AI81.62026-04
Phi 3 mini 128K instructMicrosoft81.62024-07
Nous Hermes 2 Mixtral 8x7B DPONous Research80.42024-07
Phi 3 mini 4K instructMicrosoft80.42024-07
Yi 1.5 9B Chat01.AI80.42024-07
Yi 6B01.AI79.12024-07
Yi 6B Chat01.AI79.12024-07
Mistral 7B v0.3Mistral AI78.52024-07
mistral 7B v0.3 bnb 4bitUnsloth78.52024-07
Yi 1.5 9B01.AI78.52024-07
Mixtral 8x7B v0.1Mistral AI77.32026-04
Meta Llama 3 8B InstructMeta76.72024-07
Meta Llama 3 8B InstructNous Research76.72024-07
Yi 1.5 6B Chat01.AI76.72024-07
Yi 9B01.AI76.12024-07
gemma 7B itGoogle DeepMind74.82024-07
Nous Hermes 2 SOLAR 10.7BNous Research74.82024-07
Yi 1.5 9B 32K01.AI74.22024-07
Hermes 2 Theta Llama 3 8BNous Research73.62024-07
Meta Llama 3 8BMeta73.62024-07
Meta Llama 3 8BNous Research73.62024-07
Yi 1.5 6B01.AI73.62024-07
Mistral 7B Instruct v0.2Mistral AI73.02024-07
phi 2Microsoft73.02024-07
Hermes 2 Pro Llama 3 8BNous Research71.82024-07
Qwen2 1.5B InstructAlibaba / Qwen Team68.72024-07
falcon 40BTII65.62024-07
deepseek llm 7B baseDeepSeek60.72024-07
deepseek llm 7B chatDeepSeek60.72024-07
chatglm2 6BZhipu AI49.12024-07
Qwen2 0.5B InstructAlibaba / Qwen Team47.92024-07
deepseek coder 6.7B instructDeepSeek46.02024-07
deepseek coder 6.7B baseDeepSeek42.32024-07
gemma 2BGoogle DeepMind41.12024-07
gemma 2B itGoogle DeepMind36.22024-07
deepseek coder 1.3B baseDeepSeek28.82024-07
deepseek coder 1.3B instructDeepSeek28.82024-07
OLMo 1B hfAllen AI27.02024-07

Data

This page as JSON · Edit on GitHub