MMLU: Econometrics

MMLU subject subset: Statistical methods applied to economic data: regression, estimation and hypothesis testing in an economics context.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryknowledge
Subcategoryeconomics
Page statusactive
Metricaccuracy
Directionhigher_is_better
Unit%
Dataset size114
Dataset licenceMIT
PublisherUC Berkeley (original); Center for AI Safety (current host)

What it measures

Statistical methods applied to economic data: regression, estimation and hypothesis testing in an economics context. Questions are four-option multiple-choice, drawn from the MMLU test set's "economics" subcategory within the benchmark's "social sciences" top-level group, and are graded on the single correct labelled option.

Task format

Four-option multiple-choice questions, graded on the single correct labelled option; commonly evaluated 5-shot, consistent with the rest of MMLU.

Models reporting this benchmark

These figures come from the model cards, which carry one collection date per card and no per-score attribution. They are shown as reported, not as verified evidence.
ModelProviderScoreCard as of
Meta Llama 3 70BMeta72.82024-07
Meta Llama 3 70B InstructMeta72.82026-04
Meta Llama 3 70B InstructNous Research72.82026-04
Mixtral 8x7B v0.1Mistral AI64.92026-04
Nous Hermes 2 Mixtral 8x7B DPONous Research64.92024-07
Yi 1.5 34B01.AI64.92026-04
Mixtral 8x22B Instruct v0.1Mistral AI63.22026-04
Yi 1.5 34B 32K01.AI63.22024-07
Yi 1.5 9B01.AI63.22024-07
Yi 1.5 9B 32K01.AI63.22024-07
Mixtral 8x7B Instruct v0.1Mistral AI61.42026-04
Meta Llama 3 8B InstructMeta60.52024-07
Meta Llama 3 8B InstructNous Research60.52024-07
Yi 1.5 9B Chat 16K01.AI59.62024-07
Yi 1.5 9B Chat01.AI58.82024-07
Yi 1.5 34B Chat01.AI57.92026-04
Yi 1.5 34B Chat 16K01.AI57.92024-07
Nous Hermes 2 Yi 34BNous Research57.02024-07
Yi 34B 200K01.AI56.12024-07
Hermes 2 Theta Llama 3 8BNous Research55.32024-07
Yi 34B Chat01.AI55.32024-07
Nous Hermes 2 SOLAR 10.7BNous Research54.42024-07
Yi 9B01.AI52.62024-07
Yi 1.5 6B Chat01.AI51.82024-07
Hermes 2 Pro Llama 3 8BNous Research50.92024-07
Meta Llama 3 8BMeta49.12024-07
Meta Llama 3 8BNous Research49.12024-07
Phi 3 mini 128K instructMicrosoft49.12024-07
gemma 7B itGoogle DeepMind48.22024-07
Yi 1.5 6B01.AI46.52024-07
Phi 3 mini 4K instructMicrosoft45.62024-07
Mistral 7B v0.3Mistral AI44.72024-07
mistral 7B v0.3 bnb 4bitUnsloth44.72024-07
Mistral 7B Instruct v0.2Mistral AI40.42024-07
phi 2Microsoft38.62024-07
Qwen2 1.5B InstructAlibaba / Qwen Team38.62024-07
Yi 6B01.AI35.12024-07
Yi 6B Chat01.AI35.12024-07
falcon 40BTII33.32024-07
Qwen2 0.5B InstructAlibaba / Qwen Team32.52024-07
gemma 2BGoogle DeepMind31.62024-07
chatglm2 6BZhipu AI30.72024-07
deepseek coder 6.7B instructDeepSeek30.72024-07
deepseek coder 6.7B baseDeepSeek29.82024-07
gemma 2B itGoogle DeepMind28.12024-07
deepseek llm 7B baseDeepSeek27.22024-07
deepseek llm 7B chatDeepSeek27.22024-07
deepseek coder 1.3B baseDeepSeek24.62024-07
deepseek coder 1.3B instructDeepSeek24.62024-07
OLMo 1B hfAllen AI23.72024-07

Data

This page as JSON · Edit on GitHub