{
 "body": "\n## What it measures\n\nTurkishMMLU tests Turkish-language question answering across nine high-school curriculum subjects: Biology, Chemistry, Physics, Geography, Mathematics, Turkish Language and Literature, Philosophy, History, and Religion and Ethics. Questions were written by curriculum experts specifically for Turkish high-school standards, rather than translated from an English MMLU-style source, following the general design pattern of localized MMLU variants for other languages.\n\n## How it is scored\n\nThe paper evaluates over 20 models under zero-shot and few-shot settings, plus chain-of-thought prompting, and includes a question-difficulty analysis; the harness task itself uses multiple-choice accuracy. The paper's own evaluation covers both multilingual open models (for example Gemma, Llama, mT5) and proprietary models (GPT-4o, Claude, Gemini) alongside Turkish-specialized models, confirmed directly from the abstract. No human baseline is reported in the paper. Subject aggregation should be reported explicitly, since accuracy typically varies by subject.\n\n## Dataset and licence\n\nThe harness README says over 10,000 questions in nine subjects. It says dataset access requires contacting the authors; licence and exact split counts were not established.\n\n## Who publishes it\n\nThe benchmark was introduced by Y\u00fcksel, K\u00f6ksal, \u015eenel, Korhonen, and Sch\u00fctze in 2024. EleutherAI maintains the harness integration.\n\n## Lineage\n\nTurkishMMLU is a Turkish analogue of multitask multiple-choice evaluations in the style of MMLU, but built from curriculum-expert-authored questions rather than translation. No predecessor or successor benchmark specific to Turkish was established from the sources opened.\n\n## Saturation and contamination\n\nUnknown. The paper reports model scores but no current ceiling analysis.\n\n## How to run it\n\nUse lm-evaluation-harness group task `turkishmmlu`, which aggregates nine per-subject tasks named `turkishmmlu_{subject}`, or a chain-of-thought variant named `turkishmmlu_cot_{subject}`, confirmed from the harness README. Record which variant was run (plain multiple-choice versus CoT), the zero- or few-shot setting, and the dataset access revision, since the authors distribute the question set on request rather than through an open, versioned repository.\n\n## Reading the numbers\n\nA strong score indicates Turkish curriculum knowledge and multiple-choice reasoning. It does not establish general Turkish fluency or current factual accuracy. Compare subject mix and shot/CoT protocol.\n",
 "build": {
  "built_at": "2026-09-09T16:56:50+00:00",
  "commit": "0a599558854c0e238c03a0f0d725239cb28f9d11",
  "eligibility_as_of": "2026-09-09"
 },
 "disposition": {
  "canonical_id": "turkishmmlu",
  "reasons": [],
  "status": "unassessed",
  "verified_results": []
 },
 "models_covered": [],
 "page": {
  "aliases": [
   "Turkish MMLU"
  ],
  "category": "knowledge",
  "contamination": {
   "note": "The dataset is public through code/configuration but the README says access requires contacting the authors; exposure is therefore uncertain.",
   "risk": "unknown"
  },
  "dataset": {
   "languages": [
    "Turkish"
   ],
   "license": "",
   "modalities": [
    "text"
   ],
   "public_test_set": false,
   "size": 10000,
   "size_note": "The harness README says over 10,000 questions across nine subjects.",
   "splits": "subject-specific",
   "url": "https://github.com/EleutherAI/lm-evaluation-harness/tree/main/lm_eval/tasks/turkishmmlu"
  },
  "freshness": {
   "luna-new-002": null,
   "luna-new-002 (Codex coordinated)": null,
   "researched": "2026-09-08",
   "researched_by": "GPT-5.6 Luna",
   "reviewed": "2026-09-08",
   "reviewed_by": "Claude Sonnet 5 independent review"
  },
  "harness": {
   "bigbench": "",
   "helm": "",
   "inspect_evals": "",
   "lm_eval": "turkishmmlu",
   "opencompass": "",
   "other": ""
  },
  "id": "turkishmmlu",
  "last_updated": "",
  "leaderboard_url": "",
  "lineage": {
   "family": "",
   "predecessor": "",
   "successors": [],
   "variants": []
  },
  "measures": "Questions are written by curriculum experts and cover science, mathematics, language, and social sciences and humanities.",
  "metric": {
   "baseline_note": "",
   "direction": "higher_is_better",
   "human_baseline": null,
   "max_score": 100,
   "name": "accuracy",
   "random_baseline": null,
   "unit": "percent"
  },
  "name": "TurkishMMLU",
  "page_kind": "benchmark",
  "paper": {
   "arxiv": "2407.12402",
   "title": "TurkishMMLU: Measuring Massive Multitask Language Understanding in Turkish",
   "url": "https://arxiv.org/abs/2407.12402",
   "year": 2024
  },
  "publisher": {
   "authors": [
    "Arda Y\u00fcksel",
    "Abdullatif K\u00f6ksal",
    "L\u00fctfi Kerem \u015eenel",
    "Anna Korhonen",
    "Hinrich Sch\u00fctze"
   ],
   "org": "TurkishMMLU authors",
   "url": "https://arxiv.org/abs/2407.12402"
  },
  "released": "2024",
  "repo_url": "https://github.com/EleutherAI/lm-evaluation-harness",
  "saturation": {
   "as_of": "",
   "but no current leaderboard was established.": null,
   "note": "The cited README reports paper model scores",
   "status": "unknown",
   "top_score": null
  },
  "sources": [
   {
    "accessed": "2026-09-08",
    "title": "TurkishMMLU harness README",
    "url": "https://raw.githubusercontent.com/EleutherAI/lm-evaluation-harness/main/lm_eval/tasks/turkishmmlu/README.md"
   },
   {
    "accessed": "2026-09-08",
    "title": "TurkishMMLU paper",
    "url": "https://arxiv.org/abs/2407.12402"
   }
  ],
  "status": "active",
  "subcategory": "Turkish multiple-choice knowledge",
  "summary": "TurkishMMLU evaluates Turkish-language multiple-choice knowledge across nine high-school subjects.",
  "tags": [
   "turkish",
   "multiple-choice",
   "multilingual"
  ],
  "task_format": "Multiple-choice question answering in Turkish."
 }
}