{
 "body": "\n## What it measures\n\nTMMLU+ is a family of Traditional Chinese knowledge evaluations, focused on content reflecting Taiwan's linguistic, educational, and professional context. It uses subject-specific multiple-choice questions spanning 66 subjects, from elementary school topics to professional licensing exams (law, medicine, accounting), to test factual and applied knowledge.\n\nThe family identity is more precise than treating `tmmluplus` as one score: subject composition, difficulty, and item count vary widely (from about 90 to over 1,600 rows per subject in the released dataset), so aggregation choices materially affect a reported number.\n\n## How it is scored\n\nThe harness family uses accuracy-style multiple-choice grading, following the MMLU-style format the paper describes as an improvement on the original TMMLU. Subject-level random baselines depend on the number of answer choices per item. The paper reports evaluated open-weight Chinese models (1.8B-72B parameters) scoring below both Simplified Chinese counterparts and human performance on average, but does not establish one fixed human baseline figure that this page can cite.\n\n## Dataset and licence\n\nThe dataset is hosted on Hugging Face as `ikala/tmmluplus` under an MIT licence, with 22,160 rows across 66 subject configs, each split into train (typically 5 examples), validation, and test partitions. It is described as roughly six times larger than its predecessor TMMLU, with a more balanced subject distribution and a dedicated development set that TMMLU lacked.\n\n## Who publishes it\n\niKala published TMMLU+ in a 2024 paper, \"An Improved Traditional Chinese Evaluation Suite for Foundation Model,\" by Zhi-Rui Tam, Ya-Ting Pai, Yen-Wei Lee, Jun-Da Chen, Wei-Min Chu, Sega Cheng, and Hong-Han Shuai. iKala maintains the dataset on Hugging Face; EleutherAI's lm-evaluation-harness maintains the `tmmluplus` task group used to run it.\n\n## Lineage\n\nTMMLU+ succeeds TMMLU (Taiwan Massive Multitask Language Understanding), which does not yet have its own page in this repository; the paper describes TMMLU+ as roughly six times larger with more balanced subject coverage. TMMLU+ is itself a family with 66 subject-level variants that should be reported separately unless a documented aggregation rule is given.\n\n## Saturation and contamination\n\nAggregate saturation is unknown from primary sources available here, though the paper's own results show a persistent gap between Traditional Chinese and Simplified Chinese model performance as of 2024. Contamination risk is unknown at the subject level: the pooled dataset has been publicly hosted since December 2023, but individual subjects' original publication histories were not established in this review.\n\n## How to run it\n\nRun lm-evaluation-harness task group `tmmluplus` after checking the current registry, or load `ikala/tmmluplus` directly from Hugging Face. Record selected subjects, dataset revision (the card notes a v1.1 quality update), and aggregation rule, since a pooled score across 66 unevenly sized subjects is not comparable to a single-subject score.\n\n## Reading the numbers\n\nA high score indicates knowledge on the selected Traditional Chinese subject tasks, with content weighted toward Taiwan-specific context (geography, professional licensing, Taiwanese Hokkien, and similar subjects). It does not establish broad Chinese-language proficiency or uniform competence across all 66 subjects; the paper itself shows models can trail human performance and Simplified Chinese benchmarks on average. Subject-level results are necessary for meaningful comparison, since subject sizes range from under 100 to over 1,600 items.\n\nAvoid comparing a pooled TMMLU+ score with one subject unless the weighting and dataset revision match, and note the dataset card documents at least one quality-focused revision (v1.1) since initial release.\n",
 "build": {
  "built_at": "2026-09-09T16:56:50+00:00",
  "commit": "0a599558854c0e238c03a0f0d725239cb28f9d11",
  "eligibility_as_of": "2026-09-09"
 },
 "disposition": {
  "canonical_id": "tmmluplus",
  "reasons": [],
  "status": "unassessed",
  "verified_results": []
 },
 "models_covered": [],
 "page": {
  "aliases": [
   "TMMLU Plus"
  ],
  "category": "knowledge",
  "contamination": {
   "note": "Subject datasets have different publication histories and exposure profiles; the pooled Hugging Face release itself has been public since December 2023.",
   "risk": "unknown"
  },
  "dataset": {
   "languages": [
    "Chinese"
   ],
   "license": "MIT",
   "modalities": [
    "text"
   ],
   "public_test_set": true,
   "size": 22160,
   "size_note": "22,160 rows across 66 subject configs (train+validation+test), per the Hugging Face datasets-server size endpoint for ikala/tmmluplus; per-subject counts range from about 90 to over 1,600 rows.",
   "splits": "train, validation, test (per subject)",
   "url": "https://huggingface.co/datasets/ikala/tmmluplus"
  },
  "freshness": {
   "luna-new-001 (Codex coordinated)": null,
   "researched": "2026-09-08",
   "researched_by": "GPT-5.6 Luna",
   "reviewed": "2026-09-08",
   "reviewed_by": "Claude Sonnet 5 independent review, luna-new-001"
  },
  "harness": {
   "bigbench": "",
   "helm": "",
   "inspect_evals": "",
   "lm_eval": "tmmluplus",
   "opencompass": "",
   "other": ""
  },
  "id": "tmmluplus",
  "last_updated": "",
  "leaderboard_url": "",
  "lineage": {
   "family": "",
   "predecessor": "tmmlu",
   "successors": [],
   "variants": []
  },
  "measures": "TMMLU+ evaluates subject knowledge through multiple-choice questions in Taiwanese Mandarin. It is a harness family with subject-level tasks rather than one homogeneous item set.",
  "metric": {
   "baseline_note": "Subject task baselines vary.",
   "direction": "higher_is_better",
   "human_baseline": null,
   "max_score": 100,
   "name": "accuracy",
   "random_baseline": null,
   "unit": "percent"
  },
  "name": "TMMLU+",
  "page_kind": "family",
  "paper": {
   "arxiv": "2403.01858",
   "title": "An Improved Traditional Chinese Evaluation Suite for Foundation Model",
   "url": "https://arxiv.org/abs/2403.01858",
   "year": 2024
  },
  "publisher": {
   "authors": [
    "Zhi-Rui Tam",
    "Ya-Ting Pai",
    "Yen-Wei Lee",
    "Jun-Da Chen",
    "Wei-Min Chu",
    "Sega Cheng",
    "Hong-Han Shuai"
   ],
   "org": "iKala",
   "url": "https://huggingface.co/datasets/ikala/tmmluplus"
  },
  "released": "2024-03",
  "repo_url": "https://github.com/EleutherAI/lm-evaluation-harness",
  "saturation": {
   "as_of": "",
   "note": "The paper reports closed and open-weight (1.8B-72B) Chinese LLMs still trailing Simplified Chinese counterparts and human performance on average, but no current aggregate leaderboard was established here.",
   "status": "unknown",
   "top_score": null
  },
  "sources": [
   {
    "accessed": "2026-09-08",
    "title": "An Improved Traditional Chinese Evaluation Suite for Foundation Model (arXiv abstract)",
    "url": "https://arxiv.org/abs/2403.01858"
   },
   {
    "accessed": "2026-09-08",
    "title": "Hugging Face dataset card metadata for ikala/tmmluplus",
    "url": "https://huggingface.co/api/datasets/ikala/tmmluplus"
   },
   {
    "accessed": "2026-09-08",
    "title": "Hugging Face datasets-server size endpoint for ikala/tmmluplus",
    "url": "https://datasets-server.huggingface.co/size?dataset=ikala/tmmluplus"
   },
   {
    "accessed": "2026-09-08",
    "title": "lm-evaluation-harness TMMLU+ task family",
    "url": "https://github.com/EleutherAI/lm-evaluation-harness/tree/main/lm_eval/tasks/tmmluplus"
   },
   {
    "accessed": "2026-09-08",
    "title": "lm-evaluation-harness repository",
    "url": "https://github.com/EleutherAI/lm-evaluation-harness"
   }
  ],
  "status": "active",
  "subcategory": "Taiwanese Mandarin multiple-choice knowledge",
  "summary": "TMMLU+ is an lm-evaluation-harness task family for Taiwanese Mandarin knowledge questions.",
  "tags": [
   "knowledge",
   "chinese",
   "multiple-choice",
   "taiwan"
  ],
  "task_format": "Subject-specific multiple-choice question with answer choices."
 }
}