{
 "body": "\n## What it measures\n\nWhat Is the Tao? compares stylistic elements of translations of a complex philosophical text. The task asks a model to select the best answer among alternatives.\n\n## How it is scored\n\nBIG-bench uses multiple-choice grading. The task file contains 36 examples and does not state a human baseline.\n\n## Dataset and licence\n\nThe 36 examples are embedded in the public BIG-bench task file. The translated passages are excerpted from V\u00edt Brunner's Tao Te Ching comparison website (ttc.tasuki.org). No separate split or licence is stated.\n\n## Who publishes it\n\nThe task was contributed to BIG-bench by Alice Xiang and Ekin Dogus Cubuk. No separate paper or leaderboard was established.\n\n## Lineage\n\nThis is a standalone BIG-bench task with no established predecessor or successor.\n\n## Saturation and contamination\n\nSaturation is unknown. The tiny public set makes memorization and uncertainty material. The task's own README reports that on one of its four question types (reading comprehension about the nature of the Tao), the largest models tested answered consistently correctly even when the translated passages were removed from the prompt, suggesting pre-training knowledge rather than in-context analysis drove that subset's score.\n\n## How to run it\n\nRun BIG-bench task `what_is_the_tao`; preserve task revision and answer choices.\n\n## Reading the numbers\n\nA high score indicates agreement with the task\u2019s translation-style distinctions. It does not establish broad translation quality, philosophical understanding, or cultural competence. Item-level review is essential because 36 examples are too few for precise ranking.\n\nInterpret the result as a narrow literary judgment probe, not a general language metric.\n\nTranslation style is inherently sensitive to source edition and evaluator framing. The task\u2019s small public example set cannot support stable model rankings, and no claim about philosophical expertise should be inferred from its accuracy. Preserve the exact task revision and choices when reproducing a result.\n\nThe benchmark can probe sensitivity to wording and literary register, but those are narrower constructs than translation quality in general. Human interpretation remains necessary for claims about meaning or philosophical fidelity.\n\nDifferent translations may be defensible for reasons the multiple-choice key does not capture. Error analysis should retain the full alternatives and explain which stylistic feature determined the reference answer. This matters when comparing models with different prompts or language backgrounds.\n\nThe task should therefore be cited with its exact examples and choice key.\n\nIts narrow scope is still useful for checking whether a model distinguishes register and stylistic framing in philosophical translation prompts. It should not be used as a standalone measure of translation competence.\n",
 "build": {
  "built_at": "2026-09-09T16:56:50+00:00",
  "commit": "0a599558854c0e238c03a0f0d725239cb28f9d11",
  "eligibility_as_of": "2026-09-09"
 },
 "disposition": {
  "canonical_id": "what_is_the_tao",
  "reasons": [],
  "status": "unassessed",
  "verified_results": []
 },
 "models_covered": [],
 "page": {
  "aliases": [],
  "category": "knowledge",
  "contamination": {
   "note": "Public examples may be in training data.",
   "risk": "medium"
  },
  "dataset": {
   "languages": [
    "English"
   ],
   "license": "",
   "modalities": [
    "text"
   ],
   "public_test_set": true,
   "size": 36,
   "size_note": "The task README and auto-generated header both state 36 multiple-choice queries; translated passages are excerpted from V\u00edt Brunner's Tao Te Ching comparison website (ttc.tasuki.org).",
   "splits": "",
   "url": "https://github.com/google/BIG-bench/tree/main/bigbench/benchmark_tasks/what_is_the_tao"
  },
  "freshness": {
   "luna-stream-a-001 (Codex coordinated)": null,
   "researched": "2026-09-08",
   "researched_by": "GPT-5.6 Luna",
   "reviewed": "2026-09-08",
   "reviewed_by": "Claude Sonnet 5 independent review, luna-stream-a-001"
  },
  "harness": {
   "bigbench": "what_is_the_tao",
   "helm": "",
   "inspect_evals": "",
   "lm_eval": "",
   "opencompass": "",
   "other": ""
  },
  "id": "what_is_the_tao",
  "last_updated": "",
  "leaderboard_url": "",
  "lineage": {
   "family": "bigbench",
   "predecessor": "",
   "successors": [],
   "variants": []
  },
  "measures": "The task asks a model to compare translations of a complex philosophical text and select the stylistically appropriate option. It is an English multiple-choice interpretation task.",
  "metric": {
   "baseline_note": "",
   "direction": "higher_is_better",
   "human_baseline": null,
   "max_score": 100,
   "name": "multiple_choice_grade",
   "random_baseline": null,
   "unit": "percent"
  },
  "name": "What Is the Tao?",
  "page_kind": "benchmark",
  "paper": {
   "arxiv": "2206.04615",
   "title": "Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models",
   "url": "https://arxiv.org/abs/2206.04615",
   "year": 2022
  },
  "publisher": {
   "authors": [
    "Alice Xiang",
    "Ekin Dogus Cubuk"
   ],
   "org": "Google BIG-bench",
   "url": "https://github.com/google/BIG-bench/tree/main/bigbench/benchmark_tasks/what_is_the_tao"
  },
  "released": "2022",
  "repo_url": "https://github.com/google/BIG-bench",
  "saturation": {
   "as_of": "",
   "note": "No leaderboard was established.",
   "status": "unknown",
   "top_score": null
  },
  "sources": [
   {
    "accessed": "2026-09-08",
    "title": "BIG-bench task definition",
    "url": "https://raw.githubusercontent.com/google/BIG-bench/main/bigbench/benchmark_tasks/what_is_the_tao/task.json"
   },
   {
    "accessed": "2026-09-08",
    "title": "what_is_the_tao README (36 questions, authors, data source, limitations)",
    "url": "https://raw.githubusercontent.com/google/BIG-bench/main/bigbench/benchmark_tasks/what_is_the_tao/README.md"
   },
   {
    "accessed": "2026-09-08",
    "title": "BIG-bench paper",
    "url": "https://arxiv.org/abs/2206.04615"
   }
  ],
  "status": "active",
  "subcategory": "philosophy and translation",
  "summary": "BIG-bench What Is the Tao compares stylistic elements of translations of a philosophical text.",
  "tags": [
   "philosophy",
   "translation",
   "multiple-choice"
  ],
  "task_format": "Question comparing translation styles with answer choices."
 }
}