{
 "body": "## What it measures\n\nyes_no_black_white is a benchmark task documented by the bigbench evaluation integration. The task-specific input and expected output should be taken from the cited primary registry revision. It exercises the capability named by the benchmark rather than establishing a general model property.\n\n## How it is scored\n\nThe integration uses a task-specific correctness or generation metric. The inspected registry lead does not establish a complete aggregate baseline or universal normalization, so those fields remain unknown.\n\n## Dataset and licence\n\nThe cited primary source identifies the benchmark integration. Item count, split details, and licence are left unknown where the inspected source does not state them. Answers and public exposure should be checked against the exact release.\n\n## Who publishes it\n\nThe cited harness or benchmark repository maintains the integration. A separate current leaderboard and complete author list were not established from this bounded source check.\n\n## Lineage\n\nThis page documents the exact identifier `yes_no_black_white`. Similar names and alternate harness integrations should not be merged without checking identity and protocol. No predecessor or successor was established.\n\n## Saturation and contamination\n\nSaturation is unknown. Public task code does not establish training exposure or a contamination study, so results should retain the dataset and harness revision.\n\n## How to run it\n\nUse the exact `yes_no_black_white` task in the cited harness where available. Record revision, prompt, shot count, decoding, and evaluator because protocol differences can change scores.\n\n## Reading the numbers\n\nA strong score indicates success on this specific task format and data release. It does not establish broad reasoning, translation, language, or knowledge ability beyond that protocol. Compare only matching revisions and metrics. Preserve per-example outputs when possible so formatting failures can be separated from capability failures. Unknown fields remain unknown until a primary source establishes them.\n\nThe benchmark should be treated as a measurement of the published task, not a general capability certificate. Update the source record only after checking a new primary release or harness implementation.\n\nResults should preserve prompt, language, reference, and evaluator settings. A score from a similarly named benchmark or a different release is not interchangeable.\n The exact release remains essential for reproducible comparison.\n",
 "build": {
  "built_at": "2026-09-09T16:56:50+00:00",
  "commit": "0a599558854c0e238c03a0f0d725239cb28f9d11",
  "eligibility_as_of": "2026-09-09"
 },
 "disposition": {
  "canonical_id": "yes_no_black_white",
  "reasons": [],
  "status": "unassessed",
  "verified_results": []
 },
 "models_covered": [],
 "page": {
  "aliases": [],
  "category": "knowledge",
  "contamination": {
   "note": "",
   "risk": "unknown"
  },
  "dataset": {
   "languages": [],
   "license": "",
   "modalities": [
    "text"
   ],
   "public_test_set": null,
   "size": null,
   "size_note": "",
   "splits": "",
   "url": "https://github.com/google/BIG-bench/tree/main/bigbench/benchmark_tasks/yes_no_black_white"
  },
  "freshness": {
   "researched": "2026-09-09",
   "researched_by": "GPT-5.6 Luna, luna-stream-b-002 (Codex coordinated)",
   "reviewed": "",
   "reviewed_by": ""
  },
  "harness": {
   "bigbench": "yes_no_black_white",
   "lm_eval": "yes_no_black_white"
  },
  "id": "yes_no_black_white",
  "last_updated": "",
  "leaderboard_url": "",
  "lineage": {
   "family": "",
   "predecessor": "",
   "successors": [],
   "variants": []
  },
  "measures": "yes_no_black_white is a benchmark task documented by the bigbench evaluation integration.",
  "metric": {
   "baseline_note": "",
   "direction": "higher_is_better",
   "human_baseline": null,
   "max_score": 100,
   "name": "accuracy",
   "random_baseline": null,
   "unit": "percent"
  },
  "name": "yes_no_black_white",
  "page_kind": "benchmark",
  "paper": {
   "arxiv": "",
   "title": "",
   "url": "",
   "year": null
  },
  "publisher": {
   "authors": [],
   "org": "",
   "url": "https://github.com/google/BIG-bench/tree/main/bigbench/benchmark_tasks/yes_no_black_white"
  },
  "released": "",
  "repo_url": "https://github.com/google/BIG-bench/tree/main/bigbench/benchmark_tasks/yes_no_black_white",
  "saturation": {
   "as_of": "",
   "note": "",
   "status": "unknown",
   "top_score": null
  },
  "sources": [
   {
    "accessed": "2026-09-09",
    "title": "yes_no_black_white primary task source",
    "url": "https://github.com/google/BIG-bench/tree/main/bigbench/benchmark_tasks/yes_no_black_white"
   }
  ],
  "status": "active",
  "subcategory": "benchmark task",
  "summary": "yes_no_black_white is a benchmark task documented by its cited evaluation harness.",
  "tags": [
   "benchmark"
  ],
  "task_format": "Text input with task-specific prediction or generation output."
 }
}