{
 "body": "## What it measures\n\nNYU-LLM-CTF/NYU_CTF_Bench is a benchmark task documented by the  evaluation integration. The task-specific input and expected output should be taken from the cited primary registry revision. It exercises the capability named by the benchmark rather than establishing a general model property.\n\n## How it is scored\n\nThe integration uses a task-specific correctness or generation metric. The inspected registry lead does not establish a complete aggregate baseline or universal normalization, so those fields remain unknown.\n\n## Dataset and licence\n\nThe cited primary source identifies the benchmark integration. Item count, split details, and licence are left unknown where the inspected source does not state them. Answers and public exposure should be checked against the exact release.\n\n## Who publishes it\n\nThe cited harness or benchmark repository maintains the integration. A separate current leaderboard and complete author list were not established from this bounded source check.\n\n## Lineage\n\nThis page documents the exact identifier `nyu_llm_ctf_nyu_ctf_bench`. Similar names and alternate harness integrations should not be merged without checking identity and protocol. No predecessor or successor was established.\n\n## Saturation and contamination\n\nSaturation is unknown. Public task code does not establish training exposure or a contamination study, so results should retain the dataset and harness revision.\n\n## How to run it\n\nUse the exact `nyu_llm_ctf_nyu_ctf_bench` task in the cited harness where available. Record revision, prompt, shot count, decoding, and evaluator because protocol differences can change scores.\n\n## Reading the numbers\n\nA strong score indicates success on this specific task format and data release. It does not establish broad reasoning, translation, language, or knowledge ability beyond that protocol. Compare only matching revisions and metrics. Preserve per-example outputs when possible so formatting failures can be separated from capability failures. Unknown fields remain unknown until a primary source establishes them.\n\nThe benchmark should be treated as a measurement of the published task, not a general capability certificate. Update the source record only after checking a new primary release or harness implementation.\n\nThe exact release remains essential for reproducible comparison and independent review. Additional protocol details should be retained with every reported score.\n\nThe source record should be updated only after checking a new primary release.\n",
 "build": {
  "built_at": "2026-09-09T16:56:50+00:00",
  "commit": "0a599558854c0e238c03a0f0d725239cb28f9d11",
  "eligibility_as_of": "2026-09-09"
 },
 "disposition": {
  "canonical_id": "nyu_llm_ctf_nyu_ctf_bench",
  "reasons": [],
  "status": "unassessed",
  "verified_results": []
 },
 "models_covered": [],
 "page": {
  "aliases": [],
  "category": "knowledge",
  "contamination": {
   "note": "",
   "risk": "unknown"
  },
  "dataset": {
   "languages": [],
   "license": "",
   "modalities": [
    "text"
   ],
   "public_test_set": null,
   "size": null,
   "size_note": "",
   "splits": "",
   "url": "https://github.com/NYU-LLM-CTF/NYU_CTF_Bench"
  },
  "freshness": {
   "researched": "2026-09-09",
   "researched_by": "GPT-5.6 Luna, luna-stream-b-011 (Codex coordinated)",
   "reviewed": "",
   "reviewed_by": ""
  },
  "harness": {
   "bigbench": "",
   "helm": "",
   "inspect_evals": "",
   "lm_eval": "",
   "opencompass": "",
   "other": ""
  },
  "id": "nyu_llm_ctf_nyu_ctf_bench",
  "last_updated": "",
  "leaderboard_url": "",
  "lineage": {
   "family": "",
   "predecessor": "",
   "successors": [],
   "variants": []
  },
  "measures": "NYU-LLM-CTF/NYU_CTF_Bench is a benchmark task documented by the  evaluation integration.",
  "metric": {
   "baseline_note": "",
   "direction": "higher_is_better",
   "human_baseline": null,
   "max_score": 100,
   "name": "accuracy",
   "random_baseline": null,
   "unit": "percent"
  },
  "name": "NYU-LLM-CTF/NYU_CTF_Bench",
  "page_kind": "benchmark",
  "paper": {
   "arxiv": "",
   "title": "",
   "url": "",
   "year": null
  },
  "publisher": {
   "authors": [],
   "org": "",
   "url": "https://github.com/NYU-LLM-CTF/NYU_CTF_Bench"
  },
  "released": "",
  "repo_url": "https://github.com/NYU-LLM-CTF/NYU_CTF_Bench",
  "saturation": {
   "as_of": "",
   "note": "",
   "status": "unknown",
   "top_score": null
  },
  "sources": [
   {
    "accessed": "2026-09-09",
    "title": "NYU-LLM-CTF/NYU_CTF_Bench primary task source",
    "url": "https://github.com/NYU-LLM-CTF/NYU_CTF_Bench"
   }
  ],
  "status": "active",
  "subcategory": "benchmark task",
  "summary": "NYU-LLM-CTF/NYU_CTF_Bench is a benchmark task documented by its cited evaluation harness.",
  "tags": [
   "benchmark"
  ],
  "task_format": "Text input with task-specific prediction or generation output."
 }
}