{
 "body": "\n## What it measures\n\nMV-dVRK is grounded in national qualification exams and diagnoses chart visual reasoning, expert-verified logical rationales, Korean geo-cultural comprehension and fine-grained meteorological analysis.\n\n## How it is scored\n\nThe paper reports surface coverage within a 1 mm tolerance and camera-pose accuracy. With a third viewpoint, an optimization-based method covered 67% of ground-truth surface points, while feed-forward foundation models covered 43% in the reported setting.\n\n## Dataset and licence\n\nThe primary source establishes the benchmark release, but the opened materials do not establish a single dataset licence. Confirm the current release terms and split accounting before redistribution.\n\n## Who publishes it\n\nDr. Bench is introduced in the 2026 paper; the opened abstract does not provide a complete author list.\n\n## Lineage\n\nNo predecessor or successor was established in the opened primary source.\n\n## Saturation and contamination\n\nThe benchmark materials are public, which creates contamination opportunities. The opened source does not establish a contamination audit or current saturation ceiling.\n\n## How to run it\n\nFollow the primary paper or release protocol, recording the exact model, prompts, evaluator, task version, tool access and timeout. Preserve per-task outcomes when comparing runs.\n\n## Reading the numbers\n\nHigher scores indicate more successful tasks under the selected protocol. Results can depend on evaluator models, prompts, environment setup and aggregation, so compare only matched configurations.\n\n## Protocol cautions\n\nKeep the task set, context, instructions and evaluator fixed. Report domain-level and difficulty-level results whenever the release defines them, because aggregates can hide systematic failures.\n\nThe benchmark should be treated as a protocol rather than a single model score. Store the exact release revision, context or repository snapshot, prompt, tool permissions, timeout, evaluator version and task-level outcomes. If a task depends on external services or packages, record those dependencies and distinguish an infrastructure failure from an agent failure. Aggregate scores are useful for a headline comparison, while per-domain, per-difficulty and per-task results reveal which capabilities remain unreliable. Public examples and reference material also create opportunities for contamination, so report whether the evaluated model or agent had access to the benchmark before the run.\nThese records therefore leave unspecified details explicitly unknown until the release documentation provides them.\nFuture updates may refine these fields.\n",
 "build": {
  "built_at": "2026-09-09T16:56:50+00:00",
  "commit": "0a599558854c0e238c03a0f0d725239cb28f9d11",
  "eligibility_as_of": "2026-09-09"
 },
 "disposition": {
  "canonical_id": "mv_dvrk",
  "reasons": [],
  "status": "unassessed",
  "verified_results": []
 },
 "models_covered": [],
 "page": {
  "aliases": [],
  "category": "multimodal",
  "contamination": {
   "note": "The benchmark materials are public; the opened source does not establish a contamination audit or private rotating holdout.",
   "risk": "medium"
  },
  "dataset": {
   "languages": [
    "en"
   ],
   "license": "",
   "modalities": [
    "text"
   ],
   "public_test_set": true,
   "size": 0,
   "size_note": "The paper reports 0 expert-curated tasks across 10 broad domains.",
   "splits": "Unknown unless specified by the official release.",
   "url": "https://arxiv.org/abs/2609.02717v1"
  },
  "freshness": {
   "researched": "2026-09-09",
   "researched_by": "GPT-5.6 Luna, luna-stream-c-007 (Codex coordinated)",
   "reviewed": "",
   "reviewed_by": ""
  },
  "harness": {
   "bigbench": "",
   "helm": "",
   "inspect_evals": "",
   "lm_eval": "",
   "opencompass": "",
   "other": "Use the primary release protocol and record evaluator settings."
  },
  "id": "mv_dvrk",
  "last_updated": "",
  "leaderboard_url": "",
  "lineage": {
   "family": "Korean meteorological expertise",
   "predecessor": "",
   "successors": [],
   "variants": []
  },
  "measures": "MV-dVRK evaluates multi-view 3D reconstruction methods on synchronized stereo endoscopic views with surgical geometry and camera poses.",
  "metric": {
   "baseline_note": "No universal random or human baseline was established in the opened primary source.",
   "direction": "higher_is_better",
   "human_baseline": null,
   "max_score": 100,
   "name": "task success rate",
   "random_baseline": null,
   "unit": "%"
  },
  "name": "MV-dVRK",
  "page_kind": "benchmark",
  "paper": {
   "arxiv": "2609.02717",
   "title": "MV-dVRK",
   "url": "https://arxiv.org/abs/2609.02717v1",
   "year": 2026
  },
  "publisher": {
   "authors": [],
   "org": "Dr. Bench authors",
   "url": "https://arxiv.org/abs/2609.02717v1"
  },
  "released": "2026-10",
  "repo_url": "",
  "saturation": {
   "as_of": "",
   "note": "No current saturation ceiling was established in the opened source.",
   "status": "unknown",
   "top_score": null
  },
  "sources": [
   {
    "accessed": "2026-09-09",
    "title": "MV-dVRK dataset",
    "url": "https://huggingface.co/datasets/soyeonbot/MV-dVRK"
   }
  ],
  "status": "active",
  "subcategory": "surgical multi-view 3D reconstruction",
  "summary": "MV-dVRK evaluates multi-view 3D reconstruction methods on synchronized stereo endoscopic views with surgical geometry and camera poses.",
  "tags": [
   "surgery",
   "3d",
   "reconstruction"
  ],
  "task_format": "Benchmark instances with task-specific data, visual or geometric inputs and reference annotations."
 }
}