{
 "body": "\n## What it measures\n\nVibe-Eval contains 269 visual-understanding prompts with expert-authored gold responses, including 100 hard prompts. It is designed for open-ended evaluation of multimodal chat models on day-to-day and frontier capability questions.\n\n## How it is scored\n\nThe paper discusses automatic evaluation and human judgment, reporting that automatic evaluation with Reka Core roughly correlates with human judgment. Exact judge prompts and aggregation must be recorded for reproducibility.\n\n## Dataset and licence\n\nThe primary source establishes the benchmark release, but the opened materials do not establish a single dataset licence. Confirm the current release terms and split accounting before redistribution.\n\n## Who publishes it\n\nVibe-Eval is released by Reka AI with evaluation code and data in the official `reka-ai/reka-vibe-eval` repository.\n\n## Lineage\n\nNo predecessor or successor was established in the opened primary source.\n\n## Saturation and contamination\n\nThe benchmark materials are public, which creates contamination opportunities. The opened source does not establish a contamination audit or current saturation ceiling.\n\n## How to run it\n\nFollow the official repository or paper protocol, recording the exact model, prompts, evaluator, task version, tool access and timeout. Preserve per-task outcomes when comparing runs.\n\n## Reading the numbers\n\nHigher scores indicate more successful tasks under the selected protocol. Results can depend on evaluator models, prompts, environment setup and aggregation, so compare only matched configurations.\n\n\nThe hard subset is intended to expose failures that average scores can hide. The authors describe more than half of its hard questions as incorrectly answered by all frontier models in their study, but this is a dated experimental observation and not a permanent claim about every later model.\n\nThe benchmark is open-ended: answers should be assessed against the expert reference rather than by exact string matching. Preserve the image encoding, prompt wording and evaluator version when reproducing results.\n\n## Protocol cautions\n\nAutomatic judging is not interchangeable with human annotation. Report which evaluator produced each number, whether the evaluator saw the reference answer, and how disagreements were handled. The public release makes the task reproducible, but a valid comparison still requires matching image preprocessing and prompt templates.\n",
 "build": {
  "built_at": "2026-09-09T16:56:50+00:00",
  "commit": "0a599558854c0e238c03a0f0d725239cb28f9d11",
  "eligibility_as_of": "2026-09-09"
 },
 "disposition": {
  "canonical_id": "vibe_eval",
  "reasons": [],
  "status": "unassessed",
  "verified_results": []
 },
 "models_covered": [],
 "page": {
  "aliases": [],
  "category": "multimodal",
  "contamination": {
   "note": "The benchmark materials are public; the opened source does not establish a contamination audit or private rotating holdout.",
   "risk": "medium"
  },
  "dataset": {
   "languages": [
    "en"
   ],
   "license": "",
   "modalities": [
    "image",
    "text"
   ],
   "public_test_set": true,
   "size": 269,
   "size_note": "The paper reports 269 prompts, including 100 marked hard.",
   "splits": "Unknown unless specified by the official release.",
   "url": "https://github.com/reka-ai/reka-vibe-eval"
  },
  "freshness": {
   "researched": "2026-09-09",
   "researched_by": "GPT-5.6 Luna, luna-stream-c-005 (Codex coordinated)",
   "reviewed": "",
   "reviewed_by": ""
  },
  "harness": {
   "bigbench": "",
   "helm": "",
   "inspect_evals": "",
   "lm_eval": "",
   "opencompass": "",
   "other": "Use the official release protocol and record its evaluator and prompt settings."
  },
  "id": "vibe_eval",
  "last_updated": "",
  "leaderboard_url": "",
  "lineage": {
   "family": "multimodal chat evaluation",
   "predecessor": "",
   "successors": [],
   "variants": []
  },
  "measures": "Multimodal chat-model performance on visual-understanding prompts, including difficult everyday tasks.",
  "metric": {
   "baseline_note": "No universal random or human baseline was established in the opened primary source.",
   "direction": "higher_is_better",
   "human_baseline": null,
   "max_score": 100,
   "name": "automatic or human judged answer quality",
   "random_baseline": null,
   "unit": "%"
  },
  "name": "Vibe-Eval",
  "page_kind": "benchmark",
  "paper": {
   "arxiv": "2405.02287",
   "title": "Vibe-Eval: A hard evaluation suite for measuring progress of multimodal language models",
   "url": "https://arxiv.org/abs/2405.02287",
   "year": 2024
  },
  "publisher": {
   "authors": [],
   "org": "Reka AI",
   "url": "https://github.com/reka-ai/reka-vibe-eval"
  },
  "released": "2024-05",
  "repo_url": "https://github.com/reka-ai/reka-vibe-eval",
  "saturation": {
   "as_of": "",
   "note": "No current saturation ceiling was established in the opened source.",
   "status": "unknown",
   "top_score": null
  },
  "sources": [
   {
    "accessed": "2026-09-09",
    "title": "Primary paper",
    "url": "https://arxiv.org/abs/2405.02287"
   },
   {
    "accessed": "2026-09-09",
    "title": "Official repository",
    "url": "https://github.com/reka-ai/reka-vibe-eval"
   }
  ],
  "status": "active",
  "subcategory": "multimodal chat evaluation",
  "summary": "Vibe-Eval is an open benchmark of 269 visual-understanding prompts for evaluating multimodal chat models.",
  "tags": [
   "multimodal",
   "visual",
   "chat"
  ],
  "task_format": "Open-ended multimodal prompts with expert-authored gold responses."
 }
}