{
 "body": "\n## What it measures\n\nVIBE-Bench studies profile-preference conceptual misalignment, where profile cues and query-specific preferences are not semantically aligned. It defines two psychology-grounded tasks and reports 3,504 personas and 12,239 dialogues, with a manually verified gold test set.\n\n## How it is scored\n\nThe abstract establishes benchmark tasks and a gold test set but does not specify the complete metric and aggregation protocol. Report the task-level metric and exact split used by the release.\n\n## Dataset and licence\n\nThe primary source establishes the benchmark release, but the opened materials do not establish a single dataset licence. Confirm the current release terms and split accounting before redistribution.\n\n## Who publishes it\n\nVIBE-Bench is introduced in the September 2026 arXiv paper; the opened abstract does not provide a complete author list or stable repository URL.\n\n## Lineage\n\nNo predecessor or successor was established in the opened primary source.\n\n## Saturation and contamination\n\nThe benchmark materials are public, which creates contamination opportunities. The opened source does not establish a contamination audit or current saturation ceiling.\n\n## How to run it\n\nFollow the official repository or paper protocol, recording the exact model, prompts, evaluator, task version, tool access and timeout. Preserve per-task outcomes when comparing runs.\n\n## Reading the numbers\n\nHigher scores indicate more successful tasks under the selected protocol. Results can depend on evaluator models, prompts, environment setup and aggregation, so compare only matched configurations.\n\n\nThe benchmark is aimed at a failure regime in which surface similarity between profile information and a query is an unreliable guide to the actual preference. This distinction matters when interpreting scores: retrieval of semantically related history is not sufficient evidence of preference reasoning.\n\nThe manually verified gold test set is useful for checking whether a result reflects robust cross-concept mappings rather than accidental correlations. Keep persona construction, dialogue history, query wording and test-set membership fixed when comparing systems.\n\n## Protocol cautions\n\nPersonalization results should include the available profile and dialogue history exactly as supplied by the benchmark. Do not merge the manually verified gold set into development or prompt-tuning data. Per-task errors are useful because a high aggregate can conceal systematic failures on particular conceptual mappings.\n",
 "build": {
  "built_at": "2026-09-09T16:56:50+00:00",
  "commit": "0a599558854c0e238c03a0f0d725239cb28f9d11",
  "eligibility_as_of": "2026-09-09"
 },
 "disposition": {
  "canonical_id": "vibe_bench",
  "reasons": [],
  "status": "unassessed",
  "verified_results": []
 },
 "models_covered": [],
 "page": {
  "aliases": [],
  "category": "reasoning",
  "contamination": {
   "note": "The benchmark materials are public; the opened source does not establish a contamination audit or private rotating holdout.",
   "risk": "medium"
  },
  "dataset": {
   "languages": [
    "en"
   ],
   "license": "",
   "modalities": [
    "text"
   ],
   "public_test_set": true,
   "size": 12239,
   "size_note": "The paper reports 3,504 personas and 12,239 dialogues, including a manually verified gold test set.",
   "splits": "Unknown unless specified by the official release.",
   "url": "https://arxiv.org/abs/2609.00921v1"
  },
  "freshness": {
   "researched": "2026-09-09",
   "researched_by": "GPT-5.6 Luna, luna-stream-c-005 (Codex coordinated)",
   "reviewed": "",
   "reviewed_by": ""
  },
  "harness": {
   "bigbench": "",
   "helm": "",
   "inspect_evals": "",
   "lm_eval": "",
   "opencompass": "",
   "other": "Use the official release protocol and record its evaluator and prompt settings."
  },
  "id": "vibe_bench",
  "last_updated": "",
  "leaderboard_url": "",
  "lineage": {
   "family": "personalization evaluation",
   "predecessor": "",
   "successors": [],
   "variants": []
  },
  "measures": "Whether a personalized language model can infer query-relevant preferences when profile cues and preferences occupy different concept spaces.",
  "metric": {
   "baseline_note": "No universal random or human baseline was established in the opened primary source.",
   "direction": "higher_is_better",
   "human_baseline": null,
   "max_score": 100,
   "name": "preference-reasoning task accuracy",
   "random_baseline": null,
   "unit": "%"
  },
  "name": "VIBE-Bench",
  "page_kind": "benchmark",
  "paper": {
   "arxiv": "2609.00921",
   "title": "VIBE-Bench: Evaluating Personalized Large Language Models When Profiles Don't Mean Preferences",
   "url": "https://arxiv.org/abs/2609.00921v1",
   "year": 2026
  },
  "publisher": {
   "authors": [],
   "org": "VIBE-Bench authors",
   "url": "https://arxiv.org/abs/2609.00921v1"
  },
  "released": "2026-09",
  "repo_url": "https://arxiv.org/abs/2609.00921v1",
  "saturation": {
   "as_of": "",
   "note": "No current saturation ceiling was established in the opened source.",
   "status": "unknown",
   "top_score": null
  },
  "sources": [
   {
    "accessed": "2026-09-09",
    "title": "Primary paper",
    "url": "https://arxiv.org/abs/2609.00921v1"
   },
   {
    "accessed": "2026-09-09",
    "title": "Official repository",
    "url": "https://arxiv.org/abs/2609.00921v1"
   }
  ],
  "status": "active",
  "subcategory": "personalized preference reasoning",
  "summary": "VIBE-Bench tests personalized language models under profile-preference conceptual misalignment using personas and dialogues.",
  "tags": [
   "personalization",
   "preference",
   "dialogue"
  ],
  "task_format": "Personalized dialogue and preference-reasoning tasks using user personas, histories and queries."
 }
}