{
 "body": "## What it measures\n\nMULTI-Bench evaluates spoken dialogue models in genuinely interactive multi-turn conversations. It emphasizes emotional intelligence rather than isolated speech recognition or single-turn response quality.\n\nThe basic track tests emotion understanding and reasoning. The advanced track tests emotion support and application. The benchmark contains five tasks and about 3.2K samples across eight subsets.\n\n## How it is scored\n\nThe paper reports task performance for six spoken dialogue models across the eight subsets. It finds good basic understanding but remaining gaps in advanced multi-turn interaction and emotion-related reasoning. The abstract does not establish one universal metric name or maximum; report task and track separately.\n\n## Dataset and licence\n\nThe paper establishes about 3.2K samples, five tasks, and eight subsets. It does not state a licence or public test policy in the abstract. Reproductions should record audio source, transcript, turn count, and evaluation framework version.\n\n## Who publishes it\n\nYayue Deng and eight coauthors introduced MULTI-Bench in a November 2025 arXiv submission to ICASSP 2026. No independent leaderboard is established.\n\n## Lineage\n\nMULTI-Bench is a standalone spoken emotional-intelligence benchmark. It is distinct from the Chinese multimodal MULTI family despite the shared name. The paper names no successor.\n\n## Saturation and contamination\n\nThe reported advanced-track gaps indicate an open benchmark. Training exposure is unknown.\n\n## How to run it\n\nUse the reproducible evaluation framework and preserve track, task, subset, audio input, and multi-turn interaction settings. Report model speech interface and whether transcripts are available. Do not collapse basic and advanced tracks without showing both.\n\n## Reading the numbers\n\nA strong result indicates emotional dialogue performance under the selected spoken interaction. It does not prove robust empathy in unseen languages, cultures, or high-stakes support. Compare basic and advanced tracks and inspect multi-turn failures. Audio and transcript conditions materially affect comparability.\n\nBecause the benchmark is interactive, a model may perform well on recognition while failing to sustain an emotionally appropriate response over several turns. The advanced track is therefore a more demanding deployment signal than a single emotion-label score. Report latency, interruption handling, and the availability of conversation history when those factors are part of the evaluation setup.\n",
 "build": {
  "built_at": "2026-09-09T16:56:50+00:00",
  "commit": "0a599558854c0e238c03a0f0d725239cb28f9d11",
  "eligibility_as_of": "2026-09-09"
 },
 "disposition": {
  "canonical_id": "multi_bench",
  "reasons": [],
  "status": "unassessed",
  "verified_results": []
 },
 "models_covered": [],
 "page": {
  "category": "human-preference",
  "contamination": {
   "note": "Training exposure is not established by the paper abstract.",
   "risk": "unknown"
  },
  "dataset": {
   "languages": [
    "English"
   ],
   "modalities": [
    "audio",
    "text"
   ],
   "public_test_set": null,
   "size": 3200,
   "size_note": "About 3.2K samples across five tasks and eight subsets."
  },
  "freshness": {
   "researched": "2026-09-08",
   "researched_by": "GPT-5.6 Luna, luna-stream-a-008 (Codex coordinated)",
   "reviewed": "",
   "reviewed_by": ""
  },
  "harness": {
   "other": "Reproducible MULTI-Bench evaluation framework."
  },
  "id": "multi_bench",
  "measures": "MULTI-Bench tests whether spoken dialogue models sustain interactive conversations with emotional awareness. Its basic track covers emotion understanding and reasoning; its advanced track covers emotion support and application.",
  "metric": {
   "direction": "higher_is_better",
   "name": "task performance",
   "unit": "score"
  },
  "name": "MULTI-Bench",
  "page_kind": "benchmark",
  "paper": {
   "arxiv": "2511.00850",
   "title": "MULTI-Bench: A Multi-Turn Interactive Benchmark for Assessing Emotional Intelligence ability of Spoken Dialogue Models",
   "url": "https://arxiv.org/abs/2511.00850",
   "year": 2025
  },
  "publisher": {
   "authors": [
    "Yayue Deng",
    "Guoqiang Hu",
    "Haiyang Sun",
    "Xiangyu Zhang",
    "Haoyang Zhang",
    "Fei Tian",
    "Xuerui Yang",
    "Gang Yu",
    "Eng Siong Chng"
   ],
   "org": "MULTI-Bench authors",
   "url": "https://arxiv.org/abs/2511.00850"
  },
  "released": "2025-11",
  "saturation": {
   "note": "The paper reports remaining gaps in advanced interactive dialogue and emotional reasoning.",
   "status": "open"
  },
  "sources": [
   {
    "accessed": "2026-09-08",
    "title": "MULTI-Bench paper",
    "url": "https://arxiv.org/abs/2511.00850"
   }
  ],
  "subcategory": "spoken emotional intelligence",
  "summary": "MULTI-Bench evaluates spoken dialogue models on multi-turn emotional intelligence through basic understanding and advanced support tracks.",
  "tags": [
   "speech",
   "emotion",
   "multi-turn",
   "dialogue"
  ],
  "task_format": "Multi-turn spoken dialogue across five tasks and eight subsets."
 }
}