{
 "body": "## What it measures\n\nHEART evaluates emotional-support dialogue as an interpersonal capability. It places humans and language models on the same multi-turn conversation histories, then compares their responses. The benchmark focuses on reading emotion, adapting tone, and handling resistance, frustration, and distress.\n\nIts rubric has five dimensions: Human Alignment, Empathic Responsiveness, Attunement, Resonance, and Task-Following. These dimensions separate supportive quality from general language fluency.\n\n## How it is scored\n\nResponses are evaluated by blinded human raters and an ensemble of LLM judges. The paper reports pairwise preferences and judge-human agreement. It finds about 80% alignment between LLM-judge and human pairwise preferences, while humans retain advantages in adaptive reframing and nuanced tone shifts. Exact aggregation and scale are not established in the abstract.\n\n## Dataset and licence\n\nThe paper describes multi-turn emotional-support conversations but does not state a total item count, licence, or public test policy in the abstract. Those fields remain unknown. Because the evaluation concerns sensitive dialogue, reproductions should document privacy and consent handling.\n\n## Who publishes it\n\nLaya Iyer, Kriti Aggarwal, Sanmi Koyejo, Gail Heyman, Desmond C. Ong, and Subhabrata Mukherjee introduced HEART in a 2026 arXiv paper. No public leaderboard is established.\n\n## Lineage\n\nHEART is a standalone emotional-support dialogue benchmark. It is related in subject to ES-MemEval but differs by directly comparing humans and models on shared conversations rather than testing long-term memory. The paper names no successor.\n\n## Saturation and contamination\n\nSeveral frontier models approach or surpass average human responses on perceived empathy and consistency, but humans remain stronger on nuanced adaptive behaviors. This indicates an open capability axis. Training exposure is unknown.\n\n## How to run it\n\nUse the same dialogue histories, blinded comparison design, rubric, and judge ensemble. Report human versus model comparison, judge model, turn type, and each HEART dimension separately. Avoid treating judge scores as human ratings without reporting their agreement.\n\n## Reading the numbers\n\nA favorable preference means raters judged the response more supportive under the selected rubric. It does not establish clinical effectiveness or safe crisis handling. Examine adversarial turns and dimension-level results, especially reframing and tone. Human baselines are central to interpretation because the benchmark is explicitly comparative.\n",
 "build": {
  "built_at": "2026-09-09T16:56:50+00:00",
  "commit": "0a599558854c0e238c03a0f0d725239cb28f9d11",
  "eligibility_as_of": "2026-09-09"
 },
 "disposition": {
  "canonical_id": "heart",
  "reasons": [],
  "status": "unassessed",
  "verified_results": []
 },
 "models_covered": [],
 "page": {
  "category": "human-preference",
  "contamination": {
   "note": "The paper abstract does not establish training-data exposure.",
   "risk": "unknown"
  },
  "dataset": {
   "languages": [
    "English"
   ],
   "modalities": [
    "text"
   ],
   "public_test_set": null
  },
  "freshness": {
   "researched": "2026-09-08",
   "researched_by": "GPT-5.6 Luna, luna-stream-a-007 (Codex coordinated)",
   "reviewed": "",
   "reviewed_by": ""
  },
  "harness": {
   "other": "HEART blinded human-rating and ensemble LLM-judge protocol."
  },
  "id": "heart",
  "measures": "HEART evaluates emotional-support dialogue beyond fluency. Human raters and LLM judges assess responses for human alignment, empathic responsiveness, attunement, resonance, and task-following on shared dialogue histories.",
  "metric": {
   "baseline_note": "The benchmark directly compares humans and models; no universal maximum is established.",
   "direction": "higher_is_better",
   "name": "pairwise preference",
   "unit": "preference"
  },
  "name": "HEART",
  "page_kind": "benchmark",
  "paper": {
   "arxiv": "2601.19922",
   "title": "HEART: A Unified Benchmark for Assessing Humans and LLMs in Emotional Support Dialogue",
   "url": "https://arxiv.org/abs/2601.19922",
   "year": 2026
  },
  "publisher": {
   "authors": [
    "Laya Iyer",
    "Kriti Aggarwal",
    "Sanmi Koyejo",
    "Gail Heyman",
    "Desmond C. Ong",
    "Subhabrata Mukherjee"
   ],
   "org": "HEART authors",
   "url": "https://arxiv.org/abs/2601.19922"
  },
  "released": "2026-01",
  "saturation": {
   "note": "The paper reports differences between model and human strengths, especially on adversarial turns.",
   "status": "open"
  },
  "sources": [
   {
    "accessed": "2026-09-08",
    "title": "HEART paper",
    "url": "https://arxiv.org/abs/2601.19922"
   }
  ],
  "subcategory": "emotional-support dialogue",
  "summary": "HEART compares human and LLM responses on the same multi-turn emotional-support conversations using blinded ratings and five interpersonal dimensions.",
  "tags": [
   "empathy",
   "emotional-support",
   "human-evaluation",
   "dialogue"
  ],
  "task_format": "Multi-turn emotional-support conversations with pairwise human and model response evaluation."
 }
}