{
 "body": "## What it measures\n\nES-MemEval evaluates long-term conversational memory in personalized emotional support. User information is fragmented, implicit, and continuously changing across sessions, so an agent must retrieve more than explicit facts from the latest turn.\n\nThe benchmark covers information extraction, temporal reasoning, conflict detection, abstention, and user modeling. Tasks include question answering, summarization, and dialogue generation through the EvoEmo multi-session dataset.\n\n## How it is scored\n\nThe paper reports benchmark results across open-source long-context models, commercial models, and retrieval-augmented systems. It finds that explicit long-term memory reduces hallucinations and improves personalization, while RAG improves factual consistency but struggles with temporal dynamics and evolving states. The abstract does not establish one universal metric name or maximum; report task type and memory capability with each score.\n\n## Dataset and licence\n\nThe paper names EvoEmo as a multi-session dataset capturing fragmented disclosures and evolving user states. The abstract does not provide an item count, complete split description, or licence. Those fields remain unknown until the official release is inspected.\n\n## Who publishes it\n\nTiantian Chen, Jiaqi Lu, Ying Shen, and Lin Zhang introduced ES-MemEval in a February 2026 arXiv paper accepted to The Web Conference 2026. No independent leaderboard is identified.\n\n## Lineage\n\nES-MemEval addresses limitations of long-term dialogue benchmarks that focus on static explicit fact retrieval. EvoEmo is its associated dataset and should not be treated as a separate benchmark page. The paper does not name a successor.\n\n## Saturation and contamination\n\nThe reported memory, temporal, and retrieval limitations indicate an open benchmark. The paper does not establish model training exposure or whether evaluation conversations are public, so contamination risk is unknown.\n\n## How to run it\n\nUse the multi-session conversations and the paper\u2019s memory-sensitive task prompts when released. Report session count, memory mechanism, retrieval policy, prompt, language, and task type. Compare direct long-context and RAG systems under matched context budgets.\n\n## Reading the numbers\n\nA strong score indicates effective memory for the selected task and conversation history. It does not establish safe or clinically appropriate emotional support. Inspect temporal conflicts, abstentions, and personalization separately. Retrieval gains should be read alongside evidence faithfulness and readability.\n",
 "build": {
  "built_at": "2026-09-09T16:56:50+00:00",
  "commit": "0a599558854c0e238c03a0f0d725239cb28f9d11",
  "eligibility_as_of": "2026-09-09"
 },
 "disposition": {
  "canonical_id": "es_memeval",
  "reasons": [],
  "status": "unassessed",
  "verified_results": []
 },
 "models_covered": [],
 "page": {
  "category": "long-context",
  "contamination": {
   "note": "The paper abstract does not establish training-data exposure.",
   "risk": "unknown"
  },
  "dataset": {
   "languages": [
    "English"
   ],
   "modalities": [
    "text"
   ],
   "public_test_set": null,
   "url": "https://arxiv.org/abs/2602.01885"
  },
  "freshness": {
   "researched": "2026-09-08",
   "researched_by": "GPT-5.6 Luna, luna-stream-a-006 (Codex coordinated)",
   "reviewed": "",
   "reviewed_by": ""
  },
  "harness": {
   "other": "ES-MemEval and EvoEmo evaluation protocol described by the paper."
  },
  "id": "es_memeval",
  "measures": "ES-MemEval tests whether conversational agents use fragmented, implicit, and changing user information over long interactions. It covers question answering, summarization, and dialogue generation through the EvoEmo multi-session dataset.",
  "metric": {
   "direction": "higher_is_better",
   "name": "memory capability score",
   "unit": "score"
  },
  "name": "ES-MemEval",
  "page_kind": "benchmark",
  "paper": {
   "arxiv": "2602.01885",
   "title": "ES-MemEval: Benchmarking Conversational Agents on Personalized Long-Term Emotional Support",
   "url": "https://arxiv.org/abs/2602.01885",
   "year": 2026
  },
  "publisher": {
   "authors": [
    "Tiantian Chen",
    "Jiaqi Lu",
    "Ying Shen",
    "Lin Zhang"
   ],
   "org": "ES-MemEval authors",
   "url": "https://arxiv.org/abs/2602.01885"
  },
  "released": "2026-02",
  "saturation": {
   "note": "The paper reports continuing limitations in long-term memory and evolving user-state handling.",
   "status": "open"
  },
  "sources": [
   {
    "accessed": "2026-09-08",
    "title": "ES-MemEval paper",
    "url": "https://arxiv.org/abs/2602.01885"
   }
  ],
  "subcategory": "emotional support dialogue",
  "summary": "ES-MemEval evaluates long-term conversational memory for personalized emotional support across extraction, temporal reasoning, conflict detection, abstention, and user modeling.",
  "tags": [
   "memory",
   "dialogue",
   "personalization",
   "emotional-support"
  ],
  "task_format": "Multi-session personalized dialogue with memory-sensitive QA, summarization, and generation."
 }
}