ES-MemEval evaluates long-term conversational memory for personalized emotional support across extraction, temporal reasoning, conflict detection, abstention, and user modeling.
unassessed
| Category | long-context |
|---|---|
| Subcategory | emotional support dialogue |
| Metric | memory capability score |
| Direction | higher_is_better |
| Unit | score |
| Publisher | ES-MemEval authors |
ES-MemEval tests whether conversational agents use fragmented, implicit, and changing user information over long interactions. It covers question answering, summarization, and dialogue generation through the EvoEmo multi-session dataset.
Multi-session personalized dialogue with memory-sensitive QA, summarization, and generation.
No model card in ModelSpec reports this benchmark yet.