ES-MemEval

ES-MemEval evaluates long-term conversational memory for personalized emotional support across extraction, temporal reasoning, conflict detection, abstention, and user modeling.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorylong-context
Subcategoryemotional support dialogue
Metricmemory capability score
Directionhigher_is_better
Unitscore
PublisherES-MemEval authors

What it measures

ES-MemEval tests whether conversational agents use fragmented, implicit, and changing user information over long interactions. It covers question answering, summarization, and dialogue generation through the EvoEmo multi-session dataset.

Task format

Multi-session personalized dialogue with memory-sensitive QA, summarization, and generation.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub