HEART

HEART compares human and LLM responses on the same multi-turn emotional-support conversations using blinded ratings and five interpersonal dimensions.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryhuman-preference
Subcategoryemotional-support dialogue
Metricpairwise preference
Directionhigher_is_better
Unitpreference
PublisherHEART authors

What it measures

HEART evaluates emotional-support dialogue beyond fluency. Human raters and LLM judges assess responses for human alignment, empathic responsiveness, attunement, resonance, and task-following on shared dialogue histories.

Task format

Multi-turn emotional-support conversations with pairwise human and model response evaluation.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub