HEART compares human and LLM responses on the same multi-turn emotional-support conversations using blinded ratings and five interpersonal dimensions.
unassessed
| Category | human-preference |
|---|---|
| Subcategory | emotional-support dialogue |
| Metric | pairwise preference |
| Direction | higher_is_better |
| Unit | preference |
| Publisher | HEART authors |
HEART evaluates emotional-support dialogue beyond fluency. Human raters and LLM judges assess responses for human alignment, empathic responsiveness, attunement, resonance, and task-following on shared dialogue histories.
Multi-turn emotional-support conversations with pairwise human and model response evaluation.
No model card in ModelSpec reports this benchmark yet.