HEART-Bench evaluates whether LLM agents preserve human-like personality and memory-consistent decisions across structured psychological scenarios.
unassessed
| Category | human-preference |
|---|---|
| Subcategory | personality consistency |
| Metric | decision consistency |
| Direction | higher_is_better |
| Unit | percent |
| Dataset size | 673 |
| Publisher | HEART-Bench authors |
HEART-Bench constructs 11 character profiles from orthogonal Big Five traits and pairs each with 1,000 autobiographical-style episodic memories. It tests decisions across 64 DIAMONDS scenarios for personality and value consistency.
673 human-validated multiple-choice decision questions grounded in character memories.
No model card in ModelSpec reports this benchmark yet.