HEART-Bench

HEART-Bench evaluates whether LLM agents preserve human-like personality and memory-consistent decisions across structured psychological scenarios.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryhuman-preference
Subcategorypersonality consistency
Metricdecision consistency
Directionhigher_is_better
Unitpercent
Dataset size673
PublisherHEART-Bench authors

What it measures

HEART-Bench constructs 11 character profiles from orthogonal Big Five traits and pairs each with 1,000 autobiographical-style episodic memories. It tests decisions across 64 DIAMONDS scenarios for personality and value consistency.

Task format

673 human-validated multiple-choice decision questions grounded in character memories.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub