VIBE-Bench

VIBE-Bench tests personalized language models under profile-preference conceptual misalignment using personas and dialogues.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryreasoning
Subcategorypersonalized preference reasoning
Page statusactive
Metricpreference-reasoning task accuracy
Directionhigher_is_better
Unit%
Dataset size12239
PublisherVIBE-Bench authors

What it measures

Whether a personalized language model can infer query-relevant preferences when profile cues and preferences occupy different concept spaces.

Task format

Personalized dialogue and preference-reasoning tasks using user personas, histories and queries.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub