Vibe-Eval

Vibe-Eval is an open benchmark of 269 visual-understanding prompts for evaluating multimodal chat models.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorymultimodal
Subcategorymultimodal chat evaluation
Page statusactive
Metricautomatic or human judged answer quality
Directionhigher_is_better
Unit%
Dataset size269
PublisherReka AI

What it measures

Multimodal chat-model performance on visual-understanding prompts, including difficult everyday tasks.

Task format

Open-ended multimodal prompts with expert-authored gold responses.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub