Vibe-Eval is an open benchmark of 269 visual-understanding prompts for evaluating multimodal chat models.
unassessed
| Category | multimodal |
|---|---|
| Subcategory | multimodal chat evaluation |
| Page status | active |
| Metric | automatic or human judged answer quality |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 269 |
| Publisher | Reka AI |
Multimodal chat-model performance on visual-understanding prompts, including difficult everyday tasks.
Open-ended multimodal prompts with expert-authored gold responses.
No model card in ModelSpec reports this benchmark yet.