GenAI-Bench evaluates compositional text-to-image and text-to-video generation with human ratings and automated metric comparisons.
unassessed
| Category | multimodal |
|---|---|
| Metric | human alignment rating |
| Direction | higher_is_better |
| Unit | rating |
| Dataset size | 40000 |
| Publisher | GenAI-Bench authors |
GenAI-Bench tests whether generated visuals follow compositional prompts involving attributes, relationships, logic, and comparison. It evaluates image and video generation and studies whether VQAScore agrees with human judgments.
Text prompts paired with generated images or videos and human preference ratings.
No model card in ModelSpec reports this benchmark yet.