GenAI-Bench

GenAI-Bench evaluates compositional text-to-image and text-to-video generation with human ratings and automated metric comparisons.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorymultimodal
Metrichuman alignment rating
Directionhigher_is_better
Unitrating
Dataset size40000
PublisherGenAI-Bench authors

What it measures

GenAI-Bench tests whether generated visuals follow compositional prompts involving attributes, relationships, logic, and comparison. It evaluates image and video generation and studies whether VQAScore agrees with human judgments.

Task format

Text prompts paired with generated images or videos and human preference ratings.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub