VGA-Bench scores text-to-video models on aesthetic quality, aesthetic tags, and generation quality using 1,016 prompts and 52 sub-dimensions.
unassessed
| Category | generation |
|---|---|
| Subcategory | text-to-video aesthetics and generation quality |
| Page status | active |
| Metric | dimension-averaged aesthetic score, tag accuracy, and generation level |
| Direction | higher_is_better |
| Unit | score |
| Dataset size | 1016 |
| Dataset licence | CC-BY-NC-4.0 |
| Publisher | Ant Group, Beijing Film Academy, and Beijing Institute for General Artificial Intelligence (BIGAI) |
VGA-Bench evaluates generated videos, not language-model answers. Given a text prompt that names the target aesthetic or quality attribute, a generator produces a clip, and dedicated assessors score aesthetic quality (composition, lighting, colour, and related attributes), classify aesthetic tags, and rate generation quality (prompt alignment, physical plausibility, and basic visual stability). The skill is whether the clip is both technically faithful and photographically controlled, not whether a chatbot can describe video.
Text-to-video generation from an official prompt; videos are scored by VAQA-Net, VTag-Net, and VGQA-Net (or by human labels on the annotated subset).
No model card in ModelSpec reports this benchmark yet.