VGA-Bench

VGA-Bench scores text-to-video models on aesthetic quality, aesthetic tags, and generation quality using 1,016 prompts and 52 sub-dimensions.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorygeneration
Subcategorytext-to-video aesthetics and generation quality
Page statusactive
Metricdimension-averaged aesthetic score, tag accuracy, and generation level
Directionhigher_is_better
Unitscore
Dataset size1016
Dataset licenceCC-BY-NC-4.0
PublisherAnt Group, Beijing Film Academy, and Beijing Institute for General Artificial Intelligence (BIGAI)

What it measures

VGA-Bench evaluates generated videos, not language-model answers. Given a text prompt that names the target aesthetic or quality attribute, a generator produces a clip, and dedicated assessors score aesthetic quality (composition, lighting, colour, and related attributes), classify aesthetic tags, and rate generation quality (prompt alignment, physical plausibility, and basic visual stability). The skill is whether the clip is both technically faithful and photographically controlled, not whether a chatbot can describe video.

Task format

Text-to-video generation from an official prompt; videos are scored by VAQA-Net, VTag-Net, and VGQA-Net (or by human labels on the annotated subset).

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub