VGA-BenchV2 keeps VGA-Bench's 52-dimension prompt suite and adds larger human labels, hybrid judges, and aesthetic reward-model fine-tuning.
unassessed
| Category | generation |
|---|---|
| Subcategory | text-to-video aesthetics, generation quality, and reward-model optimization |
| Page status | active |
| Metric | dimension-averaged aesthetic score, tag classification accuracy, and generation level |
| Direction | higher_is_better |
| Unit | score |
| Dataset size | 1016 |
| Dataset licence | CC-BY-NC-4.0 |
| Publisher | Ant Group, Beijing Film Academy, and Beijing Institute for General Artificial Intelligence (BIGAI) |
VGA-BenchV2 evaluates the same text-to-video task as VGA-Bench: a model generates a clip from a dimension-aligned prompt, then judges score aesthetic quality, aesthetic tags, and generation quality. The V2 contribution is supervision and judges, not a new prompt list. Human labels grow by 36,000 task-level annotations. VAQA-Net still predicts continuous aesthetic scores; VTag-Net and VGQA-Net become Qwen-based vision-language evaluators. An optional loop uses VAQA-Net as a reward model for generator fine-tuning.
Text-to-video generation from the VGA-Bench prompt suite; clips are scored by the V2 hybrid evaluators, with an optional reinforcement-learning fine-tune using VAQA-Net reward.
No model card in ModelSpec reports this benchmark yet.