PRISM-Bench

900 human-checked clips that score text-to-audio-video models on audio type and on-screen vs off-screen sound with an MLLM judge.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorymultimodal
Subcategoryaudio-centric text-to-audio-video generation
Page statusactive
MetricMLLM pairwise dimension totals
Directionhigher_is_better
Dataset size900
PublisherShanghai Artificial Intelligence Laboratory and Meituan

What it measures

PRISM-Bench tests the soundtrack of a generated video, not just the picture. Each item is a text prompt plus a human-verified reference clip. A T2AV system must emit audio and video together. Scoring factorises along audio type (speech, music, sound) and whether the source is on screen or off screen, then grades four perceptual axes: audio-visual coherence, audio quality, audio expressiveness, and prompt following, using 35 fine-grained criteria. The judge is a multimodal LLM in a blind side-by-side comparison against the reference, not a human rater on every run.

Task format

Text prompt in; audio+video out. MLLM-as-a-judge, pairwise vs ground truth. Subsets: On-screen, Off-screen, and Mixed. Speech / Music / Sound cells.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub