900 human-checked clips that score text-to-audio-video models on audio type and on-screen vs off-screen sound with an MLLM judge.
unassessed
| Category | multimodal |
|---|---|
| Subcategory | audio-centric text-to-audio-video generation |
| Page status | active |
| Metric | MLLM pairwise dimension totals |
| Direction | higher_is_better |
| Dataset size | 900 |
| Publisher | Shanghai Artificial Intelligence Laboratory and Meituan |
PRISM-Bench tests the soundtrack of a generated video, not just the picture. Each item is a text prompt plus a human-verified reference clip. A T2AV system must emit audio and video together. Scoring factorises along audio type (speech, music, sound) and whether the source is on screen or off screen, then grades four perceptual axes: audio-visual coherence, audio quality, audio expressiveness, and prompt following, using 35 fine-grained criteria. The judge is a multimodal LLM in a blind side-by-side comparison against the reference, not a human rater on every run.
Text prompt in; audio+video out. MLLM-as-a-judge, pairwise vs ground truth. Subsets: On-screen, Off-screen, and Mixed. Speech / Music / Sound cells.
No model card in ModelSpec reports this benchmark yet.