Multiple-choice video questions scored on original clips and on text-edited counterfactual versions, so the gap measures robustness rather than raw accuracy.
unassessed
| Category | multimodal |
|---|---|
| Subcategory | video MLLM robustness under counterfactual edits |
| Page status | proposed |
| Metric | accuracy (origin, edit, and drop) |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 8600 |
| Publisher | Beijing University of Posts and Telecommunications |
RO-Bench (also written Ro-Bench) tests whether a video multimodal LLM still answers correctly after a text-driven edit of the clip. Source videos come from DAVIS, TGVE, MSR-VTT, BalanceCC, and the internet. Captions are rewritten along object, action, background, and style, then a video editor renders a new clip. Questions cover action recognition, object recognition, object existence, and video captioning. Agents in the videos are grouped as human, animal, landscape, or object. The interesting number is the drop from original to edited accuracy, not the original score alone.
A video plus an English multiple-choice question. Object-existence options are yes, no, and not sure. Other tasks use a gold option from the caption plus LLM distractors.
No model card in ModelSpec reports this benchmark yet.