RO-Bench

Multiple-choice video questions scored on original clips and on text-edited counterfactual versions, so the gap measures robustness rather than raw accuracy.

Also known as: Ro-Bench, Ro-Bench: Robust Video MLLMs Benchmark

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorymultimodal
Subcategoryvideo MLLM robustness under counterfactual edits
Page statusproposed
Metricaccuracy (origin, edit, and drop)
Directionhigher_is_better
Unit%
Dataset size8600
PublisherBeijing University of Posts and Telecommunications

What it measures

RO-Bench (also written Ro-Bench) tests whether a video multimodal LLM still answers correctly after a text-driven edit of the clip. Source videos come from DAVIS, TGVE, MSR-VTT, BalanceCC, and the internet. Captions are rewritten along object, action, background, and style, then a video editor renders a new clip. Questions cover action recognition, object recognition, object existence, and video captioning. Agents in the videos are grouped as human, animal, landscape, or object. The interesting number is the drop from original to edited accuracy, not the original score alone.

Task format

A video plus an English multiple-choice question. Object-existence options are yes, no, and not sure. Other tasks use a gold option from the caption plus LLM distractors.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub