Recent advances in video generation have made it possible to fabricate realistic depictions of wars, disasters, and public emergencies, raising serious concerns about misinformation. However, existing benchmarks provide limited insight into how detectors and generators behave in such high-stakes settings.
To address this gap, researchers introduce RA-Bench, a benchmark for AI-generated video detection that uses real videos as anchors. It contains 17,886 videos, including 1,830 real-video anchors across 10 social-risk categories and 16,056 generated clips from four open-source and five closed-source generators.
The evaluation spans three dimensions. First, they assess detector generalization across seven traditional detectors, ten zero-shot multimodal models under three review settings, and two MLLMs fine-tuned on AI-generated video detection. None of the three detector families generalizes consistently across RA-Bench instances.
Second, they examine how detectability varies with generation quality, conditioning information, and sampling seeds. The results show that generation properties affect detector families differently, while source-level detection patterns remain stable across seeds.
Finally, they study human authenticity judgments and detector reliability during social dissemination. They find that videos that mislead people are also difficult for current detectors, and that social dissemination makes detection harder. Together, these findings highlight the need for detectors robust to evolving video generators.