Hugging Face researchers have unveiled Sci-VBench, a new benchmark designed to evaluate how well AI video generation models handle knowledge- and reasoning-intensive tasks in scientific domains. The benchmark comprises 1,253 expert-annotated examples spanning 60 subjects across four core disciplines: Natural Science, Healthcare, Humanities & Social Sciences, and Engineering.
Unlike existing benchmarks that focus on visual realism, Sci-VBench requires models to generate temporally rich videos that demand scientific reasoning and knowledge-grounded synthesis. The authors also establish a rubric-based evaluation protocol, showing that both non-expert human evaluators and MLLM-as-Judge systems can achieve relatively high agreement with expert judgments, enabling reproducible evaluation at scale.
The team benchmarked 16 frontier proprietary and open-source models. While automatic perceptual-quality scores clustered tightly across systems, performance on Prompt Grounding and Scientific and Causal Correctness varied substantially, with a pronounced proprietary-open-source gap. This indicates that advances in visual realism have not yet translated into reliable modeling of scientific and causal dynamics.
The study also evaluates the latest models, including Gemini-Omni-Flash, HappyHorse-1.1, and MiniMax-H3. The authors welcome feedback on the benchmark.