Embodied AI Glossary中文

VBench: Comprehensive Benchmark Suite for Video Generative Models

VBench 视频生成评测基准Advanced

An open-source benchmark that scores video-generation quality separately across 16 dimensions instead of giving one overall number.

VBench was proposed by a team from Nanyang Technological University and the Shanghai AI Laboratory, among others (first author Ziqi Huang), released in November 2023 and selected as a CVPR 2024 Highlight. Rather than a single vague score, it breaks text-to-video quality into 16 dimensions — subject consistency, background consistency, temporal flickering, motion smoothness, dynamic degree, aesthetic quality, object class, and spatial relationships among them — each with dedicated prompts and an automated evaluation method, checked against human preference annotations to see whether the scoring agrees with human judgment. The follow-up VBench++ extended it to image-to-video, long-video, and trustworthiness evaluation; VBench-2.0, from March 2025, shifted toward “intrinsic faithfulness” — commonsense reasoning, physical realism, and human motion. In embodied AI, evaluating video world models often borrows some of its metrics.

ExampleWhen evaluating a model meant to generate robot training videos, VBench's subject-consistency and motion-smoothness dimensions can check whether the robot arm deforms or jitters partway through the clip; checking whether the motion obeys physics still requires a dedicated benchmark such as Physics-IQ.

Also called
VBench, VBench++, VBench-2.0
Related
Video Generation Model · World Model · Fréchet Video Distance · Physics-IQ (Do generative video models understand physical principles?) · WorldScore: A Unified Evaluation Benchmark for World Generation · EWMBench
Sources
VBench: Comprehensive Benchmark Suite for Video Generative Models (arXiv 2311.17982)
Vchitect/VBench GitHub 仓库 (Chinese)
As of
2025-03

See it in the full glossary →