Vidar
生数 VidarAdvancedA robot manipulation model that first predicts future frames with a video diffusion model, then infers the action from them.
Vidar was released in July 2025 by Jun Zhu's lab (TSAIL) at Tsinghua University; its name stands for VIdeo Diffusion for Action Reasoning. Real-robot experiments are built on Vidu 2.0, the video model from ShengShu Technology, and the paper notes that part of the work was done at ShengShu, with reports describing it as a joint release between ShengShu and Tsinghua. The approach has two steps: a video diffusion model first 'imagines' an upcoming video from the instruction and current image; then a masked inverse dynamics model (MIDM), which infers actions from consecutive frames, translates that video into robot actions, and automatically learns to focus only on action-relevant pixels such as the robot arm, filtering out background clutter. The video model first goes through continued embodied-domain pretraining on 750,000 multi-view trajectories across three robot platforms. On a robot it has never seen, the paper reports that only about 20 minutes of human demonstration is needed to beat baseline methods, and it generalizes to new tasks, backgrounds, and camera layouts.
ExampleA new Aloha bimanual robot is adapted with only about 20 minutes of human demonstration; Vidar then generates a video of the task being completed from a new language instruction, and an inverse dynamics model translates that video segment by segment into joint actions for execution.
- Also called
- Embodied Video Diffusion Model for Generalist Manipulation, Vidar: VIdeo Diffusion for Action Reasoning
- Related
- Video Generation Model · Inverse Dynamics Model · Video Prediction Policy · Cross-Embodiment · World Action Model · ShengShu Technology
- Sources
- Vidar (arXiv 2507.12898)
Vidar 论文 HTML v4(机构与 Vidu 2.0 基座) (Chinese)
Vidar & AnyPos project page - As of
- 2025-12