Embodied AI Glossary中文

ShengShu Technology

生数科技Advanced

A Tsinghua-linked multimodal generation company behind the Vidu video model, now extending into embodied world models.

ShengShu Technology was founded in Beijing in March 2023, with a core team from Tsinghua University; co-founder and chief scientist Jun Zhu is a Tsinghua computer science professor, and Jiayu Tang is CEO. In September 2022 the team proposed the U-ViT architecture, one of the earlier works to use a Transformer as the backbone of a diffusion model. Its signature product is the video-generation model Vidu, released in April 2024 and launched globally that July. It has extended video generation into robotics: working with Tsinghua on the embodied model Vidar, which predicts future frames with a video diffusion model and then decodes them into actions, and separately building the world-action model Motus. In April 2026 it closed a Series B of nearly RMB 2 billion led by Alibaba Cloud, positioning itself as building a “general-purpose world model.”

ExampleShengShu's Vidar first generates a video of the robot completing a task, then uses an inverse-dynamics model to infer the action at each step from the video.

Also called
ShengShu
Related
Vidar · Motus · Video Generation Model · World Action Model · World Model · Diffusion Transformer
Sources
生数科技完成近20亿元B轮融资(量子位) (Chinese)
京企生数科技完成超6亿元A+轮融资(国家科技传播中心) (Chinese)
As of
2026-04

See it in the full glossary →