Embodied AI Glossary中文

LingBot-VA (Robbyant)

蚂蚁灵波 LingBot-VAAdvanced

Robbyant's open-source video-action world model that predicts future frames and outputs robot actions at the same time.

LingBot-VA is a robot control model open-sourced in January 2026 by Robbyant, the embodied-AI company under Ant Group, in a paper titled Causal World Modeling for Robot Control, accepted to RSS 2026. It follows the “world action model” approach: it generates segment by segment with autoregressive diffusion, alternately predicting future video frames and actions within the same sequence; visual and action tokens share a latent space, processed by a Mixture-of-Transformers (MoT). At execution time, each segment is corrected against real observations (closed-loop rollout), and action prediction runs asynchronously in parallel with motor execution to cut latency. Its video encoding reuses the VAE from Tongyi Wanxiang's Wan2.2. Version 2.0, released in July 2026, switched to causal pretraining from scratch with a sparse MoE backbone.

ExampleThe company reports success rates of 92.9% (easy) and 91.6% (hard) across 50 bimanual simulation tasks on RoboTwin 2.0, and an average of 98.5% on LIBERO.

Also called
LingBot-VA 2.0, Causal World Modeling for Robot Control
Related
World Action Model · World Model · Mixture-of-Transformers · Asynchronous Inference · Autoregressive Video Generation · Robbyant
Sources
arXiv 2601.21998: Causal World Modeling for Robot Control
GitHub: Robbyant/lingbot-va
arXiv 2607.08639: Native Video-Action Pretraining for Generalizable Robot Control (LingBot-VA 2.0)
As of
2026-07

See it in the full glossary →