LingBot-VLA (Robbyant)
蚂蚁灵波 LingBot-VLAAdvancedRobbyant's open-source VLA foundation model, pretrained on about 20,000 hours of bimanual real-robot data across 9 arm configurations.
LingBot-VLA is a vision-language-action (VLA) foundation model open-sourced by Robbyant in January 2026, in a paper titled A Pragmatic VLA Foundation Model, emphasizing practicality: strong generalization and low data and compute cost when adapting to a new platform. It uses Qwen2.5-VL-3B as its vision-language backbone (also supporting PaliGemma), pretrained on about 20,000 hours of real-robot data across 9 mainstream bimanual configurations, and systematically evaluated with 100 tasks each (GM-100) on 4 platforms including AgiBot G1, AgileX, and Galaxea R1 Pro. Its companion training code reaches a throughput of 261 samples per second on 8 GPUs, 1.5 to 2.8 times faster than existing VLA codebases. The July 2026 version 2.0 expanded the data to about 60,000 hours (including 10,000 hours of human first-person video), and extended the action space to the head, waist, base, and dexterous hands.
ExampleIn the GM-100 evaluation, each task gets only 130 post-training demonstrations, and success rate is then compared against other VLA models on the same real-robot platform.
- Also called
- LingBot-VLA 2.0, A Pragmatic VLA Foundation Model
- Related
- Vision-Language-Action Model · Qwen-VL · Cross-Embodiment · Post-training · π0 · Robbyant
- Sources
- arXiv 2601.18692: A Pragmatic VLA Foundation Model
Robbyant 官网:LingBot-VLA (Chinese)
arXiv 2607.06403: From Foundation to Application (LingBot-VLA 2.0) - As of
- 2026-07