Embodied AI Glossary中文

GraspVLA

银河通用 GraspVLAAdvanced

A 2025 grasping foundation model from Galbot and others, pretrained mainly on a billion frames of simulated synthetic data.

GraspVLA was proposed by Galbot together with He Wang's group at Peking University, the University of Hong Kong, and the Beijing Academy of Artificial Intelligence, released in May 2025 and published at CoRL 2025. Real-robot data is expensive and hard to scale up, so the team instead generated data at large scale in simulation: the SynGrasp-1B dataset has about 1 billion frames, spanning 240 categories and more than 10,000 object models, with heavy randomization (domain randomization) of initial pose, placement, background, lighting, and material, and photorealistic rendering to narrow the sim-to-real gap. The model uses “progressive action generation”: it autoregressively predicts a 2D detection box for the target object, then predicts a grasp pose, and finally generates an action chunk with flow matching; training mixes this synthetic data with internet image-text data, letting it grasp unseen object categories from open-vocabulary instructions. The dataset and model weights are open-source.

ExampleAfter pretraining on synthetic data alone, the model can grasp unseen objects on a real tabletop zero-shot; if a particular scene calls for a specific grasping style, a small number of real-robot demonstrations is enough for few-shot fine-tuning.

Also called
a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data
Related
SynGrasp-1B · Synthetic Data · Sim-to-Real Transfer · Domain Randomization · Flow Matching · Grasping
Sources
GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data (arXiv 2505.03233)
GraspVLA 项目页 (Chinese)
As of
2025-08

See it in the full glossary →