Embodied AI Glossary中文

BridgeVLA

Advanced

A 3D-manipulation VLA that projects point clouds into 2D images and reads off actions as heatmaps on them.

BridgeVLA was proposed in June 2025 by teams at the Institute of Automation, Chinese Academy of Sciences, and ByteDance Seed, published at NeurIPS 2025. The problem: VLMs are pretrained on 2D images and text, while 3D manipulation needs to handle point clouds and output 3D poses — the two are misaligned, so knowledge transfers poorly between them. BridgeVLA renders a point cloud into several multi-view 2D images and feeds them into PaliGemma, having the model output a 2D heatmap on each image (each pixel representing the likelihood of “act here”), then combines the multi-view heatmaps to determine the 3D position the end effector should go to; before formal training, it's also pretrained to output heatmaps using object-detection data. It raised success rate on RLBench from 81.4% to 88.2%. In August 2026 the team also open-sourced BridgeVLA++.

ExampleTested on a real Franka Research 3 arm across more than 10 tasks, given only 3 demonstrations per task, it reached an average success rate of 96.8%.

Also called
BridgeVLA++, BridgeVLA: Input-Output Alignment for Efficient 3D Manipulation Learning with Vision-Language Models
Related
3D VLA · PaliGemma · RVT-2 · RLBench · The Colosseum: A Benchmark for Evaluating Generalization for Robotic Manipulation · Keyframe Action Prediction
Sources
BridgeVLA (arXiv 2506.07961)
BridgeVLA GitHub
As of
2026-08

See it in the full glossary →