Embodied AI Glossary中文

3D VLA

Advanced

A VLA that explicitly feeds spatial information — depth, point clouds, 3D position — into the model.

A standard VLA (vision-language-action model, which looks at an image, hears an instruction, and outputs an action directly) mostly takes only 2D images as input, but a robot's actions happen in 3D space, where distance, height, and occlusion are hard to judge accurately from a 2D image alone. 3D VLA is a general term for VLAs that add depth, point clouds, or 3D positional encoding to the input or an intermediate representation. Two papers are representative. 3D-VLA, from a UMass Amherst team in March 2024, adds interaction tokens on top of a 3D large language model and uses a diffusion model to generate a goal image and goal point cloud, 'imagining' the scene after manipulation. SpatialVLA (RSS 2025), by Qu and colleagues in January 2025, is built on PaliGemma2, uses Ego3D position encoding to inject 3D spatial information into visual features, discretizes continuous actions into spatial tokens with an adaptive action grid, and is pretrained on 1.1 million real-robot trajectories from OXE and RH20T. The main goal of this line of work is more stable spatial judgment when the camera viewpoint or the robot changes.

ExampleSpatialVLA-4B encodes each image patch's 3D position into its visual token, needs no camera calibration, and can re-partition its action grid to adapt when moved to a new robot.

Also called
Spatial VLA, SpatialVLA, 3D-VLA
Related
Vision-Language-Action Model · Spatial Reasoning · Point Cloud · Depth Estimation · 3D Diffusion Policy · World Model
Sources
3D-VLA: A 3D Vision-Language-Action Generative World Model (arXiv 2403.09631)
SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model (arXiv 2501.15830)
SpatialVLA GitHub
As of
2025-01

See it in the full glossary →