Embodied AI Glossary中文

Force-aware Vision-Language-Action Model

力觉 VLAAdvanced

A VLA that takes force/torque signals as input alongside vision and language.

A force-aware VLA is a vision-language-action model that adds force signals, such as those from a six-axis force/torque sensor (measuring force along three axes and torque around three axes), as an input modality alongside image and language. An ordinary VLA relies only on cameras, but in contact-rich tasks like plugging in a cord, wiping a table, or assembly, whether contact is properly made and how much force is being applied is often invisible to a camera, leading to getting stuck or pressing too hard. The representative work, ForceVLA (Yu et al., NeurIPS 2025), builds on π0 and uses a force-aware mixture-of-experts module (FVLMoE) to fuse a force token during action decoding, reaching 23.2% higher average performance than the π0 baseline across five contact-rich tasks, while simply concatenating force into the input only gave a small improvement. ForceVLA2, from March 2026, adds hybrid force/position control and reports 48% and 35% improvements over π0 and π0.5 respectively.

ExampleWhen plugging in a cord, as the plug touches the edge of the socket the camera image barely changes, but the force sensor reading jumps; ForceVLA uses this signal to correct the pose, reaching up to 80% success on plugging tasks.

Also called
Force-aware VLA
Related
Vision-Language-Action Model · Six-Axis Force/Torque Sensor · Contact-rich Manipulation · Hybrid Force/Position Control · Vision-Tactile-Language-Action Model · Mixture of Experts
Sources
ForceVLA: Enhancing VLA Models with a Force-aware MoE for Contact-rich Manipulation (arXiv 2505.22159)
ForceVLA2: Unleashing Hybrid Force-Position Control with Force Awareness (arXiv 2603.15169)
Learning Physical Interaction: A Survey of Tactile- and Force-aware Robot Learning (arXiv 2608.07558)
As of
2026-09

See it in the full glossary →