Force-aware Vision-Language-Action Model
力觉 VLAAdvancedA VLA that takes force/torque signals as input alongside vision and language.
A force-aware VLA is a vision-language-action model that adds force signals, such as those from a six-axis force/torque sensor (measuring force along three axes and torque around three axes), as an input modality alongside image and language. An ordinary VLA relies only on cameras, but in contact-rich tasks like plugging in a cord, wiping a table, or assembly, whether contact is properly made and how much force is being applied is often invisible to a camera, leading to getting stuck or pressing too hard. The representative work, ForceVLA (Yu et al., NeurIPS 2025), builds on π0 and uses a force-aware mixture-of-experts module (FVLMoE) to fuse a force token during action decoding, reaching 23.2% higher average performance than the π0 baseline across five contact-rich tasks, while simply concatenating force into the input only gave a small improvement. ForceVLA2, from March 2026, adds hybrid force/position control and reports 48% and 35% improvements over π0 and π0.5 respectively.
ExampleWhen plugging in a cord, as the plug touches the edge of the socket the camera image barely changes, but the force sensor reading jumps; ForceVLA uses this signal to correct the pose, reaching up to 80% success on plugging tasks.
- Also called
- Force-aware VLA
- Related
- Vision-Language-Action Model · Six-Axis Force/Torque Sensor · Contact-rich Manipulation · Hybrid Force/Position Control · Vision-Tactile-Language-Action Model · Mixture of Experts
- Sources
- ForceVLA: Enhancing VLA Models with a Force-aware MoE for Contact-rich Manipulation (arXiv 2505.22159)
ForceVLA2: Unleashing Hybrid Force-Position Control with Force Awareness (arXiv 2603.15169)
Learning Physical Interaction: A Survey of Tactile- and Force-aware Robot Learning (arXiv 2608.07558) - As of
- 2026-09