Visuo-Tactile Fusion
视触觉融合AdvancedCombining what a camera sees with what a tactile sensor feels into a single, unified signal.
Visuo-tactile fusion means using both vision and touch signals together in perception or a policy, and combining them into a unified representation. Vision is good at seeing the big picture and finding a target, but a finger is often occluded the moment it touches an object, and vision can’t tell how much force is being applied or whether something is slipping; tactile sensing fills in exactly that: contact location, pressure distribution, and slip. Common approaches include encoding the two modalities separately and concatenating their features, using attention for cross-modal fusion, or projecting tactile readings into a 3D point cloud and merging it with the visual point cloud. A Stanford team used self-supervised learning to fuse vision and force for peg-in-hole insertion at ICRA 2019; 3D-ViTac (CoRL 2024) merges a flexible tactile array into a point cloud alongside a diffusion policy to handle fragile objects, clearly outperforming vision alone. Note this is distinct from a “vision-based tactile sensor,” which is a type of tactile sensor that uses a camera to photograph an elastic membrane.
ExampleWhile a robotic hand grasps an egg, the camera finds the egg and guides the hand toward it; once contact is made, the tactile array reports the pressure distribution, and the policy keeps its grip force just below the point of crushing it.
- Also called
- Visual-Tactile Fusion
- Related
- Multimodal Fusion · Tactile Sensor · Vision-Based Tactile Sensor · 3D-ViTac · Contact-rich Manipulation · Vision-Tactile-Language-Action Model
- Sources
- Making Sense of Vision and Touch: Self-Supervised Learning of Multimodal Representations for Contact-Rich Tasks (arXiv)
3D-ViTac: Learning Fine-Grained Manipulation with Visuo-Tactile Sensing (arXiv)