Vision-Tactile-Language-Action Model
视觉-触觉-语言-动作模型VTLAAdvancedA robot model that adds touch sensing to VLA, built for contact-heavy tasks that need a physical feel for the world.
VTLA refers to models that feed tactile-sensor signals into a robot policy alongside camera images and language instructions to produce actions — an extension of VLA (vision-language-action models). The name comes from a May 2025 paper by researchers at Samsung Research China, the Beijing Academy of Artificial Intelligence, and the Institute of Automation, Chinese Academy of Sciences. They built on a Qwen2-VL 7B backbone, converted readings from GelStereo visual-tactile sensors on the gripper fingertips into image-like inputs, trained on simulated peg-in-hole insertion data, and used direct preference optimization (DPO) to ease the mismatch between token-based classification and continuous control. Touch matters because in contact-rich tasks like inserting parts, twisting caps, or grasping fragile objects, the contact point is often hidden from the camera by the fingers themselves, so vision alone can't tell how much force is being applied or whether something is slipping. VTLA is now used as a general label for this class of model, and work in 2026 still explores turning off-the-shelf VLAs into VTLAs with little extra data.
ExampleIn the original VTLA paper's real-robot experiments, the model adjusted insertion position and angle step by step using camera images plus fingertip tactile images, reaching a 95% success rate inserting a square peg into a hole with 0.6mm clearance, and 95–100% on peg shapes it had never seen.
- Also called
- VTLA, Tactile VLA
- Related
- Vision-Language-Action Model · Vision-Based Tactile Sensor · Visuo-Tactile Fusion · Contact-rich Manipulation · Peg-in-Hole Insertion · Force-aware Vision-Language-Action Model
- Sources
- VTLA: Vision-Tactile-Language-Action Model with Preference Learning for Insertion Manipulation (arXiv:2505.09577)
VT-Bridge: Bridging Pretrained Foundation VLAs to VTLAs via Lightweight Residual Adaptation (arXiv:2609.22606) - As of
- 2026-09