Embodied AI Glossary中文

Tactile Representation Learning

触觉表征学习Advanced

Encoding raw tactile-sensor readings into general-purpose features that downstream tasks like slip detection can reuse.

Tactile representation learning studies how to encode raw signals from tactile sensors, such as the gel-deformation images captured by vision-based tactile sensors like GelSight or DIGIT, into compact features that downstream tasks such as slip detection, force estimation, and precision insertion can use directly. The difficulty is that tactile data is scarce, labels are even scarcer, and different sensors image very differently, so switching sensors often means retraining from scratch. Recent work borrows self-supervised learning from vision: Meta's Sparsh (CoRL 2024) pretrains on more than 460,000 tactile images using masking and self-distillation; MIT's T3 uses a shared Transformer with sensor-specific encoders; and AnyTouch (ICLR 2025) learns a unified representation across four sensor types. This is foundational to visuo-tactile fusion and tactile VLAs.

ExampleSparsh is pretrained with self-supervision on more than 460,000 unlabeled visuo-tactile images; on the six tasks of the authors' own TacBench, the paper reports it beats models trained end-to-end per task and per sensor by 95.1% on average.

Also called
Touch Representation Learning, Tactile Pretraining
Related
Vision-Based Tactile Sensor · Self-Supervised Learning · Sparsh · AnyTouch · Visuo-Tactile Fusion · Representation Learning
Sources
Higuera et al. 2024: Sparsh: Self-supervised Touch Representations for Vision-based Tactile Sensing (CoRL 2024)
Zhao et al. 2024: Transferable Tactile Transformers for Representation Learning Across Diverse Sensors and Tasks (T3)
Feng et al. 2025: AnyTouch: Learning Unified Static-Dynamic Representation across Multiple Visuo-tactile Sensors (ICLR 2025)
As of
2025-04

See it in the full glossary →