AnyTouch
AdvancedA unified tactile-representation model spanning multiple vision-based tactile sensors, from Renmin University’s Di Hu group, covering both static and dynamic touch.
AnyTouch is a visuo-tactile representation learning framework proposed by Di Hu’s group at Renmin University of China together with Bin Fang at Beijing University of Posts and Telecommunications and others, published at ICLR 2025. Vision-based tactile sensors use a camera to image the deformation of an elastic gel layer, but different sensor models image very differently, making data and models hard to share across them. The authors first collected the TacQuad dataset: four sensors — GelSight Mini, DIGIT, a custom DuraGel, and Tac3D — touching the same location on the same object, yielding 72,606 aligned frames of contact data, paired with visual images and text descriptions of tactile properties. The model takes both tactile images and tactile video as input, uses masked modeling to learn pixel-level detail, and then learns sensor-agnostic semantic features by aligning with vision and language and by matching across sensors. In February 2026 the team released AnyTouch 2, shifting the focus to dynamic tactile sensing with force information.
ExampleIn the paper’s real-robot experiment, a robot arm pours 60 grams of small beads out of a cylinder using only tactile feedback, and the error between the poured mass and the target mass is used to evaluate different tactile representations.
- Also called
- AnyTouch: Learning Unified Static-Dynamic Representation across Multiple Visuo-tactile Sensors, AnyTouch 2
- Related
- Vision-Based Tactile Sensor · Tactile Representation Learning · GelSight Mini · DIGIT · Sparsh · Visuo-Tactile Fusion
- Sources
- AnyTouch (ICLR 2025, arXiv:2502.12191)
AnyTouch 2: General Optical Tactile Representation Learning For Dynamic Tactile Perception (arXiv:2602.09617) - As of
- 2026-02