Embodied AI Glossary中文

3D-ViTac

Advanced

A system that places low-cost tactile-sensor readings and a visual point cloud in the same 3D space to learn fine manipulation.

3D-ViTac was proposed in October 2024 by researchers from Columbia University, UIUC, and the University of Washington (first author Binghao Huang, advised by Yunzhu Li), published at CoRL 2024. Vision alone isn't enough when the gripper itself blocks the camera's view of an object, or when the task requires judging how hard to squeeze. The system covers the soft gripper's fingertips with flexible piezoresistive tactile pads: each pad has a 16×16 grid of 256 sensing cells, each about 3 square millimeters and under 1 millimeter thick, costing about $20, for 1,024 cells total across a bimanual four-finger setup. The key idea is to use robot kinematics to compute each tactile cell's 3D position, turning its reading into a “tactile point,” merged into a single point cloud together with the camera's point cloud, and then trained with imitation learning using a diffusion policy. Across four long-horizon tasks — steaming an egg, preparing grapes, collecting hex keys, and handing over a sandwich — overall success reached 80–90%, versus only 45–60% for a vision-only version.

ExampleWhile gripping an egg, the camera can't tell how much force is being applied, but the tactile points directly show contact location and pressure, letting the policy hold it gently without crushing it; when reorienting a hex key in the gripper, the tactile points tell the policy the key's current orientation.

Also called
3D-ViTac: Learning Fine-Grained Manipulation with Visuo-Tactile Sensing
Related
Visuo-Tactile Fusion · Tactile Sensor · Diffusion Policy · Point Cloud · 3D Diffusion Policy · Bimanual Manipulation
Sources
3D-ViTac: Learning Fine-Grained Manipulation with Visuo-Tactile Sensing (arXiv 2410.24091)
3D-ViTac 项目页 (Chinese)
As of
2024-10

See it in the full glossary →