Transporter Networks
AdvancedA highly sample-efficient manipulation network that treats pick-and-place as 'moving one region of the image to another location.'
Transporter Networks was released by Robotics at Google (Andy Zeng, Pete Florence, and others) in October 2020, published at CoRL 2020 and a finalist for the best paper award. It treats tabletop pick-and-place tasks as a sequence of 'spatial displacements': it first predicts where to grasp from a top-down image, then crops the depth features around that grasp point and cross-correlates them against the features of the whole scene (comparing them at every location) to find the best placement position and rotation angle. This design needs no object detection step and naturally exploits translation and rotation symmetry, making it very data-efficient — a small number of demonstrations are enough to learn tasks like stacking blocks, kit assembly, rope manipulation, and pushing objects, and it extends to 6-DoF pick-and-place as well. The paper also open-sourced Ravens, a PyBullet-based simulation task suite. The later CLIPort added CLIP-based language understanding on top of it, becoming a common baseline for language-conditioned manipulation.
ExampleFor a kit-assembly task, the model looks at a top-down image to find where to grasp each part, which slot in the mold to place it in, and how much to rotate it — learning this from just a handful of demonstrations.
- Also called
- Transporter, Transporter Networks: Rearranging the Visual World for Robotic Manipulation
- Related
- CLIPort · Pick-and-Place · Rearrangement · Tabletop Manipulation · Imitation Learning · Equivariant Policy / Equivariant Neural Network
- Sources
- Transporter Networks (arXiv 2010.14406)
Transporter Networks project page - As of
- 2020-10