Push-T
CommonA 2D task where a round pusher must push a T-shaped block to a target pose, commonly used to test imitation-learning policies.
Push-T is a 2D tabletop pushing task: an agent controls a round pusher that can only move the T-shaped block on the table through point contact, and must push it into a fixed target position and orientation. It first appeared in Google's 2021 Implicit Behavioral Cloning (IBC) work, and became a standard imitation-learning benchmark after Diffusion Policy adapted it in 2023; Hugging Face later packaged it as the gym-pusht environment. The task looks simple, but the block's pose has to be adjusted bit by bit through contact, and the same situation often has several equally valid human strategies, such as pushing from the left or from the right (action multimodality), which makes it well suited for testing whether a policy can represent a multimodal action distribution. The metric is the overlap ratio between the T block and the target region; in gym-pusht, 95% overlap counts as success. Observations can be either keypoints or 96×96 images.
ExampleThe Diffusion Policy paper compared its method against IBC, LSTM-GMM, and others on simulated Push-T, and also built a real-robot version: trained with a UR5 arm and 136 human demonstrations, requiring more precise multi-stage pushing.
- Also called
- PushT, gym-pusht
- Related
- Diffusion Policy · Action Multimodality · Imitation Learning · LeRobot · Non-prehensile Manipulation · ALOHA Sim (Transfer Cube / Insertion)
- Sources
- Diffusion Policy: Visuomotor Policy Learning via Action Diffusion (arXiv 2303.04137)
huggingface/gym-pusht (GitHub)