AVDC
AVDC(从无动作视频学动作)AdvancedFirst generates a video of a robot doing the task, then uses optical flow to derive the actions — no action labels needed.
AVDC was proposed by researchers at National Taiwan University working with MIT (including Yilun Du and Joshua Tenenbaum), published at ICLR 2024 (Spotlight). What makes robot data expensive is the action labels; video with no action labels at all is far more plentiful. Given a current image and a text instruction, AVDC first uses a text-conditioned diffusion video-generation model to “imagine” a video of the task being completed; it then estimates optical flow between adjacent frames (which way each pixel moves), treats this as a dense correspondence, combines it with the first frame's depth to compute the rigid-body pose change of the object, and finally converts that into the arm's grasping and moving actions. The policy is trained using only RGB video, and was validated on Meta-World manipulation, iTHOR navigation, and a real Franka arm; the authors also open-sourced a video-model framework that trains in a day on 4 GPUs. It belongs to the same “use video generation as the policy” line of work as UniPi.
ExampleTraining a video model on just 198 videos of a human hand pushing objects — with no robot actions of any kind — and then using it, with no fine-tuning, to control a simulated robot arm on a pushing task reached 90% success over 40 trials, an example of transfer from human video to a robot with a different embodiment.
- Also called
- Actionless Video through Dense Correspondences, Learning to Act from Actionless Videos through Dense Correspondences
- Related
- Action-free Video · UniPi · Video Generation Model · Optical Flow · Imitation from Observation · Human Video Data
- Sources
- Learning to Act from Actionless Videos through Dense Correspondences (arXiv 2310.08576)
OpenReview: ICLR 2024 spotlight
AVDC 项目页 (Chinese) - As of
- 2024-01