Embodied AI Glossary中文

3D Diffuser Actor

Advanced

An imitation-learning policy that combines diffusion policies with 3D scene representations to generate robot-arm end-effector trajectories.

3D Diffuser Actor was proposed in February 2024 by Katerina Fragkiadaki's group at Carnegie Mellon University (Tsung-Wei Ke, Nikolaos Gkanatsios), published at CoRL 2024. Two earlier lines of work each had a strength: diffusion policies can represent multiple valid ways of doing something (action multimodality), while 3D policies fuse multi-view images into 3D features using depth, making them more robust to camera-viewpoint changes. This model combines both: it lifts image features into 3D points using depth, then uses a denoising Transformer with 3D relative-position attention, conditioned on the language instruction and proprioceptive state, to progressively denoise a noised trajectory of end-effector poses. At release it beat the previous best method by 18.1 absolute percentage points in the multi-view RLBench setting and 13.1 points single-view, with a 9% relative improvement on CALVIN; on a real Franka arm, it learned 12 tasks from roughly a dozen demonstrations each.

ExampleOn RLBench's “open drawer” task, it builds a 3D feature point cloud of the scene from multiple RGB-D cameras, then repeatedly denoises from random noise to produce the arm's next key pose — a position in front of the handle with a specific gripper orientation.

Also called
3D Diffuser Actor: Policy Diffusion with 3D Scene Representations
Related
Diffusion Policy · 3D Diffusion Policy · PerAct · RLBench · CALVIN Benchmark · Keyframe Action Prediction
Sources
3D Diffuser Actor: Policy Diffusion with 3D Scene Representations (arXiv 2402.10885)
3D Diffuser Actor 项目页 (Chinese)
As of
2024-07

See it in the full glossary →