GraspGen
AdvancedNVIDIA’s framework that generates 6-DoF grasp poses with a diffusion model, then scores and filters them.
GraspGen is a 6-DoF grasp-generation framework NVIDIA released publicly in July 2025 (Murali, Fox, Eppner, and colleagues). Given an object’s point cloud, it first uses a diffusion Transformer to generate a large batch of candidate grasp poses (the gripper’s 3D position and orientation), then uses a discriminator to score each grasp and filter out the poor ones; the discriminator is trained directly on grasps sampled from the generator, which the paper calls on-generator training. It was released alongside a dataset of over 53 million simulated grasps, covering the Franka gripper, the Robotiq 2F-140, suction cups, and other end effectors, aimed at the problem that learned grasping models often don’t transfer well to a new gripper or a new real scene. The paper reports state-of-the-art results on the FetchBench simulation benchmark, and it also works on noisy real-world point clouds.
ExampleA segmentation model first cuts the target object out of a depth point cloud; the object’s point cloud is fed into GraspGen to get a batch of scored grasp poses, and the highest-scoring reachable one is passed to cuRobo to plan an arm trajectory to execute.
- Also called
- GraspGen: A Diffusion-based Framework for 6-DOF Grasping with On-Generator Training
- Related
- Grasp Pose Detection · Diffusion Model · Contact-GraspNet · AnyGrasp · Grasp Planning · cuRobo (NVIDIA GPU-accelerated motion planning)
- Sources
- GraspGen (arXiv 2507.13097)
GraspGen 项目主页 (Chinese) - As of
- 2025-07