GenManip
AdvancedA manipulation simulation and evaluation platform from Shanghai AI Lab, built on Isaac Sim, that uses a large model to auto-generate tasks.
GenManip was proposed by Shanghai AI Lab together with Zhejiang University, Xi'an Jiaotong University, Nanjing University, and others, published at CVPR 2025. It's a tabletop-manipulation simulation platform built on NVIDIA Isaac Sim, focused on whether a policy can understand a wide range of language instructions. It uses a large language model to generate task-oriented scene graphs (describing which objects are in the scene and what the target relationships are), paired with 10,000 annotated 3D object assets, to automatically synthesize a large amount of diverse tasks and demonstration data. Its evaluation component, GenManip-Bench, contains 200 hand-refined scenes and tests four kinds of generalization: spatial relationships, appearance understanding, common-sense reasoning, and long-horizon tasks. The paper compares two approaches: a modular system that uses foundation models for perception and planning, and an end-to-end policy trained with behavioral cloning, finding that the former generalizes better zero-shot while the latter improves as more data is added. The code is open-sourced on GitHub.
ExampleOn GenManip-Bench, the best-performing modular system, CoPa (paired with GPT-4.5), reaches an overall success rate of 23.0%; on long-horizon tasks, the evaluated models average only 9.07%.
- Also called
- LLM-driven Simulation for Generalizable Instruction-Following Manipulation, GenManip-Bench, GenManip Suite
- Related
- NVIDIA Isaac Sim · Instruction Following · Generative Simulation · Synthetic Data · CoPa · Shanghai Artificial Intelligence Laboratory
- Sources
- GenManip: LLM-driven Simulation for Generalizable Instruction-Following Manipulation (arXiv 2506.10966)
GenManip Suite project page - As of
- 2025-06