Embodied AI Glossary中文

GenManip

Advanced

A manipulation simulation and evaluation platform from Shanghai AI Lab, built on Isaac Sim, that uses a large model to auto-generate tasks.

GenManip was proposed by Shanghai AI Lab together with Zhejiang University, Xi'an Jiaotong University, Nanjing University, and others, published at CVPR 2025. It's a tabletop-manipulation simulation platform built on NVIDIA Isaac Sim, focused on whether a policy can understand a wide range of language instructions. It uses a large language model to generate task-oriented scene graphs (describing which objects are in the scene and what the target relationships are), paired with 10,000 annotated 3D object assets, to automatically synthesize a large amount of diverse tasks and demonstration data. Its evaluation component, GenManip-Bench, contains 200 hand-refined scenes and tests four kinds of generalization: spatial relationships, appearance understanding, common-sense reasoning, and long-horizon tasks. The paper compares two approaches: a modular system that uses foundation models for perception and planning, and an end-to-end policy trained with behavioral cloning, finding that the former generalizes better zero-shot while the latter improves as more data is added. The code is open-sourced on GitHub.

ExampleOn GenManip-Bench, the best-performing modular system, CoPa (paired with GPT-4.5), reaches an overall success rate of 23.0%; on long-horizon tasks, the evaluated models average only 9.07%.

Also called
LLM-driven Simulation for Generalizable Instruction-Following Manipulation, GenManip-Bench, GenManip Suite
Related
NVIDIA Isaac Sim · Instruction Following · Generative Simulation · Synthetic Data · CoPa · Shanghai Artificial Intelligence Laboratory
Sources
GenManip: LLM-driven Simulation for Generalizable Instruction-Following Manipulation (arXiv 2506.10966)
GenManip Suite project page
As of
2025-06

See it in the full glossary →