Embodied AI Glossary中文

OmniManip

Advanced

A zero-shot manipulation framework that turns a vision-language model's reasoning into point-and-direction constraints defined in each object's own coordinate frame.

OmniManip comes from Hao Dong's lab at Peking University working with AgiBot (the PKU-AgiBot joint lab), released in January 2025 and selected as a CVPR 2025 Highlight paper. Vision-language models (VLMs) have broad commonsense knowledge but cannot reliably output precise 3D positions and orientations. OmniManip first places each object into its own canonical space — a coordinate frame aligned to the object's function rather than just its shape — and defines 'interaction primitives' inside it, such as an interaction point and an interaction direction. These primitives become spatial constraints that the VLM selects and checks. During execution, the system tracks each object's 6D pose (3D position plus 3D orientation) in real time and updates the trajectory accordingly, so both planning and execution run as closed loops. None of this requires fine-tuning the VLM, yet the method generalizes to many manipulation tasks in a zero-shot setting. It belongs to the same 'VLM plus intermediate representation' family as ReKep, VoxPoser, and CoPa; the authors also note it can be used to auto-generate simulation data.

ExampleTake 'pour tea into a cup': the VLM first recognizes the teapot and cup, picks the spout point and pouring direction in the teapot's canonical space and the rim point on the cup, and uses these to compute the end-effector pose; during execution it keeps tracking both objects' 6D poses and corrects the trajectory.

Also called
OmniManip: Towards General Robotic Manipulation via Object-Centric Interaction Primitives as Spatial Constraints
Related
ReKep · VoxPoser · CoPa · Affordance · Intermediate Representation · 6D Object Pose Estimation
Sources
arXiv 2501.03841: OmniManip
OmniManip 项目主页 (Chinese)
As of
2025-06

See it in the full glossary →