Embodied AI Glossary中文

ConceptGraphs

Advanced

Fuses multi-view images into an open-vocabulary 3D object scene graph for large-model reasoning and robot planning.

ConceptGraphs was proposed by a team from MIT, Université de Montréal, the University of Toronto, and others (lead author Qiao Gu), published at ICRA 2024. It uses a general-purpose segmentation model frame by frame to cut out object regions, extracts CLIP features for each region, projects them into 3D points using depth, and then merges regions belonging to the same object across viewpoints to obtain a set of 3D objects with semantic features attached. It then uses LLaVA to generate text descriptions for the objects and GPT-4 to infer spatial relationships between objects as edges, forming a 3D scene graph — nodes are objects, edges are relationships. The whole pipeline needs no 3D annotation and no model fine-tuning. The resulting graph can be converted to text for a large model to answer queries like “find a basketball”; the paper demonstrates object retrieval, relocalization, and navigation tasks that judge which obstacles can be pushed aside, on mobile robots like the Jackal.

ExampleA robot first drives a loop through a house to build the scene graph; when the user says “find something to prop up my phone,” a large model picks “book” from the object descriptions in the graph, and the robot navigates to that node’s 3D location.

Also called
ConceptGraphs: Open-Vocabulary 3D Scene Graphs for Perception and Planning
Related
3D Scene Graph · Open-vocabulary · Semantic Map · CLIP · LLM-based Task Planning · SayPlan
Sources
ConceptGraphs (arXiv 2309.16650)
ConceptGraphs 项目页 (Chinese)
As of
2024-05

See it in the full glossary →