Embodied AI Glossary中文

Scene Understanding

场景理解Common

Working out what's in an environment, where it is, and how things relate to each other from images, depth, or point clouds.

Scene understanding is a goal rather than a single algorithm: figuring out, from an image, depth, or point cloud, what's in an environment, where it is, and how things relate to each other. That covers geometry (where objects are, how big they are, where's walkable), semantics (what each thing is), and relations (a cup is on the table, a drawer can be pulled open). It's assembled from object detection, semantic and instance segmentation, depth estimation, 3D reconstruction, pose estimation, and affordance detection. Because robots need to navigate and manipulate in 3D space, embodied AI cares especially about 3D scene understanding: ScanNet (2017) provides 1,513 indoor scenes and about 2.5 million semantically annotated RGB-D frames, and ConceptGraphs (2023) fuses results from 2D foundation models across multiple views into open-vocabulary 3D scene graphs that a large language model can use for instruction-based planning.

ExampleEntering a kitchen, a household robot first builds a 3D scene graph: a fridge, a dining table, two bowls on the table, with the bowls sitting on the table. Given the instruction ‘put the bowls in the sink,’ it looks up the bowls' and sink's 3D positions from this graph and plans accordingly.

Also called
3D Scene Understanding
Related
3D Scene Graph · Semantic Map · 3D Vision · Semantic Segmentation · ConceptGraphs · ScanNet
Sources
ScanNet: Richly-annotated 3D Reconstructions of Indoor Scenes (arXiv 1702.04405)
ConceptGraphs: Open-Vocabulary 3D Scene Graphs for Perception and Planning (arXiv 2309.16650)

See it in the full glossary →