Embodied AI Glossary中文

Matterport3D

Matterport3D 数据集MP3DAdvanced

A dataset of RGB-D indoor scans from 90 real buildings, a standard scene library for indoor navigation research.

Matterport3D (abbreviated MP3D) was published at the 3DV conference in 2017 by researchers from Princeton University, Stanford University, and other institutions, built from data captured with Matterport's 3D scanning cameras. It covers 90 building-scale real indoor scenes, with 194,400 RGB-D images (color plus depth) stitched into 10,800 panoramas, along with surface reconstruction meshes, camera poses, and 2D/3D semantic segmentation labels. Its significance is letting researchers train and evaluate agents inside digital copies of real houses, without physically moving a robot into a real home every time: the vision-and-language navigation benchmark R2R is built on top of it, and simulation platforms such as Habitat use it as a standard set of scenes for tasks like point-goal and object-goal navigation. The later HM3D dataset expanded the scene count to 1,000. Access requires signing a terms-of-use agreement.

ExampleThe R2R benchmark annotates about 22,000 navigation instructions (averaging 29 words) across Matterport3D's 90 buildings; an agent must understand a description like “go through the kitchen and stop at the doorway by the stairs” and walk to the destination.

Also called
MP3D
Related
Habitat-Matterport 3D Dataset · Habitat · Room-to-Room · Vision-and-Language Navigation · Object-Goal Navigation · ScanNet
Sources
Matterport3D 项目主页 (Chinese)
Room-to-Room (R2R) 数据集主页 (Chinese)
As of
2017

See it in the full glossary →