SpatialLM
SpatialLM(群核空间大模型)AdvancedManycore's open-source 3D large model that reads an indoor point cloud and outputs a structured layout of walls, doors, windows, and furniture.
SpatialLM was open-sourced by Hangzhou-based Manycore Tech (owner of Kujiale) in March 2025, with a technical report in June, and was selected for NeurIPS 2025. It takes as input a point cloud of an indoor scene — which can come from phone-video reconstruction, an RGB-D camera, or LiDAR — and outputs a structured scene description: the positions of walls, doors, and windows, and 3D boxes for furniture with category and orientation. Rather than designing a separate network for each task, as earlier work did, it reuses a standard multimodal large-model structure: a point-cloud encoder turns geometric information into tokens, and a small open-source language model such as Llama 1B or Qwen 0.5B then writes out the scene item by item as text. Training used point clouds and annotations from 12,328 synthetic indoor scenes (54,778 rooms). Version 1.1, released in June, switched to the Sonata point-cloud encoder and added detection restricted to user-specified categories. SpatialLM can supply structured spatial information for robot navigation and indoor layout understanding.
ExampleA phone is used to film a walkthrough of a room, which is reconstructed into a point cloud via MASt3R-SLAM and fed into SpatialLM, yielding each wall's endpoints, the positions of doors and windows, and 3D boxes for the bed and nightstand.
- Also called
- SpatialLM 1.1, Manycore SpatialLM, SpatialLM: Training Large Language Models for Structured Indoor Modeling
- Related
- Point Cloud · 3D Object Detection · Scene Understanding · Manycore Tech · Spatial Intelligence · Multimodal Large Language Model
- Sources
- SpatialLM: Training Large Language Models for Structured Indoor Modeling (arXiv 2506.07491)
manycore-research/SpatialLM (GitHub) - As of
- 2025-09