Embodied AI Glossary中文

LEO (BIGAI)

LEO(3D 具身通才智能体)LEOAdvanced

A 3D embodied generalist model from BIGAI that understands 3D scenes and can answer questions, navigate, and manipulate objects.

LEO was proposed by the Beijing Institute for General Artificial Intelligence (BIGAI) together with Peking University, Carnegie Mellon University, and Tsinghua University, released in November 2023 and published at ICML 2024. At the time, most multimodal large models only handled 2D images, and struggled with tasks defined in a 3D scene. LEO concatenates first-person images, object-centric 3D tokens (each object's point cloud encoded by PointNet++, with a spatial Transformer then modeling relationships between objects), and a text instruction into a single sequence, fed into a Vicuna-7B fine-tuned with LoRA, with actions also output as discrete tokens. Training has two stages: first 3D vision-language alignment, then 3D vision-language-action instruction tuning, with the data generated with the help of a large model. It can do 3D captioning, question answering, embodied reasoning, navigation, and manipulation, and is an early representative example of a 3D VLA.

ExampleGiven a 3D scan of a room, and asked “what's on the table next to the sofa,” LEO answers in text; for object navigation, it looks at first-person images and outputs actions like moving forward or turning left, step by step, to find the target.

Also called
An Embodied Generalist Agent in 3D World
Related
3D VLA · 3D-LLM · Beijing Institute for General Artificial Intelligence · Embodied Agent · Object-centric Representation · LoRA
Sources
An Embodied Generalist Agent in 3D World (arXiv 2311.12871)
LEO 项目页 (Chinese)
As of
2024-07

See it in the full glossary →