Embodied AI Glossary中文

Octo

Common

A 2024 open-source generalist robot policy pretrained on 800,000 cross-embodiment trajectories, quick to fine-tune to new robots.

Octo was released in May 2024 by the Octo Model Team, a group of researchers from UC Berkeley, Stanford, CMU, and Google DeepMind, and published at RSS 2024. It's a Transformer-based generalist policy pretrained on about 800,000 trajectories from 25 sub-datasets of Open X-Embodiment (a cross-embodiment robot dataset pooled from many institutions), released in two sizes: Octo-Small (27 million parameters) and Octo-Base (93 million parameters). A task can be specified either with a language instruction or with a single goal image; the Transformer backbone reads in task and observation tokens, and a diffusion action head outputs continuous actions, able to represent multimodal action distributions. Its emphasis is on being open and easy to modify: code, weights, and the training pipeline are all public, the input and output modules are modular, and it can be fine-tuned to a new sensor setup and action space on a consumer GPU in a few hours. It beat the next-best baseline by an average of 52% across 6 fine-tuning evaluation scenarios, and later work such as OpenVLA often uses it as a comparison baseline.

ExampleIn the “Berkeley Bimanual” task, Octo — pretrained only on single-arm data — was fine-tuned onto an ALOHA bimanual platform made of two ViperX arms: the right arm picks up a marker from the table while the left arm removes its cap, with the action space switched to joint-position control.

Also called
Octo: An Open-Source Generalist Robot Policy, Octo-Small, Octo-Base
Related
Open X-Embodiment · Generalist Policy · Diffusion Action Head · Cross-Embodiment · OpenVLA · RT-X
Sources
Octo: An Open-Source Generalist Robot Policy (arXiv 2405.12213)
Octo 项目主页 (Chinese)
RSS 2024 论文页(Robotics: Science and Systems XX, p090) (Chinese)
As of
2024-07

See it in the full glossary →