Embodied AI Glossary中文

OXE Magic Soup

Magic Soup(OXE 数据配方)Advanced

The recipe Octo and OpenVLA used to select and weight subsets of Open X-Embodiment for pretraining.

Magic Soup is the name for a particular robot-data mixing recipe, taken from the Octo codebase's oxe_magic_soup configuration. The individual datasets inside Open X-Embodiment (OXE, a cross-embodiment robot dataset pooled from many institutions) vary hugely in size, quality, robot type, and camera setup, so mixing them together naively doesn't work well. Octo's (2024) training mixture drew on 25 of these datasets, about 800,000 trajectories in total, with weights roughly proportional to sample count, then doubled the weight of datasets with richer scenes and tasks and down-weighted ones with many repetitive clips. OpenVLA reused this same weighting, and its codebase also has an expanded version, Magic Soup++, which adds DROID for about 970,000 demonstrations total. It isn't new data — it's an empirically tuned recipe for which datasets to include and how much weight to give each; even the Octo paper admits it still lacks a systematic analysis of the choice.

ExampleIn OpenVLA's training mixture, Fractal (RT-1 data) and Kuka each make up about 12.7%, while UCSD Kitchen is under 0.1%; DROID was included at 10% weight but removed for the final third of training because the model was learning slowly from it.

Also called
oxe_magic_soup, Magic Soup++, oxe_magic_soup_plus
Related
Open X-Embodiment · Data Mixture · Octo · OpenVLA · DROID (Distributed Robot Interaction Dataset) · RLDS (Reinforcement Learning Datasets)
Sources
octo-models/octo: oxe_dataset_mixes.py (GitHub)
Octo: An Open-Source Generalist Robot Policy (arXiv 2405.12213)
OpenVLA: An Open-Source Vision-Language-Action Model (arXiv 2406.09246)
As of
2024-06

See it in the full glossary →