Embodied AI Glossary中文

WALL-OSS

自变量 WALL-OSSAdvanced

X Square Robot's open-source embodied foundation model that turns a vision-language model into a VLA that outputs actions directly.

WALL-OSS is an embodied foundation model open-sourced by X Square Robot (自变量机器人) in September 2025, described in the paper 'Igniting VLMs toward the Embodied Space.' It uses Qwen2.5-VL-3B as its backbone, with a tightly coupled mixture-of-experts structure (different feed-forward layers assigned to different objectives depending on the training task) that connects 'instruction to reasoning to sub-task planning to continuous action' inside one differentiable model. Training first grounds the model with embodied question-answering and FAST discrete-action data, then teaches continuous actions using flow matching, on more than 10,000 hours of data. Two versions were released, FLOW and FAST, under the Apache-2.0 license. In May 2026 the company released WALL-OSS 0.5 (about 4 billion parameters), built to deploy directly from its pretrained weights with no fine-tuning required.

ExampleLoading the WALL-OSS-FLOW weights through the official wall-x codebase, a researcher fine-tunes it on their own LeRobot-format data for a tabletop tidying task.

Also called
WALL-OSS 0.5, WALL-OSS-FLOW, WALL-OSS-FAST
Related
X Square Robot · WALL-A · Vision-Language-Action Model · Embodied Chain-of-Thought · Flow Matching · Qwen-VL
Sources
Igniting VLMs toward the Embodied Space (arXiv:2509.11766)
X-Square-Robot/wall-x (GitHub)
x-square-robot/wall-oss-0.5 (Hugging Face)
As of
2026-06

See it in the full glossary →