WALL-A
自变量 WALL-AAdvancedX Square Robot's in-house, closed-source embodied manipulation model series, running end-to-end from perception to motor control.
WALL-A is a series of in-house manipulation foundation models built by Shenzhen-based embodied AI company X Square Robot (自变量机器人); it is not open-sourced. The company's website describes it as achieving 'unified intelligence across the full pipeline, from perception and understanding to action control' — in other words, an end-to-end vision-language-action approach. According to a January 2026 report from Zhidongxi, the WALL-A series combines a VLA with a world model, using the world model to predict how the environment's state changes over time and reasoning jointly with vision to decide on actions. The WALL-OSS paper, open-sourced in September 2025, describes WALL-A's architecture as a tightly coupled design with shared self-attention and separate feed-forward layers per modality, and WALL-OSS follows the same idea, making it viewable as an open-source version of the WALL-A approach. In April 2026 the company released a further 'world unified model' called WALL-B, and its later public demonstrations — housework, logistics sorting — have mostly been driven by WALL-B. WALL-A's initial release date and parameter count could not be verified from public sources for this entry.
ExampleAccording to Zhidongxi's report, X Square Robot's 'Quantum No. 1' used its models to demonstrate mobile manipulation across both outdoor and indoor settings: breaking down and recycling a takeout box outdoors, then making its own way through a building door and elevator to deliver it indoors.
- Also called
- WALL-A Manipulation Foundation Model, WALL-A series
- Related
- X Square Robot · WALL-OSS · Vision-Language-Action Model · World Model · Embodied Foundation Model · Mixture of Experts
- Sources
- 自变量机器人官网 (Chinese)
字节、阿里、美团首次在具身智能「同框」,十亿级融资背后,自变量到底凭什么?(智东西,2026-01-12) (Chinese)
Igniting VLMs toward the Embodied Space (WALL-OSS, arXiv 2509.11766) - As of
- 2026-09