Embodied AI Glossary中文

Rho-alpha

微软 Rho-alphaραAdvanced

Microsoft's first robotics model, built from its Phi vision-language models and adding touch sensing.

Rho-alpha (written ρα) was released by Microsoft Research on January 21, 2026, Microsoft's first robotics model, derived from its Phi family of vision-language models. Microsoft calls it 'VLA+': on top of the usual vision-language-action model setup — looking at the scene, listening to an instruction, and outputting actions directly — it adds touch to perception, with force sensing still under development, and is exploring how a robot can keep adapting during use based on a person's corrective feedback. It converts natural-language instructions into control signals for two-arm manipulation. Training data includes real-robot demonstrations, simulated tasks, and web-scale visual question-answering data. At launch it was evaluated on dual UR5e arms fitted with tactile sensors and on a humanoid robot, using the BusyBox benchmark. It launched as an early-access research program, with Microsoft saying it would later become available on Microsoft Foundry.

ExampleAn operator gives a single natural-language instruction, and a pair of UR5e arms fitted with tactile sensors carries out a two-handed manipulation task, using both vision and touch to judge whether the action has succeeded.

Also called
ρα, Rho-Alpha, Microsoft Research robotics model derived from Phi
Related
Vision-Language-Action Model · Vision-Tactile-Language-Action Model · Tactile Sensor · Bimanual Manipulation · Human-in-the-Loop · Microsoft Research
Sources
Microsoft Research: Advancing AI for the physical world
As of
2026-01

See it in the full glossary →