Embodied AI Glossary中文

ChatVLA

Advanced

A unified VLA model that can both answer questions about images and directly control a robot.

ChatVLA was proposed in February 2025 by teams at Midea Group and East China Normal University, and was selected for an oral presentation at the EMNLP 2025 main conference. It targets a common problem in VLAs: fine-tuning on robot data tends to wipe out the base VLM's original image question-answering ability (the paper calls this spurious forgetting), while training control and understanding data together makes the two interfere with each other. Its fix has two parts. First, staged alignment training: the model is trained on robot data alone until it can control the robot, then image-text data is mixed back in at a 1:3 ratio to restore understanding. Second, a mixture-of-experts structure: attention layers are shared, but feed-forward layers split into an understanding expert and a control expert, routed by task. Built on Qwen2-VL-2B, it scores 47.2 on MMStar, with multimodal understanding clearly stronger than earlier VLAs, and it also beats OpenVLA and ECoT across 25 real-robot tasks.

ExampleThe same ChatVLA model can both answer questions about a photo and carry out instructions like grasp, place, push, or hang in scenes such as a bathroom, kitchen, or tabletop.

Also called
ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model
Related
Vision-Language-Action Model · Catastrophic Forgetting · Mixture of Experts · Co-training · DexVLA · Knowledge Insulation
Sources
ChatVLA (arXiv 2502.14420)
ChatVLA 项目主页 (Chinese)
As of
2025-11

See it in the full glossary →