Embodied AI Glossary中文

GRAPE

Advanced

A method that uses preference alignment on successful and failed trajectories to improve a VLA's generalization to new tasks.

GRAPE was proposed in November 2024 by researchers at UNC Chapel Hill, the University of Washington, and the University of Chicago, with experiments built on OpenVLA as the base model. Most VLA models only do behavior cloning (learning by copying demonstrations) on successful demonstrations; having never seen failure, they don't know which behaviors to avoid, and tend to make mistakes when generalizing to new tasks. GRAPE borrows the idea of preference alignment from large language models: it has the model execute a task many times, ranks the resulting trajectories from best to worst by score, and then uses trajectory-wise preference optimization (TPO) to bias the model toward the better trajectories — so even failed trajectories provide useful information. For scoring, a complex task is first broken into stages, and a vision-language model generates spatiotemporal constraints for each stage to score against; swapping in a different set of constraints lets alignment target different goals, such as “safer” or “more efficient.”

ExampleThe authors report that GRAPE improves success rate by 51.79% on in-distribution tasks and 58.20% on unseen tasks; when aligned toward safety and efficiency goals, collision rate drops 37.44% and the number of execution steps falls 11.15%.

Also called
Generalizing Robot Policy via Preference Alignment
Related
Direct Preference Optimization · Vision-Language-Action Model · OpenVLA · Behavior Cloning · Generalization · Reinforcement Fine-Tuning (RL Fine-Tuning)
Sources
GRAPE: Generalizing Robot Policy via Preference Alignment (arXiv 2411.19309)
GRAPE 项目页 (Chinese)
As of
2025-02

See it in the full glossary →