Direct Preference Optimization
直接偏好优化DPOAdvancedAligning a model directly from paired “better/worse” samples, without training a reward model or running reinforcement learning.
DPO was introduced by Stanford's Rafailov, Finn, and colleagues in 2023. Traditional RLHF (reinforcement learning from human feedback) first trains a reward model on preference data, then runs reinforcement learning with an algorithm like PPO — a long, often unstable pipeline. DPO proves that the optimal policy for a KL-constrained reward-maximization problem has a closed-form solution, which turns the problem into a classification-style loss: raise the probability of the preferred sample relative to a reference model, and lower the probability of the rejected one, with no need to sample from the model during training. It was first used to align large language models, and embodied AI has borrowed it for VLA post-training: pairing successful and failed trajectories into preferences and optimizing the policy directly.
ExampleGRAPE (2024) extends DPO from single steps to whole trajectories (called TPO) on OpenVLA: it pairs successful and failed manipulation trajectories into preferences to fine-tune the model, using a step-wise DPO variant, OpenVLA-DPO, as a comparison baseline.
- Also called
- DPO
- Related
- Reinforcement Learning from Human Feedback · Reward Model · KL Regularization · GRAPE · Post-training · Group Relative Policy Optimization
- Sources
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model (arXiv:2305.18290)
GRAPE: Generalizing Robot Policy via Preference Alignment (arXiv:2411.19309)