Embodied AI Glossary中文

Reinforcement Learning from Human Feedback

基于人类反馈的强化学习RLHFCommon

Having people compare pairs of model outputs to train a reward model, then using reinforcement learning to optimize toward human preference.

RLHF was introduced by Christiano and colleagues (OpenAI and DeepMind) in 2017: a reward function can be learned just from people comparing which of two trajectories is better, needing human feedback on less than 1% of the total interactions. OpenAI's InstructGPT (2022) applied it to large language models in three steps: supervised fine-tuning; training a reward model on human rankings of multiple candidate answers; and optimizing the model with PPO, adding a KL penalty to keep it from drifting too far from the original model — a recipe ChatGPT continued to use. It suits goals that are hard to write as a formula but easy for a person to judge at a glance; the downsides are that labeling is expensive and the model can learn to exploit the reward model's blind spots. DPO skips the reward model and trains directly on preference data instead; in embodied AI, GRAPE aligns a VLA using trajectory-level preferences.

ExampleIn InstructGPT's human evaluations, a version with only 1.3 billion parameters that went through RLHF was preferred over the 175-billion-parameter GPT-3.

Also called
RLHF, Preference Alignment
Related
Reward Model · Direct Preference Optimization · Proximal Policy Optimization · KL Regularization · Reward Hacking · GRAPE
Sources
Deep Reinforcement Learning from Human Preferences (Christiano et al., 2017)
Training language models to follow instructions with human feedback (InstructGPT)
Wikipedia: Reinforcement learning from human feedback

See it in the full glossary →