Rejection Sampling Fine-Tuning
拒绝采样微调 / 过滤式行为克隆AdvancedLetting a model attempt a task many times, keeping only the successful or high-scoring results, and training on those.
Rejection sampling fine-tuning and filtered behavior cloning describe the same basic idea: let an existing model generate multiple results for a task, use answer-checking, success detection, or return ranking to throw out the bad ones, and use only the good samples as supervised fine-tuning or behavior-cloning data. In large-model work, Yuan and colleagues' 2023 paper is often cited, using it to expand math-reasoning training data; the 2021 Decision Transformer paper also used “percentile behavior cloning” (%BC), cloning only the top X% of data by return, as a baseline. It is simple to implement and trains stably, though information in the failed samples is simply thrown away. In robotics it is commonly used to keep only successful autonomous trajectories, or, as with ByteDance's GR-RL, to use a learned task-progress function to filter out demonstration segments that made no progress on the task. Note that the abbreviation RFT is also commonly used for RL fine-tuning.
ExampleYuan and colleagues had multiple models answer GSM8K math problems repeatedly, keeping only the reasoning traces that reached the correct answer for the training set; LLaMA-7B's accuracy rose from 35.9% with plain supervised fine-tuning to 49.3%.
- Also called
- RFT, Filtered Behavior Cloning, Filtered BC, Percentile Behavior Cloning (%BC)
- Related
- Behavior Cloning · Supervised Fine-Tuning · Self-improvement · Best-of-N Sampling · Suboptimal (Noisy) Demonstrations · Advantage Conditioning
- Sources
- Yuan et al. 2023: Scaling Relationship on Learning Mathematical Reasoning with Large Language Models
Chen et al. 2021: Decision Transformer: Reinforcement Learning via Sequence Modeling
Li et al. 2025: GR-RL: Going Dexterous and Precise for Long-Horizon Robotic Manipulation - As of
- 2025-12