Embodied AI Glossary中文

Supervised Fine-Tuning

监督微调SFTCommon

Continuing supervised training on a pretrained model using paired input–correct-answer data, to teach it a specific behavior.

Supervised fine-tuning takes a pretrained model and keeps training it on a curated, high-quality labeled dataset to teach it some specific behavior; the training objective is simply to make its output match the given answer as closely as possible (cross-entropy for language output, often L1 or mean squared error for continuous actions). The term became popular through OpenAI's InstructGPT (2022): first run SFT on human-written example answers, then train a reward model and run RLHF — the now-classic three-step recipe for post-training a large model. In robotics, SFT usually means fine-tuning a foundation model like a VLA on teleoperation demonstrations for a specific robot and set of tasks, which is essentially behavior cloning. It's simple and stable, but it only imitates the demonstrations, so it tends to err on states the demonstrations never showed — which is why SFT is often followed by RL fine-tuning.

ExampleStanford's OpenVLA-OFT improves on how OpenVLA is fine-tuned — parallel decoding, action chunking, and continuous action output with L1 regression — raising the success rate on the LIBERO simulation benchmark from 76.5% to 97.1% while increasing action-generation throughput 26-fold.

Also called
SFT
Related
Fine-tuning · Behavior Cloning · Post-training · Instruction Tuning · Reinforcement Fine-Tuning (RL Fine-Tuning) · Reinforcement Learning from Human Feedback
Sources
Training language models to follow instructions with human feedback (InstructGPT, arXiv 2203.02155)
Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success (OpenVLA-OFT, arXiv 2502.19645)

See it in the full glossary →