Parameter-Efficient Fine-Tuning
参数高效微调PEFTAdvancedFreezing most of a large model's parameters and training only a small added or selected subset for a new task.
Parameter-efficient fine-tuning is an umbrella term for fine-tuning methods that keep a pretrained model's parameters mostly frozen and train only a small subset to adapt it to a new task. Four approaches are common: inserting small modules (adapters); prepending learnable vectors to the input (prompt tuning); representing weight changes with low-rank matrices (LoRA); and unfreezing only a few of the original parameters, such as just the biases. It addresses the high cost of full fine-tuning: memory, compute, and the storage cost of keeping a full copy of the weights per task all drop sharply, while performance often comes close to full fine-tuning. Hugging Face's PEFT library bundles many of these methods together. When adapting an open-source VLA to one's own robot, LoRA is the most commonly used choice.
ExampleIn the OpenVLA paper, using LoRA to update only about 1.4% of the parameters, trained for 10 to 15 hours on a single A100, reached a 68.2% task success rate, close to full fine-tuning's 69.7%, which needed two GPUs and about 163 GB of memory.
- Also called
- PEFT
- Related
- LoRA · Adapter · Prompt Tuning / Soft Prompt · Full Fine-Tuning · Backbone Freezing · Fine-tuning
- Sources
- Hugging Face PEFT 文档 (Chinese)
Han et al. 2024: Parameter-Efficient Fine-Tuning for Large Models: A Comprehensive Survey
Kim et al. 2024: OpenVLA: An Open-Source Vision-Language-Action Model