Adapter
适配器AdvancedSmall modules inserted between a frozen large model's layers, with only those small modules trained to adapt it to a new task.
Adapters are a family of parameter-efficient fine-tuning methods, introduced by Houlsby and colleagues in the 2019 paper “Parameter-Efficient Transfer Learning for NLP”: a small bottleneck network — project down, then back up, with a residual connection — is inserted into each Transformer layer, and fine-tuning freezes the original model and trains only these adapters. On 26 text-classification tasks with BERT, adding just 3.6% extra parameters per task matched full fine-tuning within 0.4%. Adapters save memory and storage, let one base model carry adapters for many tasks at once, and reduce catastrophic forgetting; LoRA, introduced later, continues the same basic idea. In robotics, TAIL (ICLR 2024) compares bottleneck adapters, P-Tuning, and LoRA for imitation learning, and this family of methods is commonly used to adapt a VLA to a new task or new robot.
ExampleTAIL fine-tunes a pretrained decision-making model on a small number of demonstrations and finds that LoRA, training only about 1% of the parameters, matches full fine-tuning's performance while avoiding catastrophic forgetting during continual learning.
- Also called
- Adapter Tuning, Bottleneck Adapter
- Related
- Parameter-Efficient Fine-Tuning · LoRA · Backbone Freezing · Catastrophic Forgetting · Full Fine-Tuning · Prompt Tuning / Soft Prompt
- Sources
- Parameter-Efficient Transfer Learning for NLP (Houlsby et al., arXiv 1902.00751)
TAIL: Task-specific Adapters for Imitation Learning with Large Pretrained Models (arXiv 2310.05905)