Instruction Tuning
指令微调AdvancedFine-tuning a model on many tasks written as instruction-and-answer pairs so it learns to follow instructions.
Instruction tuning is supervised fine-tuning of a pretrained model on a large set of tasks phrased as natural-language instructions paired with the expected output. Google's 2021 FLAN work instruction-tuned a 137-billion-parameter model across more than 60 NLP tasks, and its zero-shot performance beat GPT-3 on 20 of 25 evaluation tasks. It turns a model from one that merely continues text into one that follows requests, making it a basic building block of chat assistants. In multimodal models, LLaVA (2023) used GPT-4 to generate image-text instruction data for “visual instruction tuning”; many VLA backbones go through a comparable step, and this kind of data is often mixed into VLA training to reduce forgetting.
ExampleLLaVA (2023) used GPT-4 to rewrite image captions into “look at the image and answer the question”-style instruction data, then fine-tuned a vision-encoder-plus-language-model combination on it to produce an assistant that can converse about images.
- Also called
- Instruction Fine-tuning
- Related
- Supervised Fine-Tuning · Large Language Model · Vision-Language Model · LLaVA · Zero-shot · Post-training
- Sources
- Finetuned Language Models Are Zero-Shot Learners (FLAN, arXiv:2109.01652)
Visual Instruction Tuning (LLaVA, arXiv:2304.08485)