Model Merging
模型合并AdvancedCombining several related models' parameters by weighted averaging, with no extra training required.
Model merging combines multiple models that share the same architecture — usually all fine-tuned from the same pretrained model — directly at the parameter level, needing no extra training and adding no inference cost. The simplest version is linear interpolation: new parameters equal (1−α) times the pretrained parameters plus α times the fine-tuned parameters. Notable examples include WiSE-FT, which interpolates a zero-shot model with a fine-tuned one to preserve robustness; Model Soups, which averages the results of fine-tuning runs with different hyperparameters; and task arithmetic, which treats “fine-tuned minus pretrained” as a task vector that can be added or subtracted. In robotics it is used to fight the catastrophic forgetting caused by fine-tuning: Sergey Levine's team's 2025 RETAIN interpolates a fine-tuned π0-FAST-DROID with the original model, keeping both the new skill and the model's original generalization. Other work, however, has found that directly merging VLAs each fine-tuned on a different task can push success rate close to zero.
ExampleRETAIN fine-tunes π0-FAST-DROID on a “wipe the whiteboard with an eraser” task using a few dozen demonstrations, then interpolates it with the original model at an α between 0.25 and 0.75; the merged model clearly beats the plain fine-tuned version on out-of-distribution variations such as new objects and new camera angles.
- Also called
- Weight Interpolation, Weight Averaging, Model Soups, Task Arithmetic
- Related
- Catastrophic Forgetting · Fine-tuning · Continual Learning · LoRA · Exponential Moving Average · Checkpoint
- Sources
- Yadav et al. 2025: Robust Finetuning of Vision-Language-Action Robot Policies via Parameter Merging (RETAIN)
Wortsman et al. 2022: Model soups
Ilharco et al. 2022: Editing Models with Task Arithmetic - As of
- 2025-12