Catastrophic Forgetting
灾难性遗忘CommonA neural network's sharp loss, or outright loss, of old abilities after it learns something new.
Catastrophic forgetting, also called catastrophic interference, was first systematically reported by McCloskey and Cohen in 1989, when a backpropagation network trained on one set of addition problems and then another lost much of its performance on the first. Learning new material changes the same weights that encoded old knowledge, so performance on the old task drops sharply. Mitigations include replaying old data, elastic weight consolidation (EWC, which protects weights that matter most for old tasks), freezing part of the network, and parameter-efficient fine-tuning. In embodied AI, converting a pretrained VLM into a VLA and fine-tuning it only on robot data often erodes the model's original language understanding and general knowledge. That's why models like π0.5 co-train on a mix of web image-text data and robot data, and why Physical Intelligence's knowledge insulation approach blocks the action expert's gradients from flowing back into the VLM.
ExamplePhysical Intelligence's knowledge insulation paper found that directly attaching a continuous action expert to a VLM and training the whole thing together noticeably erodes the model's pretrained knowledge and weakens its ability to understand language instructions; they mitigate this with gradient blocking plus co-training on image-text QA data.
- Also called
- Catastrophic Interference
- Related
- Continual Learning · Backbone Freezing · Co-training · Knowledge Insulation · Stop-Gradient · Parameter-Efficient Fine-Tuning
- Sources
- Wikipedia: Catastrophic interference
Knowledge Insulating Vision-Language-Action Models (arXiv 2505.23705)
π0.5: a Vision-Language-Action Model with Open-World Generalization (arXiv 2504.16054) - As of
- 2025-05