Embodied AI Glossary中文

Overfitting

过拟合Essential

When a model memorizes the training data too closely, so its performance drops on new data it hasn't seen.

Overfitting means a model has fit the training data too tightly, memorizing even its noise and incidental details, so its predictions become unreliable on data it hasn't seen before. The classic sign is that training loss keeps falling while validation loss starts rising; it's most likely to happen when data is scarce, the model is large, or training runs for too many epochs. Common remedies include adding more data or using data augmentation, regularization, dropout, early stopping, and picking the checkpoint that performs best on a validation set. This problem is especially prominent in embodied AI: real-robot demonstrations often number only in the dozens to low hundreds, so a policy can end up memorizing the object placement, lighting, and background it saw during data collection — and fail once an object moves a few centimeters or the tablecloth changes — which is why papers specifically test positional and scene generalization. The opposite problem, where a model fails to fit even the training data well, is called underfitting.

ExampleFor instance, training a cup-grasping policy on just 50 demonstrations where the cup always sits in the same spot: after enough training, it succeeds almost every time at that original spot but grasps at empty air once the cup moves 10cm away — the policy memorized that one fixed trajectory instead of learning to “find the cup, then grasp it.”

Related
Underfitting · Generalization · Regularization · Early Stopping · Data Augmentation · Training / Validation / Test Set
Sources
Google Machine Learning Glossary: overfitting
Wikipedia: Overfitting

See it in the full glossary →