Embodied AI Glossary中文

Training / Validation / Test Set

训练集 / 验证集 / 测试集Essential

Splitting a dataset into three parts: one to learn from, one to tune and select a model, and one saved for a final, honest check.

This is the standard way machine learning splits its data. The training set is used to update the model's parameters. The validation set evaluates the model during development, guiding the choice of hyperparameters and which checkpoint (a saved snapshot of the model's weights partway through training) to keep. The test set is touched only at the very end, to report how the model performs on data it has truly never seen. Google's Machine Learning Crash Course uses a 70/15/15 split as an example, but the exact ratio isn't fixed; what matters is that the validation and test sets are large enough to give reliable conclusions. The three sets must not share samples — overlap is equivalent to peeking at the answer key and inflates the reported score — and a validation set that gets reused for many rounds of decisions can gradually lose its value as an honest check. In robotics, real-robot or simulation evaluation often plays the role of the test set, and it's just as important that the test scenes and objects never appeared in training.

ExampleTo train a grasping policy: train on data collected in 8 kitchens, use data from a 9th kitchen as validation to decide how long to train, then run real-robot trials in a never-before-seen 10th kitchen and report that success rate.

Also called
Train/Val/Test Split, Training Set, Validation Set, Test Set, Held-out Set
Related
Overfitting · Hyperparameter · Checkpoint · Generalization · Out-of-Distribution · Data Leakage / Test-Set Contamination
Sources
Google Machine Learning Crash Course: Dividing the original dataset
Google Machine Learning Glossary: validation set

See it in the full glossary →