Embodied AI Glossary中文

Data Leakage / Test-Set Contamination

数据泄漏 / 测试集污染Advanced

When information that should only appear at test time leaks into training, inflating evaluation scores beyond real-world performance.

Data leakage means a model was exposed during training to information it shouldn't have access to at test or deployment time, so it looks good on offline evaluation but performs much worse in real use. IBM divides this into two types: target leakage, where a feature secretly contains the “answer” that wouldn't be available at prediction time, and train-test contamination, where test data leaks into training, or where preprocessing steps like normalization are computed on the full dataset before splitting. In robot learning, common cases include splitting train and validation sets by individual frame rather than by whole trajectory, training on simulation benchmarks using scenes and initial states nearly identical to the test set, and letting evaluation questions leak into a large model's pretraining corpus. LIBERO-PRO found that VLA (vision-language-action) models scoring over 90% success on the original LIBERO benchmark could drop to 0% once objects, initial positions, or instructions were changed — evidence that they had mostly memorized the training set. The fix is to split data by trajectory or scene, and evaluate under out-of-distribution settings.

ExampleIf all the frames from a 30-second demonstration are shuffled randomly before being split into train and validation sets, the action error measured on the validation set will look very low, because the model has essentially seen the neighboring frames of nearly every validation frame; splitting by whole trajectory instead gives an error that reflects real generalization ability.

Also called
Train-Test Contamination, Test-Set Leakage
Related
Training / Validation / Test Set · Overfitting · Out-of-Distribution · LIBERO-PRO · Benchmark Saturation · Leaderboard Chasing
Sources
What is Data Leakage in Machine Learning?(IBM)
LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization (arXiv 2510.03827)
As of
2026-05

See it in the full glossary →