Embodied AI Glossary中文

Random Seed and Reproducibility

随机种子与可复现性Common

Fixing the random number generator's starting point so runs can be repeated, and checking results across several seeds to rule out luck.

Many parts of deep learning involve randomness: network initialization, data shuffling, dropout, reinforcement learning's exploration noise, and a simulator's domain randomization. A random seed is the starting point for a random number generator; fixing the seed in Python, NumPy, and PyTorch lets the same code produce close to the same result, though PyTorch's own documentation notes exact reproducibility isn't guaranteed across versions or platforms. Reproducibility also has a second meaning: whether a conclusion can be trusted. Henderson and colleagues (2018) found that in deep reinforcement learning, changing only the random seed can shift results enough to flip which algorithm looks better. Because of this, results should be reported as a mean and confidence interval across multiple seeds; Agarwal and colleagues (2021) further recommend using the interquartile mean (IQM) and bootstrap confidence intervals.

ExampleIn PyTorch, fixing seeds looks like torch.manual_seed(0), np.random.seed(0), random.seed(0), plus setting torch.backends.cudnn.benchmark=False and torch.use_deterministic_algorithms(True); when comparing two reinforcement-learning algorithms, each is run once per seed across 5 different seeds, and the mean and confidence interval are reported.

Also called
Random Seed, Reproducibility, Multi-seed Experiments
Related
Hyperparameter · Ablation Study · Statistical Rigor in Policy Evaluation (Confidence Intervals / Sequential Testing / Multiple Seeds) · Benchmark · Training / Validation / Test Set
Sources
Deep Reinforcement Learning that Matters (Henderson et al., AAAI 2018)
PyTorch Docs: Reproducibility
Deep Reinforcement Learning at the Edge of the Statistical Precipice (NeurIPS 2021)

See it in the full glossary →