Embodied AI Glossary中文

Time-Contrastive Networks

时间对比学习TCNAdvanced

Self-supervised video representation learning: pull same-moment multi-view frames together, push nearby-but-different moments apart.

Time-Contrastive Networks were proposed by Pierre Sermanet and colleagues at Google Brain in 2017 (ICRA 2018), an early landmark in learning visual representations from unlabeled video. The multi-view version uses several cameras filming the same process at once: frames from different viewpoints at the same moment are pulled together in feature space, while frames from the same viewpoint that look similar but sit at different points in the task are pushed apart, trained with a triplet loss. The resulting features are insensitive to viewpoint and lighting, yet can still distinguish task-relevant state, such as how far a cup has tilted. The paper used it to let a robot imitate pouring water and human poses from a single third-person human video, treating feature distance to the demonstration as a reinforcement-learning reward. Later robot visual pretraining such as R3M also used time-contrastive learning as one of its objectives.

ExampleA robot watches a third-person video of a person pouring water, uses the distance between its own footage and the demonstration in TCN feature space as a reward, and learns the pouring motion through reinforcement learning.

Also called
TCN, Time-Contrastive Learning
Related
Contrastive Learning · Self-Supervised Learning · R3M · Representation Learning · Imitation from Observation · Pre-trained Visual Representation
Sources
Sermanet et al. 2017: Time-Contrastive Networks: Self-Supervised Learning from Video
项目页:Time-Contrastive Networks (Chinese)
Nair et al. 2022: R3M: A Universal Visual Representation for Robot Manipulation

See it in the full glossary →