Embodied AI Glossary中文

InfoNCE Loss

InfoNCE 损失Advanced

A contrastive-learning loss that trains a model to pick out the one true positive among many candidates.

InfoNCE was introduced by Aaron van den Oord and colleagues at DeepMind in their 2018 Contrastive Predictive Coding (CPC) paper. It pairs an anchor sample with one positive (say, a different augmentation of the same image, or the next moment in a sequence) and several negatives, computes similarity scores, and applies softmax classification so the positive scores highest — which is really cross-entropy underneath. The paper shows that minimizing it is equivalent to maximizing a lower bound on mutual information (how much two variables share), and the bound gets tighter as the number of negatives grows. It is the core training objective behind contrastive methods such as CLIP and SimCLR, and robot visual representations like R3M use a similar time-contrastive version of it.

ExampleWhen CLIP trains on a batch of N image-text pairs, each image treats only its own caption as the positive and the other N−1 captions as negatives, computing InfoNCE in both the image-to-text and text-to-image directions.

Also called
Contrastive Loss, InfoNCE Objective
Related
Contrastive Learning · CLIP · Self-Supervised Learning · Time-Contrastive Networks · Cross-Entropy · Representation Learning
Sources
Representation Learning with Contrastive Predictive Coding (arXiv:1807.03748)
CPC 论文 HTML 版(第 2.3 节 InfoNCE Loss and Mutual Information Estimation) (Chinese)

See it in the full glossary →