InfoNCE Loss
InfoNCE 损失AdvancedA contrastive-learning loss that trains a model to pick out the one true positive among many candidates.
InfoNCE was introduced by Aaron van den Oord and colleagues at DeepMind in their 2018 Contrastive Predictive Coding (CPC) paper. It pairs an anchor sample with one positive (say, a different augmentation of the same image, or the next moment in a sequence) and several negatives, computes similarity scores, and applies softmax classification so the positive scores highest — which is really cross-entropy underneath. The paper shows that minimizing it is equivalent to maximizing a lower bound on mutual information (how much two variables share), and the bound gets tighter as the number of negatives grows. It is the core training objective behind contrastive methods such as CLIP and SimCLR, and robot visual representations like R3M use a similar time-contrastive version of it.
ExampleWhen CLIP trains on a batch of N image-text pairs, each image treats only its own caption as the positive and the other N−1 captions as negatives, computing InfoNCE in both the image-to-text and text-to-image directions.
- Also called
- Contrastive Loss, InfoNCE Objective
- Related
- Contrastive Learning · CLIP · Self-Supervised Learning · Time-Contrastive Networks · Cross-Entropy · Representation Learning
- Sources
- Representation Learning with Contrastive Predictive Coding (arXiv:1807.03748)
CPC 论文 HTML 版(第 2.3 节 InfoNCE Loss and Mutual Information Estimation) (Chinese)