Embodied AI Glossary中文

Fréchet Inception Distance

弗雷歇初始距离FIDCommon

A metric for how close a batch of generated images is to real images in overall distribution; lower is better.

FID was introduced by Heusel and colleagues in a NeurIPS 2017 paper to evaluate the image quality of GANs and other generative models. It works by feeding a batch of real images and a batch of generated images through a pretrained Inception v3 network (an image classifier), taking the 2048-dimensional features from its final pooling layer, fitting each set to a multivariate Gaussian distribution, and computing the Fréchet distance between the two Gaussians. Because it compares overall distributions rather than pixel by pixel, it captures both image quality and diversity at once; lower scores are better, and 0 means the two sets are statistically identical. Its limitations are that it depends on Inception's learned features, doesn't always track human judgment, and is biased when the sample size is small. The video version is called FVD. In embodied AI, FID, FVD, and PSNR are often reported together when evaluating the generated frames of world models and video-generation models.

ExampleTo evaluate a robot video world model, researchers take a few thousand frames of its generated manipulation footage and a few thousand real frames from the same scene, run both through Inception v3, and compute FID; a lower score means the generated footage looks more like real data overall.

Also called
FID, FID Score
Related
Fréchet Video Distance · Peak Signal-to-Noise Ratio / Structural Similarity Index / Learned Perceptual Image Patch Similarity · Video Generation Model · World-Model-based Policy Evaluation · Generative Adversarial Network · Generative Model
Sources
GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium (arXiv 1706.08500)
Fréchet inception distance - Wikipedia

See it in the full glossary →