Embodied AI Glossary中文

Feature-wise Linear Modulation

FiLM 特征调制FiLMAdvanced

Using conditioning information to compute a per-channel scale and shift that modulates a network's intermediate features.

FiLM was proposed by Ethan Perez, Aaron Courville, and colleagues (AAAI 2018) as a general-purpose layer for injecting conditioning information into a neural network. The idea is simple: a small network computes a scale γ and a shift β for each feature channel from the condition (such as a language-instruction vector), and the intermediate feature is transformed into γ·x + β. The original paper roughly halved the best error rate at the time on the CLEVR visual-reasoning benchmark. It's common in robotics: RT-1 uses FiLM to inject the language instruction into its pretrained EfficientNet image encoder, zero-initializing the layers that produce γ and β so FiLM starts out as an identity transform and doesn't disturb the pretrained weights; the CNN version of Diffusion Policy also uses FiLM at every convolutional layer to inject observation features.

ExampleIn RT-1, the instruction 'pick up the coke can' is first turned into a vector by the Universal Sentence Encoder, then modulates EfficientNet's layer features through FiLM, so the same image produces different visual features under different instructions.

Also called
FiLM, FiLM Layer, FiLM Conditioning
Related
Adaptive Layer Normalization · Language-conditioned Policy · RT-1 · Diffusion Policy · EfficientNet · Cross-Attention
Sources
FiLM: Visual Reasoning with a General Conditioning Layer (Perez et al., AAAI 2018)
RT-1: Robotics Transformer for Real-World Control at Scale (arXiv 2212.06817)
Diffusion Policy: Visuomotor Policy Learning via Action Diffusion (arXiv 2303.04137)

See it in the full glossary →