Embodied AI Glossary中文

torch.compile

Advanced

The one-line model-compilation acceleration API that PyTorch has offered since version 2.0.

torch.compile is a compilation feature introduced in PyTorch 2.0 (2023). Normally, PyTorch executes one operator at a time, eagerly — every step carries Python dispatch overhead, and there's no way to optimize across operators. torch.compile uses TorchDynamo to capture the computation graph inside a piece of Python code at runtime, then hands it to the default backend, TorchInductor, which fuses operators together and generates Triton or C++ kernels (a kernel being the low-level function that actually executes on the GPU or CPU). Using it usually just means wrapping the model in one call; the first call takes time to compile, and training or inference is faster after that. In embodied AI it's commonly used to cut per-step inference latency for large models like VLAs, and can be paired with CUDA Graphs to further reduce kernel-launch overhead.

ExampleWhen deploying a diffusion policy, write policy = torch.compile(policy, mode=“max-autotune”); after a few warm-up calls, the time spent per denoising step drops.

Also called
PyTorch 2 compilation
Related
PyTorch · Inference Latency · CUDA Graphs · Operator / Kernel · NVIDIA TensorRT · Inference Deployment
Sources
torch.compile — PyTorch documentation
PyTorch 2.x overview

See it in the full glossary →