Embodied AI Glossary中文

CUDA Graphs

CUDA GraphAdvanced

Recording a sequence of GPU operations as a single graph, so it can be replayed in one submission with less launch overhead.

CUDA Graphs is an execution model in CUDA: a sequence of kernel launches, memory copies, and their dependencies within a computation are first recorded as a graph, and afterward only that graph needs to be submitted each time, letting the driver schedule the whole thing in one go. It addresses CPU-side launch overhead — when a model has many small operators, the CPU time spent launching each kernel individually can exceed the time the GPU actually spends computing, becoming the bottleneck. This is especially useful for robot-policy inference: vision-language-action models and diffusion policies repeatedly run a forward pass with a fixed structure within tens of milliseconds, with unchanging tensor shapes, which is exactly the situation graph capture is suited for. PyTorch exposes this capability through graph capture and torch.compile's low-overhead mode.

ExampleCapture a diffusion action head's multi-step denoising loop as a CUDA Graph, reducing kernel-launch overhead within a single frame of inference.

Also called
CUDA graph
Related
CUDA · Inference Latency · torch.compile · NVIDIA TensorRT · Operator / Kernel · Policy Inference Frequency
Sources
Getting Started with CUDA Graphs (NVIDIA Technical Blog)
CUDA semantics - PyTorch Documentation

See it in the full glossary →