NVIDIA TensorRT
TensorRTTRTCommonNVIDIA's inference-acceleration library that compiles a trained model into a faster GPU runtime engine.
TensorRT is NVIDIA's deep-learning inference optimizer and runtime. It reads in a model in a format such as ONNX, performs layer fusion, picks the fastest available GPU kernels, and lowers numerical precision (FP16, INT8, FP8, and so on) to produce an inference “engine” file compiled for a specific GPU. A robot policy has to produce an action within tens of milliseconds, and running it directly in PyTorch is often too slow, so compiling with TensorRT typically cuts latency noticeably — making it a standard step for both Jetson and server-side deployment. Note that an engine is tied to a specific GPU model and TensorRT version, so switching hardware requires recompiling; large language models have their own dedicated variant, TensorRT-LLM.
ExampleExport a trained diffusion policy to ONNX, then compile it into an FP16 engine with trtexec to cut per-step inference latency on a Jetson Orin.
- Also called
- TRT
- Related
- Open Neural Network Exchange (ONNX) · NVIDIA TensorRT-LLM · Inference Latency · Post-Training Quantization · NVIDIA JetPack SDK · Inference Deployment
- Sources
- NVIDIA TensorRT
- As of
- 2025-09