Common Training and Inference GPUs (RTX 4090 / A100 / H100 / B200)
常用 GPU 型号(RTX 4090 / A100 / H100 / B200)CommonThe handful of NVIDIA GPUs most often used to train and run embodied-AI models, differing mainly in memory and compute.
These are the NVIDIA GPUs that show up most often in embodied-AI papers and code repositories. The RTX 4090 (Ada architecture, 24 GB of memory) is a consumer gaming card, often used for single-machine inference and small-scale fine-tuning. The A100 (Ampere architecture, 40 or 80 GB) and H100 (Hopper architecture, 80 GB, with FP8 support) are data-center cards, used in multi-GPU clusters for pretraining and full-parameter fine-tuning. The B200 (Blackwell architecture) has up to about 180 GB of memory per card, and an 8-GPU DGX B200 system has roughly 1.4 TB of memory combined. When choosing a card, memory capacity comes first, since it determines how large a model and batch size fit; compute throughput, supported numeric precisions, and inter-GPU interconnect bandwidth matter next. Models running on the robot itself typically run on an edge chip like a Jetson rather than on any of these cards.
ExampleThe openpi repository's guidance: running π0 inference needs at least 8 GB of memory and LoRA fine-tuning needs at least 22.5 GB, both of which an RTX 4090 can handle; full-parameter fine-tuning needs at least 70 GB, requiring an 80 GB A100 or an H100.
- Also called
- RTX 4090, A100, H100, B200
- Related
- GPU Memory (VRAM) · NVIDIA Jetson · LoRA · Full Fine-Tuning · Numerical Precision Formats (FP32 / FP16 / BF16 / FP8 / INT8 / INT4) · AutoDL
- Sources
- openpi README(Hardware Requirements)
NVIDIA DGX B200 User Guide: Introduction - As of
- 2026-09