Embodied AI Glossary中文

Slurm Workload Manager

Slurm 集群调度Advanced

The scheduling system that queues jobs and allocates GPUs on a shared compute cluster.

Slurm is an open-source Linux cluster job-scheduling system, originally from Lawrence Livermore National Laboratory in the US and now maintained primarily by the company SchedMD; it's used by a large share of supercomputing centers, universities, and companies' GPU clusters. When many people share a pool of machines, instead of logging into a specific machine and running a job directly, everyone submits jobs to Slurm, which queues and allocates nodes and GPUs by resource availability and priority. Common commands include sbatch (submit a script), srun (run directly), squeue (check the queue), and scancel (cancel a job). Training large models like VLAs, or doing multi-node, multi-GPU distributed training, can hardly avoid it.

ExampleWrite a train.sh that declares a need for 2 nodes with 8 GPUs each, then submit it with sbatch train.sh and check whether it's been scheduled with squeue.

Also called
Slurm, SLURM
Related
Distributed Training · Distributed Data Parallel (DDP) · Ray · Docker · Secure Shell (SSH)
Sources
Slurm Workload Manager - Documentation

See it in the full glossary →