Embodied AI Glossary中文

Ray

Ray 分布式框架Advanced

Open-source distributed computing framework for scaling Python programs across many machines and GPUs.

Ray originated at UC Berkeley's RISELab, with its paper published at OSDI in 2018; it's now developed primarily by Anyscale and remains free and open source. It lets an ordinary Python function or class become a parallel unit — a stateless task or a stateful actor — just by adding a decorator, letting a cluster run it in parallel while Ray handles scheduling, data transfer, and fault tolerance. Built on top of this core are libraries such as Ray Train (distributed training), Ray Tune (hyperparameter search), RLlib (reinforcement learning), Ray Data, and Ray Serve. Reinforcement learning needs to schedule large numbers of simulation, inference, and training processes at once, and large-model RL frameworks such as veRL use Ray to orchestrate these components.

ExampleWhen fine-tuning a VLA with reinforcement learning, use Ray actors to run simulation rollout, policy inference, and parameter updates on separate GPUs, all coordinated by one driver process.

Related
veRL (Volcano Engine Reinforcement Learning) · Distributed Training · Actor-Learner Architecture (Distributed RL) · Slurm Workload Manager · Massively Parallel Reinforcement Learning
Sources
Ray 官网 (Chinese)
Ray: A Distributed Framework for Emerging AI Applications (arXiv)

See it in the full glossary →