Embodied AI Glossary中文

vLLM

Advanced

An open-source, high-throughput inference and serving engine for large language models.

vLLM originated at UC Berkeley's Sky Computing Lab; its core technique, PagedAttention, was published at SOSP 2023. During LLM inference, every request has to keep a KV cache (the attention keys and values for content already generated), and its length is unpredictable; the traditional approach reserves memory for the maximum possible length up front, which wastes a great deal of it. PagedAttention borrows the idea of paging from operating systems, allocating the KV cache in small blocks on demand, and combines this with continuous batching (new requests can be inserted into an already-running batch at any time), so a single GPU can serve far more concurrent requests. It supports mainstream open-source LLMs and VLMs and can start an OpenAI-compatible API with one command. In embodied AI it's often used to serve the large model that does task planning, and is also used by frameworks such as veRL as the generation engine for reinforcement learning.

ExampleRun vllm serve on a server to load a Qwen2.5-VL model; the robot sends an image and an instruction over HTTP and gets back a list of decomposed subtasks.

Related
Key-Value Cache · Large Language Model · Vision-Language Model · NVIDIA TensorRT-LLM · veRL (Volcano Engine Reinforcement Learning) · Inference Deployment
Sources
vllm-project/vllm GitHub
Efficient Memory Management for Large Language Model Serving with PagedAttention (arXiv)

See it in the full glossary →