Embodied AI Glossary中文

Multi-Agent Reinforcement Learning

多智能体强化学习MARLAdvanced

Reinforcement learning where multiple agents learn at once in a shared environment, cooperating or competing.

Multi-agent reinforcement learning studies settings where several learning decision-makers share one environment; based on the reward relationship, it splits into fully cooperative (a shared reward), fully competitive (a zero-sum game, like chess), and mixed settings, such as self-driving cars that each want to reach their own destination while all avoiding collisions. The biggest difficulty compared to single-agent RL is non-stationarity: every other agent keeps changing its policy too, so from any one agent's perspective the environment never stops shifting, which breaks the convergence guarantees single-agent algorithms rely on; credit assignment and partial observability are additional challenges. A common approach is “centralized training, decentralized execution”: during training a critic sees global information, while at execution time each agent sees only its own local observation, as in 2017's MADDPG and 2021's MAPPO. Multi-robot collaboration, robot soccer, and adversarial self-play training all draw on this field.

ExampleMAPPO controls each unit in a battle with its own PPO-based agent, and reaches results on par with or better than off-policy methods on benchmarks such as the StarCraft Multi-Agent Challenge (SMAC) and Google Research Football.

Also called
MARL
Related
Reinforcement Learning · Multi-robot Collaboration · Self-Play · Swarm Intelligence · Partially Observable Markov Decision Process · Proximal Policy Optimization
Sources
Albrecht, Christianos, Schäfer: Multi-Agent Reinforcement Learning: Foundations and Modern Approaches (MIT Press, 2024)
Wikipedia: Multi-agent reinforcement learning
Yu et al. 2021: The Surprising Effectiveness of PPO in Cooperative Multi-Agent Games (MAPPO)

See it in the full glossary →