Part VI: Distributed AI and Multi-Agent Systems
Chapter 30: Multi-Agent Reinforcement Learning

Multi-Agent Reinforcement Learning

Chapter 28 gave you the strategic mathematics of agents with goals of their own, and Chapter 29 built the running societies in which those agents perceive, speak, and act, but in both the behavior was handed to the agent by a designer: a payoff matrix to optimize, a protocol to follow, an architecture to embody. This chapter takes the designer out of the loop. It lets the agents learn. Multi-agent reinforcement learning is what happens when each agent in a shared world discovers its own policy through trial and reward instead of having that policy written for it, and the moment more than one agent learns at once the ground shifts under everyone's feet, because the environment each agent is trying to master now contains other agents who are themselves changing. Ten sections build the discipline from that single complication. They begin by carrying single-agent reinforcement learning across the bridge into many agents, then formalize the setting as a Markov game that generalizes both the Markov decision process and the matrix game of Chapter 28, then sort the field by whether the agents share a goal, oppose one another, or sit somewhere in between. From there the sections work through the algorithmic toolkit the field has assembled to cope with learning amid other learners: independent learners that ignore the problem and sometimes get away with it, centralized training with decentralized execution that uses global information at training time and discards it at deployment, value decomposition that factors a team's joint value into per-agent pieces, policy gradient methods that scale to continuous and high-dimensional action spaces, and the credit-assignment machinery that tells one agent how much of a shared reward it actually earned. The two hardest threads, non-stationarity and the distributed systems that train these populations at scale, get sections of their own. The thread that runs through all ten is the one game theory named and reinforcement learning must now survive: every other agent is part of your environment, and your environment will not hold still. Where Chapter 29 built agents whose coordination you designed, this chapter builds agents that learn their coordination, and it does so on the distributed reinforcement learning infrastructure of Chapter 20, now driving not one policy but a whole society of them.

Conceptual illustration for Chapter 30: Multi-Agent Reinforcement Learning

"I trained my policy to perfection against a fixed opponent. Then the opponent started learning too, and every lesson I had carefully memorized became wrong at exactly the same speed I learned it."

A Q-Function Chasing a Moving Target

Chapter Overview

This is the learning chapter of Part VI, and its subject is the gap between an agent that follows a designed policy and an agent that discovers one through reward, in a world where it is not the only one learning. The ten sections develop that subject in order. They begin by carrying single-agent reinforcement learning into the multi-agent setting and naming what breaks, then formalize the setting as a Markov game, then sort the field by the structure of the agents' rewards. From there they build the algorithmic toolkit: independent learning, centralized training with decentralized execution, value decomposition, policy gradients, and credit assignment, before confronting non-stationarity head-on and closing with the distributed systems that train these populations at scale. The through-line is the moving target. Every agent learns against an environment that includes other learners, so the optimization problem each one faces refuses to stand still.

Read in order, the ten sections take you from "here is reinforcement learning for one agent" to a working understanding of how a population of agents learns to act together or against each other: carry RL across the bridge, formalize the game, classify the reward structure, try the naive thing, centralize the training but decentralize the execution, decompose the value, scale with policy gradients, solve credit assignment, confront the moving target, and run the whole thing across a cluster. The argument is cumulative and it carries the strategic toolkit of Chapter 28 into learning dynamics: equilibria become the solution concepts that learning may or may not reach, and the cooperative and competitive games become the reward structures that decide which algorithms apply. The agent societies of Chapter 29 become populations that learn rather than follow scripts, and the thread runs straight out of the chapter into the collective behavior of Chapter 31, where coordination emerges from many simple agents rather than from learned policies over joint state.

Prerequisites

This chapter stands on the two chapters immediately before it in Part VI and on the reinforcement learning foundations the rest of the book assumes. From Chapter 28: Game-Theoretic Foundations for Multi-Agent AI you carry the strategic vocabulary that frames every learning problem here: what a game is, what an equilibrium means, the difference between cooperative and competitive interaction, and why an agent's best response depends on what everyone else does, because the Markov game of Section 30.2 is exactly the dynamic, state-bearing generalization of those matrix games. From Chapter 29: Multi-Agent Systems you carry the engineering picture of a society of autonomous agents acting on partial views inside a shared environment, since MARL is what happens when those agents learn their behaviors rather than have them designed. The chapter also assumes basic single-agent reinforcement learning: Markov decision processes, value functions, Q-learning, policy gradients, and the actor-critic pattern, the machinery this chapter generalizes from one agent to many. For the systems half of the chapter, it builds directly on Chapter 20: Distributed Reinforcement Learning Infrastructure, whose actor-learner architectures, distributed experience collection, and replay systems return in Section 30.10 as the engine that trains multi-agent populations at scale. Light formal notation and the ability to read pseudocode are assumed throughout, as in the rest of the book. No prior exposure to multi-agent learning is required; Section 30.1 builds the multi-agent setting from the single-agent case before anything is built on it.

Learning Objectives

Chapter Roadmap

What's Next?

This chapter built the learning core of Part VI: how an agent discovers a policy by trial and reward when its environment contains other agents who are learning too, from the Markov game that formalizes the setting through value decomposition, policy gradients, credit assignment, and the distributed systems that train it all at scale. The agents here learn rich policies over joint state, and the central struggle is the non-stationarity that learning amid other learners creates. Chapter 31: Swarm Intelligence and Collective Behavior turns the problem inside out. Instead of a small number of agents each learning a sophisticated policy, it asks how coordinated, robust, global behavior can emerge from very many agents following very simple local rules, with no learned value function and no joint optimization at all. Where this chapter gave you populations that learn their coordination through reward, the next gives you populations whose coordination is an emergent property of simple interaction, flocking, stigmergy, and self-organization. Read Chapter 31 next, and watch coordination appear from the bottom up rather than being learned from the top down.

Bibliography & Further Reading

Foundations and Frameworks

Littman, M. L. "Markov Games as a Framework for Multi-Agent Reinforcement Learning." Proceedings of the Eleventh International Conference on Machine Learning (ICML), 1994. sciencedirect.com

The paper that cast multi-agent reinforcement learning as learning in a Markov game and introduced minimax-Q, the formal foundation for the Markov game setting of Section 30.2.

📄 Paper

Tan, M. "Multi-Agent Reinforcement Learning: Independent vs. Cooperative Agents." Proceedings of the Tenth International Conference on Machine Learning (ICML), 1993. mit.edu

The early study contrasting independent Q-learners with cooperative agents that share information, the historical anchor for the independent-learning baseline of Section 30.4.

📄 Paper

Centralized Training and Policy Gradients

Lowe, R., Wu, Y., Tamar, A., Harb, J., Abbeel, P., Mordatch, I. "Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments (MADDPG)." arXiv:1706.02275, 2017. arxiv.org

The multi-agent actor-critic with centralized critics over the joint action that defined centralized training with decentralized execution, central to Sections 30.5 and 30.7.

📄 Paper

Foerster, J., Farquhar, G., Afouras, T., Nardelli, N., Whiteson, S. "Counterfactual Multi-Agent Policy Gradients (COMA)." arXiv:1705.08926, 2017. arxiv.org

The counterfactual baseline that isolates each agent's contribution to a shared reward, the centerpiece of the credit-assignment treatment in Section 30.8.

📄 Paper

Yu, C., Velu, A., Vinitsky, E., Gao, J., Wang, Y., Bayen, A., Wu, Y. "The Surprising Effectiveness of PPO in Cooperative Multi-Agent Games (MAPPO)." arXiv:2103.01955, 2021. arxiv.org

The study showing that a centralized-critic PPO is a strong, simple baseline across cooperative benchmarks, a workhorse method for the policy gradients of Section 30.7.

📄 Paper

de Witt, C. S., Gupta, T., Makoviichuk, D., Makoviychuk, V., Torr, P. H. S., Sun, M., Whiteson, S. "Is Independent Learning All You Need in the StarCraft Multi-Agent Challenge? (IPPO)." arXiv:2011.09533, 2020. arxiv.org

The result that independent PPO learners are competitive with centralized methods on hard cooperative tasks, sharpening the independent-learner discussion of Section 30.4.

📄 Paper

Value Decomposition

Sunehag, P., Lever, G., Gruslys, A., Czarnecki, W. M., et al. "Value-Decomposition Networks for Cooperative Multi-Agent Learning (VDN)." arXiv:1706.05296, 2017. arxiv.org

The additive factorization of a team's joint action-value into per-agent terms, the first value-decomposition method developed in Section 30.6.

📄 Paper

Rashid, T., Samvelyan, M., de Witt, C. S., Farquhar, G., Foerster, J., Whiteson, S. "QMIX: Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement Learning." arXiv:1803.11485, 2018. arxiv.org

The monotonic mixing network that generalizes additive decomposition while keeping decentralized greedy action consistent with the centralized value, central to Section 30.6.

📄 Paper

Benchmarks and Social Dilemmas

Samvelyan, M., Rashid, T., de Witt, C. S., Farquhar, G., et al. "The StarCraft Multi-Agent Challenge (SMAC)." arXiv:1902.04043, 2019. arxiv.org

The cooperative micromanagement benchmark on which value decomposition and centralized-critic methods are standardly evaluated, the reference testbed for Sections 30.5 through 30.8.

📄 Paper

Leibo, J. Z., Zambaldi, V., Lanctot, M., Marecki, J., Graepel, T. "Multi-Agent Reinforcement Learning in Sequential Social Dilemmas." arXiv:1702.03037, 2017. arxiv.org

The study of how cooperation and defection emerge among learning agents in mixed-motive games, the empirical backbone of the mixed-settings discussion in Section 30.3.

📄 Paper

Large-Scale Systems

Vinyals, O., Babuschkin, I., Czarnecki, W. M., et al. "Grandmaster Level in StarCraft II Using Multi-Agent Reinforcement Learning (AlphaStar)." Nature, 575, 2019. nature.com

The system that reached grandmaster play through league-based multi-agent training, a landmark for the distributed competitive training discussed in Section 30.10.

📄 Paper

OpenAI, Berner, C., Brockman, G., et al. "Dota 2 with Large Scale Deep Reinforcement Learning (OpenAI Five)." arXiv:1912.06680, 2019. arxiv.org

The large-scale self-play system that trained a five-agent team for months across thousands of machines, the flagship example of the distributed MARL training of Section 30.10.

📄 Paper