[Paper Review] A Survey of Learning in Multiagent Environments: Dealing with Non-Stationarity
This survey reviews how learning in multiagent environments copes with non-stationarity and introduces a five-category framework for categorizing approaches.
The key challenge in multiagent learning is learning a best response to the behaviour of other agents, which may be non-stationary: if the other agents adapt their strategy as well, the learning target moves. Disparate streams of research have approached non-stationarity from several angles, which make a variety of implicit assumptions that make it hard to keep an overview of the state of the art and to validate the innovation and significance of new works. This survey presents a coherent overview of work that addresses opponent-induced non-stationarity with tools from game theory, reinforcement learning and multi-armed bandits. Further, we reflect on the principle approaches how algorithms model and cope with this non-stationarity, arriving at a new framework and five categories (in increasing order of sophistication): ignore, forget, respond to target models, learn models, and theory of mind. A wide range of state-of-the-art algorithms is classified into a taxonomy, using these categories and key characteristics of the environment (e.g., observability) and adaptation behaviour of the opponents (e.g., smooth, abrupt). To clarify even further we present illustrative variations of one domain, contrasting the strengths and limitations of each category. Finally, we discuss in which environments the different approaches yield most merit, and point to promising avenues of future research.
Motivation & Objective
- Synthesize how opponent-induced non-stationarity is handled across bandits, reinforcement learning, and game theory.
- Introduce a coherent framework to categorize non-stationarity handling in multiagent learning.
- Classify state-of-the-art algorithms using environment and opponent adaptation factors.
- Discuss strengths, limitations, and future research directions in non-stationary multiagent learning.
Proposed method
- Review formal models from multi-armed bandits, reinforcement learning, and game theory to frame non-stationarity.
- Propose a new framework with five categories for handling non-stationarity: ignore, forget, respond to target models, learn models, theory of mind.
- Illustrate categories with domain examples to highlight strengths and limitations.
- Provide a taxonomy of algorithms by category and environment/adaptation characteristics.
- Discuss open questions and future research directions.
Experimental results
Research questions
- RQ1How does non-stationarity arise in multiagent learning across different domains (bandits, RL, game theory)?
- RQ2What framework best captures the progression of sophistication in handling non-stationarity?
- RQ3Which algorithms align with which categories under various observability and opponent-adaptation assumptions?
- RQ4What are the key open problems and promising avenues for future work in non-stationary multiagent learning.
Key findings
- Proposes a five-category framework for coping with non-stationarity: ignore, forget, respond to target opponents, learn opponent models, and theory of mind.
- Provides a taxonomy classifying state-of-the-art algorithms from MABs, RL, and game theory by category and environment/opp. adaptation.
- Uses illustrative variations to contrast strengths and limitations of each category.
- Analyzes environments where different approaches yield the most benefit and outlines promising future research directions.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.