[Paper Review] Social diversity and social preferences in mixed-motive reinforcement learning
This paper introduces Social Value Orientation (SVO) into multi-agent reinforcement learning to model diverse social preferences in mixed-motive games. By endowing agents with varying SVOs—ranging from selfish to prosocial—the study demonstrates that heterogeneous populations outperform homogeneous ones in terms of collective performance and policy generalization, particularly under equality-sensitive metrics.
Recent research on reinforcement learning in pure-conflict and pure-common interest games has emphasized the importance of population heterogeneity. In contrast, studies of reinforcement learning in mixed-motive games have primarily leveraged homogeneous approaches. Given the defining characteristic of mixed-motive games--the imperfect correlation of incentives between group members--we study the effect of population heterogeneity on mixed-motive reinforcement learning. We draw on interdependence theory from social psychology and imbue reinforcement learning agents with Social Value Orientation (SVO), a flexible formalization of preferences over group outcome distributions. We subsequently explore the effects of diversity in SVO on populations of reinforcement learning agents in two mixed-motive Markov games. We demonstrate that heterogeneity in SVO generates meaningful and complex behavioral variation among agents similar to that suggested by interdependence theory. Empirical results in these mixed-motive dilemmas suggest agents trained in heterogeneous populations develop particularly generalized, high-performing policies relative to those trained in homogeneous populations.
Motivation & Objective
- To investigate the impact of social preference diversity on multi-agent reinforcement learning in mixed-motive environments.
- To address the lack of attention to opponent heterogeneity in mixed-motive Markov games, unlike in pure-conflict or pure-common interest settings.
- To model human-like social preferences using Social Value Orientation (SVO), a formalization of intrinsic motivations for reward distribution.
- To evaluate whether SVO diversity leads to more robust, high-performing, and generalized agent policies compared to homogeneous populations.
- To bridge multi-agent RL with social psychology by grounding agent motivations in interdependence theory and SVO frameworks.
Proposed method
- Endowed reinforcement learning agents with Social Value Orientation (SVO), a parameterized representation of preferences over self and other outcomes.
- Applied SVO to two mixed-motive Markov games: HarvestPatch and a modified Prisoner’s Dilemma variant.
- Trained agents using multi-agent reinforcement learning algorithms with SVO as an intrinsic motivation component.
- Varied SVO distributions across populations to create homogeneous (e.g., all altruistic) and heterogeneous (e.g., mixed selfish, prosocial, competitive) training regimes.
- Used outcome transformation processes from interdependence theory to map raw payoffs into effective matrices reflecting subjective preferences.
- Evaluated performance using both collective reward and equality-sensitive metrics, such as the Gini coefficient of reward distribution.
Experimental results
Research questions
- RQ1How does population heterogeneity in social preferences (SVO) affect learning outcomes in mixed-motive Markov games?
- RQ2Can SVO-based diversity in intrinsic motivations lead to higher collective performance than homogeneous SVO populations?
- RQ3Do agents trained in heterogeneous SVO populations develop more generalized and robust policies compared to those in homogeneous settings?
- RQ4To what extent do SVO-based preferences align with predictions from interdependence theory in multi-agent interactions?
- RQ5Can SVO diversity mitigate the inefficiencies that arise when both agents adopt prosocial transformations simultaneously?
Key findings
- Heterogeneous SVO populations achieved significantly higher collective returns than homogeneous altruistic populations, particularly under equality-sensitive metrics.
- Agents trained in heterogeneous populations developed more generalized policies, demonstrating robustness to diverse opponent behaviors.
- In the HarvestPatch environment, agents with higher SVO were more likely to clean the river even when apples remained unharvested, indicating prosocial coordination.
- Homogeneous altruistic populations produced hyper-specialized agents that exploited either intrinsic or extrinsic motivations, leading to suboptimal group outcomes.
- The results align with interdependence theory’s prediction that mutual prosocial transformations can lead to deficient outcomes when both players adopt them.
- SVO diversity enabled populations to avoid the inefficiencies of symmetric prosociality, supporting the emergence of stable, high-performing conventions.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.