[Paper Review] Modeling the Formation of Social Conventions in Multi-Agent Populations.
This paper proposes a Control-based Reinforcement Learning (CRL) architecture within Distributed Adaptive Control (DAC) theory that integrates reactive sensorimotor control with model-free reinforcement learning to enable multi-agent systems to learn and form social conventions. The CRL framework successfully achieves optimal coordination in game-theoretic tasks, replicating human experimental data on reward efficiency, fairness, and convention stability in both discrete and continuous time.
In order to understand the formation of social conventions we need to know the specific role of control and learning in multi-agent systems. To advance in this direction, we propose, within the framework of the Distributed Adaptive Control (DAC) theory, a novel Control-based Reinforcement Learning architecture (CRL) that can account for the acquisition of social conventions in multi-agent populations that are solving a benchmark social decision-making problem. Our new CRL architecture, as a concrete realization of DAC multi-agent theory, implements a low-level sensorimotor control loop handling the agent's reactive behaviors (pre-wired reflexes), along with a layer based on model-free reinforcement learning that maximizes long-term reward. We apply CRL in a multi-agent game-theoretic task in which coordination must be achieved in order to find an optimal solution. We show that our CRL architecture is able to both find optimal solutions in discrete and continuous time and reproduce human experimental data on standard game-theoretic metrics such as efficiency in acquiring rewards, fairness in reward distribution and stability of convention formation.
Motivation & Objective
- To understand how control and learning mechanisms drive social convention formation in multi-agent systems.
- To address the gap in modeling how agents acquire conventions through coordinated decision-making in dynamic environments.
- To develop a concrete implementation of DAC theory that supports both reactive behavior and long-term reward maximization.
- To validate the model against human behavioral data in standard game-theoretic tasks.
- To demonstrate robustness of convention formation across discrete and continuous time settings.
Proposed method
- The CRL architecture implements a hierarchical control structure with a low-level sensorimotor control loop for reactive behaviors.
- A model-free reinforcement learning layer operates in parallel to maximize long-term cumulative rewards.
- The integration of reactive control and learning enables adaptive responses to environmental feedback.
- The framework is applied to a benchmark multi-agent game-theoretic task requiring coordination for optimal outcomes.
- The system is evaluated using standard game-theoretic metrics such as reward efficiency, fairness, and stability of conventions.
- The architecture is tested in both discrete and continuous time domains to assess generalization.
Experimental results
Research questions
- RQ1How can a control-theoretic framework enable multi-agent systems to learn and form social conventions in coordination tasks?
- RQ2To what extent does the CRL architecture reproduce human experimental data on reward efficiency and fairness?
- RQ3How stable and consistent are the conventions formed by the CRL agents across repeated trials?
- RQ4Can the CRL framework achieve optimal solutions in both discrete and continuous time settings?
- RQ5What role does the integration of reactive control and reinforcement learning play in convention formation?
Key findings
- The CRL architecture successfully achieves optimal solutions in both discrete and continuous time settings.
- The system reproduces human experimental data on efficiency in acquiring rewards, indicating strong behavioral alignment.
- The framework demonstrates fairness in reward distribution, matching observed human tendencies in social coordination tasks.
- Convention formation in the CRL system exhibits high stability over time, consistent with human data.
- The integration of sensorimotor control and reinforcement learning enables robust and adaptive convention acquisition.
- The CRL model generalizes across time domains, showing consistent performance in diverse temporal settings.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.