Skip to main content
QUICK REVIEW

[Paper Review] Competitive MA-DRL for Transmit Power Pool Design in Semi-Grant-Free NOMA Systems.

Muhammad Fayaz, Wenqiang Yi|arXiv (Cornell University)|Jun 21, 2021
Advanced Wireless Communication Technologies35 references8 citations
TL;DR

This paper proposes a competitive multi-agent deep reinforcement learning (MA-DRL) framework to design a dynamic transmit power pool (PP) for semi-grant-free non-orthogonal multiple access (SGF-NOMA) in massive IoT networks. By modeling resource selection as a stochastic Markov game and employing Dueling DDQN to enhance learning efficiency, the approach achieves 22.2% higher system throughput than pure grant-free protocols and 17.5% gain over fixed-power control, with reduced training time via invalid action pruning.

ABSTRACT

In this paper, we exploit the capability of multi-agent deep reinforcement learning (MA-DRL) technique to generate a transmit power pool (PP) for Internet of things (IoT) networks with semi-grant-free non-orthogonal multiple access (SGF-NOMA). The PP is mapped with each resource block (RB) to achieve distributed transmit power control (DPC). We first formulate the resource (sub-channel and transmit power) selection problem as stochastic Markov game, and then solve it using two competitive MA-DRL algorithms, namely double deep Q network (DDQN) and Dueling DDQN. Each GF user as an agent tries to find out the optimal transmit power level and RB to form the desired PP. With the aid of dueling processes, the learning process can be enhanced by evaluating the valuable state without considering the effect of each action at each state. Therefore, DDQN is designed for communication scenarios with a small-size action-state space, while Dueling DDQN is for a large-size case. Our results show that the proposed MA-Dueling DDQN based SGF-NOMA with DPC outperforms the SGF-NOMA system with the fixed-power-control mechanism and networks with pure GF protocols with 17.5% and 22.2% gain in terms of the system throughput, respectively. Moreover, to decrease the training time, we eliminate invalid actions (high transmit power levels) to reduce the action space. We show that our proposed algorithm is computationally scalable to massive IoT networks. Finally, to control the interference and guarantee the quality-of-service requirements of grant-based users, we find the optimal number of GF users for each sub-channel.

Motivation & Objective

  • Address the challenge of efficient resource allocation in massive IoT networks with semi-grant-free NOMA by enabling distributed transmit power control.
  • Overcome the limitations of fixed-power control and pure grant-free protocols in maintaining quality-of-service and spectral efficiency.
  • Design a scalable, low-latency power pool mechanism that supports massive connectivity and interference management.
  • Optimize the number of grant-free users per sub-channel to balance throughput and QoS for grant-based users.
  • Enhance learning efficiency in multi-agent environments through dueling network architecture and invalid action elimination.

Proposed method

  • Formulate the joint sub-channel and transmit power selection problem as a stochastic Markov game to model interactions among grant-free users.
  • Implement two competitive MA-DRL algorithms—DDQN and Dueling DDQN—where each grant-free user acts as an independent agent optimizing its own power and sub-channel choice.
  • Integrate dueling networks to decouple state value and advantage estimation, improving learning stability and convergence speed.
  • Reduce the action space by eliminating high transmit power levels that are invalid or impractical, thus accelerating training and improving scalability.
  • Map the resulting optimal power levels to a transmit power pool (PP) per resource block for distributed power control across the network.
  • Determine the optimal number of grant-free users per sub-channel to minimize interference and ensure QoS for grant-based users.

Experimental results

Research questions

  • RQ1How can multi-agent deep reinforcement learning be effectively applied to design a dynamic transmit power pool in semi-grant-free NOMA IoT networks?
  • RQ2What is the impact of using Dueling DDQN versus standard DDQN on learning efficiency and system throughput in large-scale action-state spaces?
  • RQ3To what extent does pruning invalid actions improve training time and scalability in massive IoT deployments?
  • RQ4How does the proposed MA-DRL-based power pool design compare to fixed-power control and pure grant-free protocols in terms of system throughput?
  • RQ5What is the optimal number of grant-free users per sub-channel that maximizes system throughput while maintaining QoS for grant-based users?

Key findings

  • The proposed MA-Dueling DDQN-based SGF-NOMA with distributed power control achieves a 22.2% higher system throughput compared to pure grant-free protocols.
  • The system outperforms fixed-power-control mechanisms by 17.5% in terms of system throughput, demonstrating the benefit of adaptive power allocation.
  • Pruning invalid high-power actions significantly reduces the effective action space, leading to faster training convergence and improved computational scalability.
  • The dueling architecture enhances learning stability and performance, particularly in large-scale action-state environments, by decoupling state value and advantage estimation.
  • The optimal number of grant-free users per sub-channel was determined to balance spectral efficiency and interference, ensuring QoS for grant-based users.
  • The overall framework is computationally scalable and suitable for deployment in massive IoT networks with high user density.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.